URLs and redirects

Serve the merchant's own URLs, keep a listing's state in its URL, and carry the old shop's links over.

OrbSearch imposes no URL: the catalog and routes.json say where each page sits. In the home store:

PageGermanEnglish
The category wohnen/sofas (living/sofas)/wohnen/sofas/en/living/sofas
The product SOFA-ASKA-2/p/aska-2-sitzer-sofa/en/p/aska-two-seater-sofa
Search/suche/en/search

A category's path and a product's url come from the catalog, the search page's path from routes.json. A brand's page sits at the brand's url while the release has a brand facet.

A listing's state in its URL

A listing keeps its query, filters, sort and page in its URL, so a link, the back button and a search engine see what the shopper sees. Each state has one spelling, and any other answers 301 to it:

Terminal
for url in '/wohnen/sofas?color=Grau' '/wohnen/sofas?COLOR=grey' '/wohnen/sofas?width.lte=220&color=grey'; do
  curl -s -G http://127.0.0.1:7800/v1/resolve --data-urlencode "url=$url" -d results=0 | jq -c '{status, location}'
done
200 · 1.5 ms · release b25a6aaa
{"status":301,"location":"/wohnen/sofas?color=grey"}
{"status":301,"location":"/wohnen/sofas?color=grey"}
{"status":301,"location":"/wohnen/sofas?color=grey&width.lte=220"}

A value reads by its label, slug or alias in any case: Grau is grey's German label. Filters stand in alphabetical order with their values sorted. A range is the field with .gte and .lte, a toggle 1.

A parameter nothing reads, such as utm_source, is ignored, and the 301 keeps it. A value the page cannot read, such as an unknown sort, leads to the URL without it. The reserved q, sort, page, from, ignore and redirect may stand anywhere, so the client's links need no redirect.

Reference: parameters

Resolve a URL

Ask resolve before rendering a byte. It says what the page is, the status to answer with, its robots and, for a listing, the search it runs:

Terminal
curl -s -G http://127.0.0.1:7800/v1/resolve --data-urlencode 'url=/wohnen/sofas?color=grey' -d results=0 \
  | jq -c '{type, status, robots, entity}, .breadcrumbs, .request'
200 · 3.0 ms · release b25a6aaa
{"type":"category","status":200,"robots":"noindex,follow","entity":"category:wohnen/sofas"}
[{"title":"Wohnen","url":"/wohnen"},{"title":"Sofas","url":"/wohnen/sofas"}]
{"on":"category_page","channel":"de","locale":"de","category":"wohnen/sofas","filters":{"color":"grey"}}

The call itself answers 200; status is the storefront's to answer with. Without results=0, the listing's results come in the same call. A filter state such as grey is noindex,follow; the sofa page alone is index,follow, with its path as canonical.

Resolve tries the legacy table, the search page, categories and brands, filter pages, then products. It answers one of these types:

  • category or brand: a listing, with its request, results and a category's breadcrumbs.
  • search: the search page, always noindex,follow.
  • product: a product's page, with product as GET /v1/product answers it.
  • redirect: location names the final URL, never a chain.
  • gone: 410, a URL removed on purpose.
  • not_found: 404, no page OrbSearch knows.

Once its results are known, a listing past its last page, or an indexable one that shows nothing, answers 404 with its body. A search whose query names a whole page, such as /suche?q=graues+sofa (grey sofa), redirects there with 302.

Reference: GET /v1/resolve · type · robots

routes.json

routes.json spells the merchant's URLs, and every field is optional. The home store's names only its search page:

tests/fixtures/home/routes.json (trimmed)
{
    "search": { "de": "/suche", "en": "/en/search" }
}

Other fields rename parameters, such as color to farbe, set trailing slashes, and say where a search lands that names a whole page.

Reference: routes.json · search · landing · trailing_slash

Filter pages

A filter page gives a combination of filters a clean path that a search engine can rank. The home store lists none; with these lines in routes.json:

routes.json
{
    "filter_pages": {
        "facets": ["brand", "color"],
        "combinations": [["brand"], ["color"], ["brand", "color"]],
        "before_category": { "brand": "marken" }
    }
}

the sofas get these URLs:

The sofas withOne URL
Halvard/marken/halvard/wohnen/sofas
grey/wohnen/sofas/color--grau
Halvard in grey/marken/halvard/wohnen/sofas/color--grau
grey or anthracite/wohnen/sofas?color=anthracite&color=grey

A segment is the parameter's name, -- and the value's slug in the page's locale; marken means brands. Other spellings of these states answer 301 to the path. A filter page with fewer than min_products products, 4 by default, stays noindex, as the two grey sofas do. POST /api/v1/preview/page tries a URL under a draft's files before you publish them.

Reference: filter_pages · min_products

redirects.json

redirects.json carries the old shop's URLs over. It is part of the release and held in memory, so an old URL costs no engine call:

redirects.json
{
    "redirects": [
        { "from": "/sofas.html", "to": "category:wohnen/sofas" },
        { "from": "/index.php", "query": { "cat": "12" }, "to": "category:wohnen" },
        { "from": "/outlet-2019", "status": 410 },
        { "pattern": "^/kategorie/(.+?)/?$", "url": "/wohnen/$1" }
    ]
}

/sofas.html answers 301 to /wohnen/sofas, wherever the catalog later moves the sofa page, and /outlet-2019 answers 410. Exact entries answer first, then the patterns in order, and a chain of old URLs is followed to its end.

Reference: redirects.json · from · pattern · status

A storefront on resolve

One catch-all route behind the shop's own pages, such as its cart, answers every other path:

app/routes/page.tsx
import { ,  } from "@orbsearch/client";
import { , , type , type  } from "react-router";

const  = "https://shop.example";

export async function ({ ,  }: ) {
    const  = ..("user-agent");
    const  = ({ : "https://search.shop.example", ...( && {  }) });
    const  = await .(. + ., { : "de" }, { : . });

    const  = ();
    if ( !== null) throw (, .);
    if (. === "gone" || . === "not_found") throw (null, { : . });
    // A listing or a product, with 404 for an empty indexable listing or a page past its last.
    return (, { : . });
}

export const : <typeof > = ({ :  }) => [
    ...(?. ? [{ : "robots", : . }] : []),
    ...(?. ? [{ : "link", : "canonical", :  + . }] : []),
];

locationOf(page) percent-encodes a redirect's path, such as /möbel, for the Location header. A listing renders from page.results, a product from page.product, with the React components or your own.

Robots

GET /v1/robots answers the Disallow lines that keep crawlers off every state the release does not make indexable:

Terminal
curl -s http://127.0.0.1:7800/v1/robots | jq -r '.lines[] | select(test("\\*&|color=|width|suche"))'
200 · 2.8 ms · release b25a6aaa
Disallow: /*?*&
Disallow: /*?color=
Disallow: /*?width.
Disallow: /*?width=
Disallow: /suche$
Disallow: /suche?

Any URL with a second parameter, each parameter no indexable page uses, and the search page are off limits; ?page=N alone stays crawlable. Serve the lines in the shop's own robots.txt, under User-agent: *.

Reference: GET /v1/robots

Limits today

  • URLs are paths. The storefront puts its own origin in front, as above. Listings carry no hreflang alternates; a product carries its own in alternates.
  • A product's URL matches only as the catalog writes it, not in another case or with a trailing slash.
  • The search page wins where its path is also a category's.
  • The client writes parameters by their codes, onto the page's own path. So a storefront on the React components keeps parameters and filter_pages empty: a renamed parameter loses a choice after one click, and a filter page's own filter cannot be taken back.

Next

On this page