URLs and redirects
Serve the merchant's own URLs, keep a listing's state in its URL, and carry the old shop's links over.
OrbSearch imposes no URL: the catalog and routes.json say where each page sits. In the home store:
| Page | German | English |
|---|---|---|
The category wohnen/sofas (living/sofas) | /wohnen/sofas | /en/living/sofas |
The product SOFA-ASKA-2 | /p/aska-2-sitzer-sofa | /en/p/aska-two-seater-sofa |
| Search | /suche | /en/search |
A category's path and a product's url come from the catalog, the search page's path from routes.json. A brand's page sits at the brand's url while the release has a brand facet.
A listing's state in its URL
A listing keeps its query, filters, sort and page in its URL, so a link, the back button and a search engine see what the shopper sees. Each state has one spelling, and any other answers 301 to it:
for url in '/wohnen/sofas?color=Grau' '/wohnen/sofas?COLOR=grey' '/wohnen/sofas?width.lte=220&color=grey'; do
curl -s -G http://127.0.0.1:7800/v1/resolve --data-urlencode "url=$url" -d results=0 | jq -c '{status, location}'
done{"status":301,"location":"/wohnen/sofas?color=grey"}
{"status":301,"location":"/wohnen/sofas?color=grey"}
{"status":301,"location":"/wohnen/sofas?color=grey&width.lte=220"}A value reads by its label, slug or alias in any case: Grau is grey's German label. Filters stand in alphabetical order with their values sorted. A range is the field with .gte and .lte, a toggle 1.
A parameter nothing reads, such as utm_source, is ignored, and the 301 keeps it. A value the page cannot read, such as an unknown sort, leads to the URL without it. The reserved q, sort, page, from, ignore and redirect may stand anywhere, so the client's links need no redirect.
Reference: parameters
Resolve a URL
Ask resolve before rendering a byte. It says what the page is, the status to answer with, its robots and, for a listing, the search it runs:
curl -s -G http://127.0.0.1:7800/v1/resolve --data-urlencode 'url=/wohnen/sofas?color=grey' -d results=0 \
| jq -c '{type, status, robots, entity}, .breadcrumbs, .request'{"type":"category","status":200,"robots":"noindex,follow","entity":"category:wohnen/sofas"}
[{"title":"Wohnen","url":"/wohnen"},{"title":"Sofas","url":"/wohnen/sofas"}]
{"on":"category_page","channel":"de","locale":"de","category":"wohnen/sofas","filters":{"color":"grey"}}The call itself answers 200; status is the storefront's to answer with. Without results=0, the listing's results come in the same call. A filter state such as grey is noindex,follow; the sofa page alone is index,follow, with its path as canonical.
Resolve tries the legacy table, the search page, categories and brands, filter pages, then products. It answers one of these types:
categoryorbrand: a listing, with itsrequest,resultsand a category'sbreadcrumbs.search: the search page, alwaysnoindex,follow.product: a product's page, withproductasGET /v1/productanswers it.redirect:locationnames the final URL, never a chain.gone:410, a URL removed on purpose.not_found:404, no page OrbSearch knows.
Once its results are known, a listing past its last page, or an indexable one that shows nothing, answers 404 with its body. A search whose query names a whole page, such as /suche?q=graues+sofa (grey sofa), redirects there with 302.
Reference: GET /v1/resolve · type · robots
routes.json
routes.json spells the merchant's URLs, and every field is optional. The home store's names only its search page:
{
"search": { "de": "/suche", "en": "/en/search" }
}Other fields rename parameters, such as color to farbe, set trailing slashes, and say where a search lands that names a whole page.
Reference: routes.json · search · landing · trailing_slash
Filter pages
A filter page gives a combination of filters a clean path that a search engine can rank. The home store lists none; with these lines in routes.json:
{
"filter_pages": {
"facets": ["brand", "color"],
"combinations": [["brand"], ["color"], ["brand", "color"]],
"before_category": { "brand": "marken" }
}
}the sofas get these URLs:
| The sofas with | One URL |
|---|---|
| Halvard | /marken/halvard/wohnen/sofas |
| grey | /wohnen/sofas/color--grau |
| Halvard in grey | /marken/halvard/wohnen/sofas/color--grau |
| grey or anthracite | /wohnen/sofas?color=anthracite&color=grey |
A segment is the parameter's name, -- and the value's slug in the page's locale; marken means brands. Other spellings of these states answer 301 to the path. A filter page with fewer than min_products products, 4 by default, stays noindex, as the two grey sofas do. POST /api/v1/preview/page tries a URL under a draft's files before you publish them.
Reference: filter_pages · min_products
redirects.json
redirects.json carries the old shop's URLs over. It is part of the release and held in memory, so an old URL costs no engine call:
{
"redirects": [
{ "from": "/sofas.html", "to": "category:wohnen/sofas" },
{ "from": "/index.php", "query": { "cat": "12" }, "to": "category:wohnen" },
{ "from": "/outlet-2019", "status": 410 },
{ "pattern": "^/kategorie/(.+?)/?$", "url": "/wohnen/$1" }
]
}/sofas.html answers 301 to /wohnen/sofas, wherever the catalog later moves the sofa page, and /outlet-2019 answers 410. Exact entries answer first, then the patterns in order, and a chain of old URLs is followed to its end.
Reference: redirects.json · from · pattern · status
A storefront on resolve
One catch-all route behind the shop's own pages, such as its cart, answers every other path:
import { , } from "@orbsearch/client";
import { , , type , type } from "react-router";
const = "https://shop.example";
export async function ({ , }: ) {
const = ..("user-agent");
const = ({ : "https://search.shop.example", ...( && { }) });
const = await .(. + ., { : "de" }, { : . });
const = ();
if ( !== null) throw (, .);
if (. === "gone" || . === "not_found") throw (null, { : . });
// A listing or a product, with 404 for an empty indexable listing or a page past its last.
return (, { : . });
}
export const : <typeof > = ({ : }) => [
...(?. ? [{ : "robots", : . }] : []),
...(?. ? [{ : "link", : "canonical", : + . }] : []),
];locationOf(page) percent-encodes a redirect's path, such as /möbel, for the Location header. A listing renders from page.results, a product from page.product, with the React components or your own.
Robots
GET /v1/robots answers the Disallow lines that keep crawlers off every state the release does not make indexable:
curl -s http://127.0.0.1:7800/v1/robots | jq -r '.lines[] | select(test("\\*&|color=|width|suche"))'Disallow: /*?*&
Disallow: /*?color=
Disallow: /*?width.
Disallow: /*?width=
Disallow: /suche$
Disallow: /suche?Any URL with a second parameter, each parameter no indexable page uses, and the search page are off limits; ?page=N alone stays crawlable. Serve the lines in the shop's own robots.txt, under User-agent: *.
Reference: GET /v1/robots
Limits today
- URLs are paths. The storefront puts its own origin in front, as above. Listings carry no
hreflangalternates; a product carries its own inalternates. - A product's URL matches only as the catalog writes it, not in another case or with a trailing slash.
- The search page wins where its path is also a category's.
- The client writes parameters by their codes, onto the page's own path. So a storefront on the React components keeps
parametersandfilter_pagesempty: a renamed parameter loses a choice after one click, and a filter page's own filter cannot be taken back.
Next
- Category pages: what a category's listing answers.
- Product pages: what a product's page answers.
- React components: pages that render resolve's answer.