Formats we read
The exports OrbSearch reads as your shop writes them, how it recognizes each one, and OrbSearch JSON Lines, its own format.
OrbSearch reads the export your shop already writes, and recognizes its format from the file's first bytes. Here is the home store's export as a CSV, an Excel workbook and a Merchant Center feed, each sent with the same content type:
for file in export.csv export.xlsx export.xml; do
draft=$(curl -s -X POST -H "Authorization: Bearer $key" -H 'Content-Type: application/octet-stream' \
--data-binary @tests/fixtures/feeds/home/$file "$api/imports?name=$file" | jq -r .import.id)
curl -s -H "Authorization: Bearer $key" $api/imports/$draft/inspection | jq -c '{format: .feed.format, products}'
curl -s -X DELETE -H "Authorization: Bearer $key" $api/imports/$draft >/dev/null
done{"format":{"kind":"csv","delimiter":";","encoding":"windows-1252","header":1},"products":23}
{"format":{"kind":"xlsx","sheet":"Artikel","header":3},"products":23}
{"format":{"kind":"xml","encoding":"utf-8"},"products":23}Each file becomes an import that changes nothing live, and goes again. All three make the same 23 products, though each was sent as application/octet-stream: the content type only names the file.
The CSV is what Excel on Windows writes for German: semicolons and Windows-1252. The workbook's column names sit in row 3 of its sheet "Artikel" (articles), under two title rows.
What it reads
Each format is told by its first bytes:
- Excel, an
.xlsxworkbook, is a zip. Its first sheet is read, unlessfeed.jsonnames another. - Google Merchant Center XML starts with
<: RSS 2.0 with an<item>per product, or Atom with an<entry>.g:pricereads as the columnprice. - JSON starts with
[: an array of records in a shape of your own. A nested object becomes columns such asdimensions.width. - OrbSearch JSON Lines starts with
{: one entity per line, below. - CSV, TSV and TXT are anything else, delimited by comma, semicolon, tab or pipe, in UTF-8, UTF-16 or Windows-1252. Merchant Center's TSV is one of them.
The delimiter comes from the first 50 rows, the encoding from a byte order mark or the bytes themselves, and the column names from the first row that looks like a header. Any format may arrive gzipped, up to 2 GB unpacked.
Reference: format.kind · POST /api/v1/imports
Pin what it would guess wrong
Where recognizing would guess wrong, the release's feed.json pins the format:
{ "format": { "kind": "json" } }JSON records written one object per line need this pin: they start with {, so they would read as OrbSearch JSON Lines. An .xls workbook from before Excel 2007 is not read; save it as .xlsx.
Reference: format
OrbSearch JSON Lines
The format of its own holds one entity per line, as JSON, and needs no mapping. A category, a product and a guide of the home store:
{"type":"category","id":"wohnen/sofas","parent":"wohnen","title":{"de":"Sofas","en":"Sofas"},"path":{"de":"/wohnen/sofas","en":"/en/living/sofas"}}
{"type":"product","id":"SOFA-ASKA-2","title":{"de":"Aska 2-Sitzer-Sofa aus Samt","en":"Aska two-seater velvet sofa"},"brand":"Halvard","categories":["wohnen/sofas"],"attributes":{"width":"164 cm","upholstery":"Samt"},"varies_by":["color"],"variants":[{"id":"SOFA-ASKA-2-GRAU","attributes":{"color":"Grau"},"price":749.0,"currency":"EUR","availability":"in_stock"},{"id":"SOFA-ASKA-2-SALBEI","attributes":{"color":"Salbei"},"price":749.0,"sale_price":649.0,"currency":"EUR","availability":"in_stock"}]}
{"type":"content","id":"guide-samt-pflege","kind":"guide","title":"Samt richtig pflegen","url":"/ratgeber/samt-pflege","body":"Saugen Sie Samt mit einer weichen Bürste in Strichrichtung ab."}- The category
wohnen/sofas(living, sofas) names its parent and its path per language. - The sofa
SOFA-ASKA-2writes its attributes as the shop does, "164 cm" and "Samt" (velvet). It varies by color, and each variant carries its own price: Grau (grey) at 749 €, Salbei (sage) on sale at 649 €. - The guide "Samt richtig pflegen" (caring for velvet) is content, with its full text in
body.
A plain text is in the first language of the shop's first channel; an object holds one value per locale. What a value means comes from attributes.json (Types and attributes).
Reference: entity.json · variants · offers
A row that does not fit
A row with a broken field is left out whole, and the rest goes live. Line 27 of the home store's catalog writes a price as text:
{
"type": "product",
"id": "TEPPICH-SILO",
"variants": [
{ "id": "TEPPICH-SILO-120", "price": 159.0 },
{ "id": "TEPPICH-SILO-160", "price": "249,00 €" }
]
}So the rug TEPPICH-SILO (Teppich: rug) is left out with both its sizes, and the index run's report names the line, the field and why.
In the CSV export each row is one variant. It writes "ab 249 €" (from 249 €) for the larger size, so only that size is left out. A value that does not fit its attribute is dropped instead, and its entity kept.
Next
- Importing and mapping: see what an export becomes and correct how its columns are read.
- The import report: what every index run took in and what it left out.
- Types and attributes: what the values of a catalog mean.