# Importing and mapping

Keep an export as an import, see what it becomes, correct how its columns are read, then publish it.

An [import](/docs/reference/glossary#import) keeps your shop's [export](/docs/reference/glossary#export) without making it live. You see the products it becomes, correct how its columns are read, and only then [publish](/docs/reference/glossary#publish) it. The home store's export, a German Excel CSV, needs two corrections.

## See what an export becomes

The import reads the export with the published [release](/docs/reference/glossary#release) and shows where each column goes, why, and what it would leave out.

**In the panel:**

Open **Catalog → Imports** and drop the file anywhere on the page, or **Choose a file**:

The import opens on its first step, **Columns**:

- the counts of **Products**, **Variants**, **Categories**, **Rows** and **Problems**;
- each problem, the costliest first;
- every column with its values, where it goes and why, below the **Essentials** a product tile needs;
- **Set up for you**, the settings derived from the file, each with **Why**.

Beside them, the first products show as tiles.

**Through the API:**

Start the import, then read its inspection. Here are four of its 43 columns:

```sh title="Terminal"
draft=$(curl -s -X POST -H "Authorization: Bearer $key" -H 'Content-Type: text/csv' \
  --data-binary @tests/fixtures/feeds/home/export.csv "$api/imports?name=export.csv" | jq -r .import.id)
curl -s -H "Authorization: Bearer $key" $api/imports/$draft/inspection | jq -c '.columns[]
  | select(.name | IN("Vater", "Preis", "Lieferstatus", "Verkäufe 30 Tage")) | {name, mapping, reason}'
```

```json title="Answer: 200 · 33 ms"
{"name":"Vater","mapping":"group","reason":{"by":"alias","alias":"vater"}}
{"name":"Preis","mapping":"sale_price","reason":{"by":"reduced_price","beside":"Streichpreis"}}
{"name":"Lieferstatus","mapping":"availability","reason":{"by":"alias","alias":"lieferstatus"}}
{"name":"Verkäufe 30 Tage","mapping":"attributes.verkaeufe_30_tage","reason":{"by":"name"}}
```

And what it would leave out:

```sh title="Terminal"
curl -s -H "Authorization: Bearer $key" $api/imports/$draft/inspection | jq -c '{products, rejected},
  (.rejections[], .values.rejections[] | {column, count, value: .examples[0].value})'
```

```json title="Answer: 200 · 33 ms"
{"products":23,"rejected":1}
{"column":"Preis","count":1,"value":"ab 249 €"}
{"column":"Lieferstatus","count":3,"value":"in Kürze lieferbar"}
```

Each column goes where its name or values say. "Vater" (parent) groups the rows of one product into its [variants](/docs/reference/glossary#variant), and "Preis" (price) is the reduced price beside "Streichpreis" (struck-through price). "Verkäufe 30 Tage" (sales in 30 days) matches nothing, so it would become an [attribute](/docs/reference/glossary#attribute) named after the column.

The import finds 23 products and four problems. The rug `TEPPICH-SILO` writes "ab 249 €" (from 249 €) as a price, so that row is left out. "In Kürze lieferbar" (available soon) is no availability OrbSearch knows, so three variants would go without one.

**Reference:** [`GET .../inspection`](/docs/reference/management-api#inspect-import) · [`columns`](/docs/reference/management-api#inspect-import.answer.columns) · [`rejections`](/docs/reference/management-api#inspect-import.answer.rejections)

## Correct the mapping

A correction changes how the file is read, never the file. Map the shop's word to an availability, and the sales column to a ranking signal, which the import's own rules never pick.

**In the panel:**

In the problem "in Kürze lieferbar", choose what it means: **Backorder**. Then open where **Verkäufe 30 Tage** goes and choose `sales_30d` under **Ranking signals**. Each choice is saved into the draft at once, and the counts follow.

**Through the API:**

Take the proposed mapping from the inspection, correct it, and send it back as the draft's `feed.json`:

```sh title="Terminal"
version=$(curl -s -H "Authorization: Bearer $key" $api/imports/$draft | jq -r .version)
curl -s -H "Authorization: Bearer $key" $api/imports/$draft/inspection \
  | jq --arg version "$version" '{version: $version, files: {"feed.json": (.mapping
      | .columns.Lieferstatus = {to: "availability", values: {"in Kürze lieferbar": "backorder"}}
      | .columns["Verkäufe 30 Tage"] = "signals.sales_30d")}}' \
  | curl -s -X PATCH -H "Authorization: Bearer $key" -H 'Content-Type: application/json' -d @- $api/imports/$draft >/dev/null
curl -s -H "Authorization: Bearer $key" $api/imports/$draft/inspection \
  | jq -c '{mapping: .feed.mapping, rejected, values_rejected: (.values.rejected // 0)}'
```

```json title="Answer: 200 · 33 ms"
{"mapping":"file","rejected":1,"values_rejected":0}
```

`version` is the draft as you read it: a draft someone changed since refuses the write with `409`, so nobody overwrites another.

The mapping is now the draft's own `feed.json`, and the three variants keep their availability. The rug's price stays a problem until the shop's export writes a number.

**Reference:** [`PATCH /api/v1/imports/{import}`](/docs/reference/management-api#change-import) · [`columns.values`](/docs/reference/catalog/feed#columns.values) · [`columns.to`](/docs/reference/catalog/feed#columns.to)

### Suggestions from a model

Once `.env` gives the Query API an `AI_GATEWAY_API_KEY`, a decision model suggests where columns go; without one, nothing leaves your stack. The Query API asks `typesafe-ai/jev` through Vercel AI Gateway once per export, with the column names and values from the first rows.

A suggestion it is at least 98 % sure of, for a column the import could only guess, goes into the draft's `feed.json`. The panel shows the others as **AI suggests**, each with **Use**; the API answers them at `GET $api/imports/$draft/suggestions`.

**Reference:** [`GET .../suggestions`](/docs/reference/management-api#suggest-mapping)

## Review and publish

Before it goes live, compare the draft with what is live.

**In the panel:**

**Review and go live** shows the products that are **New**, **Gone** and **Staying**, and the attributes and filters that change. **Go live** publishes the draft. **Discard**, next to **Review and go live**, drops it instead.

**Through the API:**

```sh title="Terminal"
curl -s -H "Authorization: Bearer $key" $api/imports/$draft/changes | jq -c '.products, .attributes'
```

```json title="Answer: 200 · 36 ms"
{"draft":23,"live":23,"added":0,"removed":0}
{"added":["anzahl_bewertungen","bewertung","erschienen_am"],"removed":[]}
```

Publish it with the version you reviewed:

```sh title="Terminal"
version=$(curl -s -H "Authorization: Bearer $key" $api/imports/$draft | jq -r .version)
curl -s -X POST -H "Authorization: Bearer $key" "$api/imports/$draft/publish?version=$version" | jq -r .next
```

Or discard it:

```sh title="Terminal"
curl -s -X DELETE -H "Authorization: Bearer $key" $api/imports/$draft | jq -r .state
```

```text title="Answer: 200 · 18 ms"
discarded
```

All 23 products stay. Three columns would still become attributes of their own: "Bewertung" (rating), "Anzahl Bewertungen" (number of ratings) and "Erschienen am" (released on). The home store's `feed.json` maps them to ranking signals as well.

A draft that removes half the live products or more waits until you confirm the number it removes. Publishing queues an [index run](/docs/reference/glossary#index-run), whose [report](/docs/import-report) says what went live.

**Reference:** [`GET .../changes`](/docs/reference/management-api#show-import-changes) · [`POST .../publish`](/docs/reference/management-api#publish-import) · [`DELETE /api/v1/imports/{import}`](/docs/reference/management-api#discard-import)

## The mapping file

`feed.json` is a release file that says where each column goes. With it, the mapping is exactly what it says: a column it does not name is reported as unmapped, never guessed, so a new column in the export cannot change the shop on its own.

```json title="tests/fixtures/feeds/home/feed.json (trimmed)"
{
    "columns": {
        "Vater": "group",
        "Kategorie": { "to": "categories", "levels": " > " },
        "Lieferstatus": { "to": "availability", "values": { "in Kürze lieferbar": "backorder" } },
        "Breite (cm)": { "to": "attributes.width", "unit": "cm" },
        "Name EN": { "to": "title", "locale": "en" },
        "Verkäufe 30 Tage": "signals.sales_30d"
    }
}
```

A column says where a value goes, and `attributes.json` what it means: "Breite" (width) goes to `width`, whose unit and filters live there ([Types and attributes](/docs/types-and-attributes)).

**Reference:** [`feed.json`](/docs/reference/catalog/feed) · [`columns`](/docs/reference/catalog/feed#columns)

## On the command line

`orbsearch feed inspect` prints the same inspection for a file on your machine, without a stack. It needs the Rust toolchain:

```sh title="Terminal"
cargo run -p orbsearch -- feed inspect tests/fixtures/feeds/home/export.csv --release tests/fixtures/home
```

It reads the export with the folder's release files. `--write` saves the mapping there as `feed.json`, and `--json` prints the inspection for scripts and agents.

## Without a look

Automation that needs no look sends the export straight to `PUT /api/v1/catalog`, and its index run takes it live:

```sh title="Terminal"
curl -s -X PUT -H "Authorization: Bearer $key" -H 'Content-Type: text/csv' \
  --data-binary @tests/fixtures/feeds/home/export.csv $api/catalog | jq -r .next
```

The file is read with the published release's `feed.json` and kept byte for byte as a [snapshot](/docs/reference/glossary#snapshot).

<Callout type="warn">
    An upload replaces the whole catalog, and nothing holds it back: a product missing from the file is gone once its
    run is live. Only an import and a [feed](/docs/feeds) wait when half the live products would go.
</Callout>

**Reference:** [`PUT /api/v1/catalog`](/docs/reference/management-api#upload-catalog)

## Next

- [The import report](/docs/import-report): what every index run took in and what it left out.
- [Feeds on a schedule](/docs/feeds): fetch the export from your shop's URL.
- [Formats we read](/docs/formats): the files an import takes, and OrbSearch JSON Lines.
