# Insights

What shoppers search and click, what finds nothing and the filters they choose, from a search log the Query API keeps off the read path, masked before it is stored.

Every search already passes through the Query API, so the merchant's most useful numbers need no shop integration: how many searches, the queries searched most, those that found nothing, and the filters chosen on each page. The Query API logs what it answers, the control plane rolls the log up per day, and the management API reports it.

## What is logged

The Query API logs every answer on its public routes: `POST` and `GET /v1/search`, and `GET /v1/resolve` when it answers a listing with its results. It logs after the answer, never inside it, so a preview in the panel, which shares the search, is never counted, and neither is a request for totals alone (`count_only`), which a filter sheet sends while the shopper chooses. One row per answer holds:

| Field                    | What                                                                                                                                    |
| ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------- |
| Query ID and time        | The `query_id` of the response, which clicks name, and when it was answered, which the ID carries                                       |
| Release and snapshot     | What answered                                                                                                                           |
| Channel, locale, surface | Where: `search`, `category_page` or `autocomplete`, with the category of a category page                                                |
| Path                     | The merchant's URL a resolved page was asked for, without its query string                                                              |
| Query                    | As typed, and its words as the engine reads them                                                                                        |
| Filters, sort and page   | The shopper's choices                                                                                                                   |
| Results                  | Each section's total, the listed section and its total, whether it redirected, and whether nothing was found: no result and no redirect |
| Placements and sponsored | The rules whose banners the answer placed, and each sponsored product it served with its campaign                                       |
| Took                     | How long the Query API took, from the request's arrival to its answer                                                                   |
| Page view and traffic    | Which page load asked, and whose traffic it is                                                                                          |

Logging never makes a search wait. The handler puts the answer into a buffer of 10,000 without waiting, and one writer inserts up to 500 rows at a time, or what arrived within a second, as the Query API's own database role, which may insert into the log and do nothing else there. A full buffer, or a write Postgres refuses, loses those searches; the log counts them, and every report says how many went unrecorded. A crash loses what the buffer held, a few seconds of searches at most.

### Personal data

Before a query or a path is stored, email addresses and phone numbers in it are replaced with `[email]` and `[phone]`. Email addresses are found by [linkify](https://github.com/robinst/linkify), phone numbers by libphonenumber's metadata ([phonenumber](https://github.com/whisperfish/rust-phonenumber)), written internationally or nationally in a country the channel sells to, never by patterns of OrbSearch's own; sizes such as `160 200` are no phone numbers. A masked query never reaches a report. Names and street addresses are the known miss, which the rule below covers: a query reaches the reports only once two page views typed it that day.

Nothing is stored on the shopper's device. The storefront names each page load with a random page view id, in memory, and sends it with that page's searches; `@orbsearch/client` makes one per client, so a storefront makes one client per page load, and on a server one per request; a page rendered on the server hands its page view to the browser's client, as the demo does, so the page load counts once.

## Clicks and views

A search with results is not yet a good search: clicks say which searches fail although they find something, and where in the list shoppers find what they want. A storefront reports them to `POST /v1/events`, on the Query API's public port like the search, with no key, no cookie and nothing stored on the shopper's device:

```json title="POST /v1/events"
{
    "page_view": "0199c4a2-7b1e-4d3a-9f2c-5e8b1a4d6c70",
    "events": [
        {
            "type": "click",
            "query_id": "0199c4a2-7c02-7d3a-9f2c-5e8b1a4d6c70",
            "entity": "product:SOFA-LUND-3",
            "position": 3
        },
        { "type": "view", "query_id": "0199c4a2-7c02-7d3a-9f2c-5e8b1a4d6c70", "placement": "sofa-week" }
    ]
}
```

- **`query_id`** is the one the answer carried, of the search's first page: a hit a later page brought names the search it continues, and its `position` counts from 1 across the pages, as a grid banner's does. The ID does not go into product links; the results page knows it when the shopper clicks.
- **`entity`** names a hit, whose type names its section, or **`placement`** the rule whose banner it was, never both. A sponsored hit carries its `campaign`.
- **`type`** is `click`, or `view` once half of the hit or banner was visible for a second; a storefront reports one view per query ID and hit or banner.
- **`surface`** is `storefront` by default, or `agent` for an agent's tool call, such as WebMCP. `page_view` and `traffic` are the page's, as its searches sent them, so a test's events are left out with its searches and a bot's user agent marks them as a bot's: a click joins only a search of its own traffic.

A batch holds 1 to 100 events (`LIMITS.events_per_call`) in at most 64 kB, what a browser's `navigator.sendBeacon` may carry: a click goes at once, views together when the page is hidden. The body is read as JSON whatever its content type, since a beacon posts `text/plain` and so needs no preflight. The answer is `202` with `{ "accepted": 2, "late": 0 }`, which a beacon never reads.

The Query API checks a batch by its query IDs alone and looks nothing up: a query ID is a UUIDv7 that carries the millisecond it was answered. A batch is refused whole with `400` when an event names both or neither of an entity and a placement, a position past the 1,000 hits pages reach, or its query ID is not a UUIDv7 or names a time still to come. An event whose answer is older than 48 hours (`EVENT_HOURS`) is refused alone and counted as late; an event whose search is not logged yet is taken, and joined when the log rolls up. Events pass the same buffer as searches, so a full buffer drops and counts them too. Taking a batch has its own latency budget, 2 ms at p95 and 5 ms at p99, since it only reads the body and hands it over; nothing new runs on the search's path.

### Tests, benchmarks and bots

A request says whose traffic it is with `traffic`: `test` for end-to-end tests and `bench` for a benchmark, `shopper` by default. The Query API marks a shopper's request as a bot's when [uap-core](https://github.com/ua-parser/uap-core)'s maintained user agent rules call its `User-Agent` a spider. A storefront that searches from its server passes the shopper's user agent on, since its own would hide every crawler:

```ts title="On the server, once per request"
const search = createSearch({ endpoint, userAgent: request.headers.get("user-agent") ?? undefined });
```

Reports count shoppers by default, and say how many answers to the others they left out.

## How it is counted

The control plane rolls the log up about once a minute, so a search or a click shows in the reports within a minute or two. A day is counted again from the raw log until a roll-up a quarter of an hour after the last click on its searches may arrive, 48 hours after its end; from then on its numbers stand. Days are UTC days, and an event counts on the day of the answer it names.

- **A search** is the first page of an answer: on the search page only one with words typed, since a storefront's own rails ask the search page without any. The pages after the first, category pages and autocomplete are counted apart.
- **No results** is the searches that found nothing, of all searches, and every report gives both numbers.
- **Click-through** is the searches with a click, of all searches; a click joins its search by query ID, whenever within the 48 hours it arrives.
- **No click** is the searches with results and no click, of the searches with results: a search that found nothing or sent the shopper elsewhere could not be clicked, so it does not count here.
- **Average position** is where the clicked hits of the listed section stood, counted from 1 across the pages; a click on a banner or a hit beside the listing counts as a click, not here.
- **A query** is its words as the engine reads them, so `Sofa` and `sofa` are one. It is listed on a day once two distinct page views typed it that day, and never when it is longer than 200 characters; until then it counts among the searches, never by its words.
- **Filters per page** count the searches on a category page, or on the search page, that chose each filter, of the searches on that page.
- **Took** is the 50th, 95th and 99th percentile per day of the answers on the search page, the later pages included.

The raw log and the events live in daily partitions, kept 60 days and then dropped; the daily rollups stay.

The same roll-up counts each banner, by the rule that placed it, and each campaign of [sponsored products](/docs/sponsored-products), per day and traffic:

- **Served** is the answers that placed it, every page counted.
- **Views** is the query IDs it was seen in: a view or a click names the first page's query ID wherever the tile stood, and counts once per query ID and banner or product, since a shopper clicks only what they saw.
- **Clicks** is the views that drew a click, so the click-through of a campaign is its clicks of its views.

An event counts only where an answer of its page and its traffic placed what it names, a sponsored product with that campaign or the banner's rule: the answer its query ID names, or a later page of the same page view, since events name the first page's query ID wherever the tile stood. An event that names a product or a campaign no answer placed counts for no one.

## In the panel

**Insights**, above Releases in the sidebar, shows the report for shoppers over today, the last 7, 30 or 90 days, in two tabs: **Searches** and **Campaigns**, which keep the days when you move between them. On **Searches**, on top are the searches, the share that found nothing and the click-through, each with the numbers it is made of, the distinct queries and a line of searches per day, with the searches that drew a click as a lighter line under it. Before the first click in the range the click-through says **No clicks yet** instead of 0 %, since the storefront may not report clicks yet. Below, every table ends in what to do about a row:

- **Found nothing** lists the queries that found nothing most often. **Add a synonym** opens Synonyms with the query's words on its line, waiting for the word your catalog uses; the list icon starts a rule for the query on Rules, in a draft, and the flask opens the search in the Playground.
- **Found something, no click** lists the queries whose searches had results and drew no click most often, under the share of searches with results that went without one. The flask opens the search in the Playground to see what shoppers saw, and the list icon starts a rule that puts the right products first. Without a click in the range it says instead that clicks count once the storefront reports them, since until then every query with results is listed.
- **Searched most** lists the queries searched most with what their latest search listed, each opening in the Playground on the search page of the channel and language in the top bar. Once shoppers click, it adds the share of each query's searches that drew a click and, on a wide screen, the average position of the results clicked, counted from 1, with the average over every search under its title.
- **Filters per page** lists the category pages and the search page with the share of searches that chose a filter and the filters chosen most, each opening the page in Pages.

A line under the tables says what each rule left out: the searches the Query API could not record always, and clicks and views that came late, matched no search of these shoppers or could not be recorded, once there are any. Before the first roll-up the screen says that searches show within a minute or two; a range up to today is read again with each roll-up while the screen is open. **Overview** shows the same four numbers for the last 30 days and the three queries that found nothing most often, each leading to its synonym.

## Campaigns

`GET /api/v1/insights/campaigns` answers each campaign of sponsored products between two UTC days, both included, from the rollups, so a campaign's report reaches back past the 60 days of raw searches: by default the last 30 days and shoppers, the most viewed campaign first.

```sh title="Terminal"
curl -s -H "Authorization: Bearer $key" "$api/insights/campaigns?from=2026-10-01&to=2026-10-31"
```

```json
{
    "from": "2026-10-01",
    "to": "2026-10-31",
    "traffic": "shopper",
    "rolled_up_at": "2026-10-31T10:41:00.402113Z",
    "campaigns": [
        {
            "campaign": "Michelin spring",
            "served": 4210,
            "views": 3120,
            "click_through": { "count": 132, "of": 3120 },
            "days": [{ "day": "2026-10-01", "served": 140, "views": 101, "clicks": 4 }]
        }
    ]
}
```

`days` lists the days a campaign was served or seen. With `Accept: text/csv` the same report comes as the CSV a merchant sends the brand, one row per campaign: `campaign`, `from`, `to`, `days` (how many), `served`, `views`, `clicks` and `click_through` as a share such as `0.0423`, empty without views. A campaign name that starts like a formula is written with a leading `'`, so a spreadsheet shows it as text.

**Insights → Campaigns** in the panel shows this report for shoppers over the same days as **Searches**: a row per campaign, the most viewed first, with how often answers served it, its views and clicks, its click-through beside the views it is a share of, and the days it was served or seen; a phone keeps the campaign, its views and its click-through. **Download CSV** downloads the CSV above for the same days, written by the control plane, so it always says what the table says. Without a campaign in the range the screen says to mark a pinned product as sponsored in Rules, and that its numbers show within a minute or two of shoppers seeing it; a range up to today is read again with each roll-up.

## The report

`GET /api/v1/insights` answers what one kind of traffic searched between two UTC days, both included: by default the last 30 days, shoppers, and 20 rows in each list.

```sh title="Terminal"
curl -s -H "Authorization: Bearer $key" "$api/insights?from=2026-10-01&to=2026-10-09&limit=3"
```

```json
{
    "from": "2026-10-01",
    "to": "2026-10-09",
    "traffic": "shopper",
    "rolled_up_at": "2026-10-09T10:41:00.402113Z",
    "releases": ["5f0c…", "9a71…"],
    "searches": 1240,
    "no_results": { "count": 87, "of": 1240 },
    "click_through": { "count": 702, "of": 1240 },
    "no_click": { "count": 421, "of": 1123 },
    "average_position": 4.2,
    "category_pages": 3310,
    "autocomplete": 5402,
    "distinct_queries": 214,
    "days": [
        {
            "day": "2026-10-01",
            "searches": 131,
            "no_results": 9,
            "clicked": 77,
            "category_pages": 360,
            "autocomplete": 570,
            "took_ms": { "p50": 4.1, "p95": 7.9, "p99": 11.2 }
        }
    ],
    "queries": [
        {
            "query": "sofa",
            "searches": 96,
            "no_results": 0,
            "clicked": 71,
            "with_results": 96,
            "no_click": 25,
            "average_position": 3.1,
            "page_views": 88,
            "results": 23
        }
    ],
    "no_result_queries": [
        {
            "query": "couch",
            "searches": 31,
            "no_results": 31,
            "clicked": 0,
            "with_results": 0,
            "no_click": 0,
            "page_views": 29,
            "results": 0
        }
    ],
    "no_click_queries": [
        {
            "query": "sofa grau",
            "searches": 40,
            "no_results": 0,
            "clicked": 2,
            "with_results": 40,
            "no_click": 38,
            "average_position": 11,
            "page_views": 37,
            "results": 40
        }
    ],
    "pages": [
        {
            "category": "wohnen/sofas",
            "searches": 820,
            "filtered": 301,
            "filters": [{ "field": "color", "searches": 190 }]
        }
    ],
    "left_out": {
        "traffic": { "test": 0, "bench": 0, "bot": 412 },
        "rare_queries": 380,
        "personal_data": 2,
        "not_recorded": 0,
        "events": { "late": 3, "unmatched": 1, "not_recorded": 0 }
    }
}
```

`days` lists every day of the range, a day without searches included. `results` is what the query's latest search listed. A page without `category` is the search page. `no_click_queries` lists the queries whose searches had results and drew no click most often, and `average_position` is left out where nothing of the listing was clicked. `left_out` says what each rule kept out: answers to the other kinds of traffic, searches whose query fewer than two page views typed, searches whose query held personal data, and searches the Query API could not log; under `events`, the events refused as late by the day they arrived, those naming an answer the log holds nothing of their traffic for (a preview's, an unrecorded search's, another traffic's or a made-up ID), and those the Query API could not log. Before the first roll-up, `rolled_up_at` is left out and every number is 0.
