# Context pack: How-to guides

Source: https://nordvec.com/cs/docs/guides/how-to
Pack: https://nordvec.com/cs/docs/packs/how-to

This pack bundles one Nordvec guide with the guides it builds on and the guides it links to, in reading order, so an assistant reading it meets no reference it cannot follow.

## Contents

1. [How-to guides](https://nordvec.com/cs/docs/guides/how-to) (this guide)
2. [Push documents from your own systems](https://nordvec.com/cs/docs/guides/how-to/push-documents) (linked from this guide)
3. [Filter and refine search](https://nordvec.com/cs/docs/guides/how-to/filter-search) (linked from this guide)
4. [List and retrieve documents](https://nordvec.com/cs/docs/guides/how-to/list-documents) (linked from this guide)

---

# How-to guides
Source: https://nordvec.com/cs/docs/guides/how-to

Step-by-step recipes for the common API tasks, from searching and listing documents to pushing your own.



Each guide takes one task from the first request to a working result, with the
request fields it uses and the response you get back. For every field of every
operation, see the [API reference](/docs/api).

- [Push documents from your own systems](https://nordvec.com/cs/docs/guides/how-to/push-documents): Create a datasource, push documents into it with an indexing API key, choose who can read them, and pause or delete it when the source changes.
- [Filter and refine search](https://nordvec.com/cs/docs/guides/how-to/filter-search): Narrow a document search with datasource, provider, type and date filters, and read the ranked results.
- [List and retrieve documents](https://nordvec.com/cs/docs/guides/how-to/list-documents): Page through your documents with cursor pagination, filter and sort them, and fetch one or many by id.


---

# Push documents from your own systems
Source: https://nordvec.com/cs/docs/guides/how-to/push-documents

Create a datasource, push documents into it with an indexing API key, choose who can read them, and pause or delete it when the source changes.



The push API indexes documents from systems Nordvec has no connector for: an
internal wiki export, a ticket archive, a database of notes. You send the text
and who may read it; Nordvec stores it in the EU, indexes it, and makes it
searchable and citable like any other document. Every pushed document lands in
a **datasource**, a named container in your workspace that a workspace admin
creates first. A push that names a datasource which does not exist, or one that
is paused, is refused.

## Create a datasource [#create-a-datasource]

Open **Workspace settings > Datasources** and choose **Create datasource**.
Workspace admins and owners can do this; in a personal workspace, that is you.

| Field | Notes                                                                                                                                             |
| ----- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
| Name  | What people see in the settings list. Up to 200 characters.                                                                                       |
| Slug  | What every push names. Lowercase letters, digits, `-` and `_`, starting with a letter or digit, up to 200 characters. It cannot be changed later. |

The slug `confluence-export` is used in the examples below.

## Create an indexing API key [#create-an-indexing-api-key]

Pushes authenticate with an API key of the **Indexing** class that carries the
`index:write` scope; add `index:status` to track ingestion and `index:delete`
to remove documents or replace a whole datasource. Create one under **Workspace settings > API keys**; the
raw key starts with `nv_eu_idx_` and is shown once. See
[Authentication](/docs/guides/authentication). Each request also names your
workspace id as `tenantId`, the id in your workspace's address in the app
(`/w/<workspace id>/...`), and it must be the workspace the key belongs to.

## Push one document [#push-one-document]

`/documents/push` creates the document, or updates it when a document with the
same `id` already exists in the datasource.

```bash
curl https://nordvec.com/api/v1/documents/push \
  -H "Authorization: Bearer $NORDVEC_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: page-4711-2026-09-28" \
  -d '{
    "tenantId": "YOUR_WORKSPACE_ID",
    "document": {
      "id": "page-4711",
      "title": "Travel expense policy",
      "datasource": "confluence-export",
      "body": { "mimeType": "text/markdown", "content": "# Travel expenses\n..." },
      "permissions": {},
      "sourceUrl": "https://wiki.example.com/pages/4711",
      "type": "policy"
    }
  }'
```

```json
{ "documentId": "page-4711", "status": "queued", "updated": false }
```

* `id` is your stable id for the document within the datasource. Pushing the
  same `id` again updates it; unchanged content is recognised by its hash and
  not indexed twice.
* `body.mimeType` is one of `text/plain`, `text/markdown`, `text/html`,
  `application/pdf`, or the Word, Excel and PowerPoint (`.docx`, `.xlsx`,
  `.pptx`) types. Binary content is sent base64-encoded.
* `sourceUrl` becomes the "jump to source" link on every citation of the
  document. Omit it on a re-push to keep the stored one, or send `null` to
  clear it.
* `type` sets the document's `content_type`, which search and list filter on.

The whole request body is capped at 1 MB, so a large file or a big batch
answers `413`; split it.

## Choose who can read it [#choose-who-can-read-it]

`permissions` is required on every push, so a sharing decision is never made
by leaving a field out. In a datasource visible to the workspace:

| `permissions`                             | Who can read the document                                  |
| ----------------------------------------- | ---------------------------------------------------------- |
| `{}`                                      | Every member of the workspace                              |
| `{ "allowedUsers": ["ana@example.com"] }` | Only the people listed                                     |
| `{ "allowedGroups": ["GROUP_ID"] }`       | Members of those workspace groups, including nested groups |
| `{ "allowAllTenantMembers": false }`      | Refused: a document nobody can read is a delete            |

To change who can read a document without sending its content again, use
`POST /documents/push/permissions`. Making an already-restricted document
visible to the whole workspace additionally needs the `index:acl-widen` scope,
so a routine sync cannot quietly undo a restriction someone set by hand.

## Push in batches [#push-in-batches]

`/documents/push/bulk` takes up to 100 documents for one datasource per call.
The answer counts `accepted` and `rejected` and gives a result per document, so
one bad document does not fail the batch. The 1 MB body cap applies per call,
so split large uploads into several calls under the same `uploadId`.

```bash
curl https://nordvec.com/api/v1/documents/push/bulk \
  -H "Authorization: Bearer $NORDVEC_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "tenantId": "YOUR_WORKSPACE_ID",
    "uploadId": "nightly-2026-09-28",
    "datasource": "confluence-export",
    "documents": [
      { "id": "page-4711", "title": "Travel expense policy", "datasource": "confluence-export",
        "body": { "mimeType": "text/plain", "content": "..." }, "permissions": {} }
    ]
  }'
```

## Replace a whole datasource [#replace-a-whole-datasource]

When your system can list everything a datasource should hold, send the full
listing as one **upload session**, and the documents it no longer contains are
moved to the trash when the session closes. Sessions need an indexing API key
holding `index:delete` as well as `index:write`, because the close removes
documents; the key that opens one is the only key that can continue it.

1. Send the first page with `"isFirstPage": true`. It is page `0`.
2. Send every further page with its `pageIndex` (`1`, `2`, ...), in any order.
   A page sent twice counts once, so a retry is always safe.
3. Send the last page with `"isLastPage": true` and its `pageIndex`. A listing
   that fits in one page sends `isFirstPage` and `isLastPage` together. The
   last page may carry no documents.

Every page uses the same `uploadId`, and each answer carries the session's
progress under `upload`. The session closes only when every page from `0` to
the last has arrived. Closing it moves to the trash each document in the
datasource that no page of the session named and that existed before the
session opened. Any other push to the datasource while the session runs keeps
the document it names: a single push, a batch without session fields, a
permissions update, and a re-push of unchanged content alike. The trash keeps
what the close moved there for 30 days; pushing a document again brings it
back, and so does restoring the whole session (see below).

A session that receives no page for 24 hours expires and closes without
removing anything. A refused page is answered with `409 Conflict`, writes
nothing, and its `data.reason` says why:

| `reason`                                               | What to do                                                                                                                                                                                                                                                                                                  |
| ------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `upload_incomplete`                                    | Send the pages listed in `missingPageIndexes`, then the last page again                                                                                                                                                                                                                                     |
| `deletion_confirmation_required`                       | The close would trash more than 20% of the datasource. If that is right, send the last page again with `"confirmDeletions"` set to `wouldTombstone`                                                                                                                                                         |
| `deletion_confirmation_too_large`                      | `confirmDeletions` is larger than the number of documents the datasource held when the session opened. Send the count you expect to remove                                                                                                                                                                  |
| `upload_in_progress`                                   | A session is open on this datasource. If it is your key's, finish it, wait for it to expire, or start over with `"forceRestartUpload": true` on your first page. If another key opened it, `forceRestartUpload` replaces it only once it has received no page for an hour, from the time in `restartableAt` |
| `upload_expired`, `upload_missing`, `upload_restarted` | The session is gone; start a new one with a new `uploadId`                                                                                                                                                                                                                                                  |
| `upload_closed`, `upload_id_reused`                    | The `uploadId` is spent; use a new one                                                                                                                                                                                                                                                                      |
| `page_index_required`                                  | Your key has a session open on this datasource; send `pageIndex` with the page                                                                                                                                                                                                                              |

To resume after a crash, read the session with
`GET /documents/push/upload?tenantId=...&datasource=...&uploadId=...` (scope
`index:status`). Its `missingPageIndexes` lists the pages still to send.

### Undo a session's close [#undo-a-sessions-close]

If a session removed documents it should not have, for example because the
listing it sent was cut short, restore them in one call:

```bash
curl https://nordvec.com/api/v1/documents/push/upload/restore \
  -H "Authorization: Bearer $NORDVEC_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "tenantId": "YOUR_WORKSPACE_ID", "datasource": "confluence-export", "uploadId": "nightly-2026-09-28" }'
```

The key that opened the session can restore it, and so can a workspace admin
signed in to Nordvec, for a session opened by any key. Every document the close
moved to the trash comes back with the content it had, and the answer counts
them: `restored` are live again, `purged` had already been deleted for good by
the trash, and `skipped` had changed since the close (pushed again, or removed
again) and were left as they are. Restoring a session twice answers with the
first restore's counts and `"replayed": true`, and queues any restored document
still waiting to be indexed, so repeating a restore that failed to answer is
safe. A session can be restored for 35 days after it closed, and for as long
as the trash still holds any document it removed. A refused restore is answered
with `409 Conflict` and its `data.reason`:

| `reason`                 | What it means                                                                                                                                                                                      |
| ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `upload_not_closed`      | The session never closed, so it removed nothing                                                                                                                                                    |
| `upload_in_progress`     | A session is open on the datasource. Restore once it has closed or expired                                                                                                                         |
| `restore_purged`         | More than 30 days have passed, and the trash has deleted every one of the documents. Push them again                                                                                               |
| `workspace_not_entitled` | The workspace's plan does not currently allow restoring from the trash                                                                                                                             |
| `corpus_cap_exceeded`    | Bringing the documents back would pass the workspace's document limit, so none came back. `data.wouldRestore` is how many it needs and `data.headroom` how many fit. Free room, then restore again |

## Track ingestion [#track-ingestion]

A push answers as soon as the document is queued. Ask for its progress with
`GET /documents/push/status` (scope `index:status`), filtered by datasource or
document id. A document moves from `queued` through `processing` to
`completed`, or to `failed` with an `error`.

## When a push is refused [#when-a-push-is-refused]

A push that names an unknown or paused datasource is answered with
`422 Unprocessable Content`. The message names the slug and links to
**Workspace settings > Datasources** in your workspace, and the error's `data`
says why and what to do:

```json
{
  "defined": true,
  "code": "UNPROCESSABLE_CONTENT",
  "status": 422,
  "message": "Datasource \"confluence-export\" is paused and accepts no documents. A workspace admin resumes it under Workspace settings > Datasources: https://nordvec.com/w/YOUR_WORKSPACE_ID/workspace/settings?tab=datasources",
  "data": {
    "why": "The datasource \"confluence-export\" is paused",
    "fix": "Resume it at https://nordvec.com/w/YOUR_WORKSPACE_ID/workspace/settings?tab=datasources, then retry the push",
    "link": "https://nordvec.com/docs/guides/how-to/push-documents"
  }
}
```

Do not retry these automatically: they succeed only after an admin creates or
resumes the datasource.

## Remove a document [#remove-a-document]

`POST /documents/push/delete` (scope `index:delete`) removes a pushed document
by its `datasource` and `id`. Documents you stop pushing are not removed on
their own: delete each one you retire, or send the datasource's full listing as
an upload session, described above.

## Pause, resume and delete [#pause-resume-and-delete]

* **Pause** refuses every further push into the datasource. Its documents stay
  searchable. A push already being written when you pause completes.
* **Resume** accepts pushes again.
* **Delete** removes the datasource and every document pushed to it, with
  their search index. Your own system keeps its copy, so pushing again after
  you recreate the datasource restores them. A delete cannot be undone.

If another admin changed the datasource after your list loaded, the action is
refused and the list reloads, so you decide again against what is there now.
Every create, pause, resume and delete is recorded in the workspace audit log.

<Callout>
  The settings list shows who each datasource is visible to. Who can read a
  pushed document is decided by the `permissions` sent with it; creating,
  pausing or deleting a datasource never widens access to anything.
</Callout>

<Callout>
  Push writes are idempotent: repeat the same `Idempotency-Key` on every retry of
  one write, and a duplicate is answered from the first attempt instead of being
  applied twice. See [Errors and rate limits](/docs/guides/errors-and-rate-limits).
</Callout>

## Next steps [#next-steps]

<Cards>
  <Card title="List and retrieve documents" href="/docs/guides/how-to/list-documents" />

  <Card title="Filter and refine search" href="/docs/guides/how-to/filter-search" />

  <Card title="API reference" href="/docs/api" />
</Cards>


---

# Filter and refine search
Source: https://nordvec.com/cs/docs/guides/how-to/filter-search

Narrow a document search with datasource, provider, type and date filters, and read the ranked results.



`/documents/search` runs full-text search over the title and the whole text of
your documents and returns the best matches, each with a relevance score and the
source document it came from. This guide shows how the query matches, how to
narrow the results with filters, and how to read the response.

## The request [#the-request]

Only `query` is required. Everything else narrows or bounds the results.

```bash
curl https://nordvec.com/api/v1/documents/search \
  -H "Authorization: Bearer $NORDVEC_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "query": "renewal terms",
    "limit": 20,
    "datasource": "contracts",
    "sourceProvider": "google",
    "createdAfter": "2026-01-01T00:00:00Z",
    "createdBefore": "2026-07-01T00:00:00Z"
  }'
```

| Field            | Type    | Notes                                                                                                                                    |
| ---------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------- |
| `query`          | string  | Required. 1 to 500 characters. Quoted phrases, `or` and a leading `-` to exclude a word are understood.                                  |
| `limit`          | integer | Optional. 1 to 50, default 20. How many results to return.                                                                               |
| `datasource`     | string  | Optional. Restrict to one datasource by its slug (up to 200 characters).                                                                 |
| `sourceProvider` | string  | Optional. Restrict to one connector provider, for example `google`, `sharepoint` or `slack`.                                             |
| `createdAfter`   | string  | Optional. ISO 8601 timestamp with an offset; only documents created at or after it.                                                      |
| `createdBefore`  | string  | Optional. ISO 8601 timestamp with an offset; only documents created at or before it.                                                     |
| `contentType`    | string  | Optional. Restrict to one knowledge type, the `type` a pushed document or a knowledge file declares (for example `policy` or `runbook`). |

<Callout>
  Every filter is combined with AND: a document must match the query **and** every
  filter you supply. Leave a filter out to widen the search.
</Callout>

## How the query matches [#how-the-query-matches]

* **The whole document is searched.** The title and every passage of the text
  count, however long the document is.
* **Words match in the form you write them, in any language.** There is no
  stemming: `invoice` does not match `invoices`, and `tilbagebetaling` does not
  match `tilbagebetalingen`. To catch several forms, join them with `or`.
* **Accents are ignored on both sides.** `cafe` finds `café`, and `børnehave`
  and `bornehave` find each other. Case is ignored too.
* **Operators.** Put words in double quotes to match them as a phrase, write
  `or` between words to match either, and put `-` before a word to leave out
  documents that contain it.

## Try it [#try-it]

Signed in, you can run a search over your own documents from this page. Change
the query in the API reference to try your own.

Try it in the API reference: [`POST /api/v1/documents/search`](https://nordvec.com/docs/api#tag/documents/POST/documents/search) (Search documents).

## The response [#the-response]

```json
{
  "results": [
    {
      "id": "0198f2a4-6c1e-7d30-b6a1-2f9d54c08a11",
      "title": "Acme Corp Master Services Agreement",
      "snippet": "Automatic **renewal**: the agreement continues for successive twelve month **terms** unless…",
      "score": 0.82,
      "datasource": "contracts",
      "source_provider": "google",
      "source_type": null,
      "mime_type": "application/pdf",
      "content_type": null,
      "status": "indexed",
      "created_at": "2026-02-14T09:00:00Z",
      "updated_at": "2026-02-14T09:00:00Z"
    }
  ],
  "totalCount": 7
}
```

Each result is one document. Its `snippet` is cut from the passage that matched
best, wherever that passage is in the text, with the matched words wrapped in
`**`; a word you typed without its accents is found and ranked, but may appear
unmarked in the snippet. The `score` runs from 0 up to, but never reaching, 1
(higher is more relevant), and the source document's metadata lets you trace
the result back. `totalCount` is how many documents matched in total, which may
be larger than the number of `results` you asked for with `limit`.

## Reading the results [#reading-the-results]

* **Results are ranked by relevance**, most relevant first. A document ranks by
  its best-matching passage. Use `score` to drop weak matches within one set of
  results; scores from different queries are not on the same scale.
* **`totalCount` vs `results.length`**: `results` holds up to `limit` items;
  `totalCount` is the full match count. If `totalCount` is much larger than your
  `limit`, tighten the `query` or add a filter; there is no second page of
  search results.
* **`status` tells you where the document is in processing.** A document is
  matched on its stored text, so one still `processing` can appear; `indexed`
  means every step finished. See
  [Documents & search](/docs/guides/concepts/documents) for the lifecycle.

<Callout>
  Search only ever returns documents the caller is allowed to see. Access is
  enforced in the database, not in application code, so a filter can never widen
  what the caller sees. For an API key that is what the workspace shares; see
  [Who sees a document](/docs/guides/concepts/documents#who-sees-a-document).
</Callout>

## Next steps [#next-steps]

<Cards>
  <Card title="Documents & search" href="/docs/guides/concepts/documents" />

  <Card title="List and retrieve documents" href="/docs/guides/how-to/list-documents" />

  <Card title="API reference" href="/docs/api" />
</Cards>


---

# List and retrieve documents
Source: https://nordvec.com/cs/docs/guides/how-to/list-documents

Page through your documents with cursor pagination, filter and sort them, and fetch one or many by id.



Where [search](/docs/guides/how-to/filter-search) ranks documents by relevance to
a query, listing walks your whole corpus in order. Use it to sync, audit, or
build your own index over what Nordvec holds. Listing returns metadata only, no
document content.

## List with cursor pagination [#list-with-cursor-pagination]

`/documents/list` returns a page of documents plus an opaque `nextCursor`. Pass
that cursor back to get the next page, and stop when `hasMore` is `false`.

```bash
curl "https://nordvec.com/api/v1/documents/list?limit=50" \
  -H "Authorization: Bearer $NORDVEC_API_KEY"
```

```json
{
  "items": [
    { "id": "…", "title": "Q3 Financial Report", "status": "indexed", "datasource": "finance", "created_at": "2026-04-01T10:00:00Z" }
  ],
  "nextCursor": "eyJrIjoi…",
  "hasMore": true
}
```

Signed in, you can list the first page of your own documents from here:

Try it in the API reference: [`GET /api/v1/documents/list`](https://nordvec.com/docs/api#tag/documents/GET/documents/list) (List documents).

To walk every page, loop until `hasMore` is `false`, passing the previous
response's `nextCursor` each time, with the same `sort` and `direction`:

```bash
curl "https://nordvec.com/api/v1/documents/list?limit=50&cursor=eyJrIjoi…" \
  -H "Authorization: Bearer $NORDVEC_API_KEY"
```

<Callout>
  The cursor is opaque, do not parse or construct it. Pass back exactly what the
  previous response returned. An invalid cursor, or one from a different sort
  order, is rejected.
</Callout>

## Filter and sort [#filter-and-sort]

All filters are optional and combine with AND. Sorting defaults to newest first.

| Field                            | Type    | Notes                                                                |
| -------------------------------- | ------- | -------------------------------------------------------------------- |
| `limit`                          | integer | 1 to 200 (default 50).                                               |
| `cursor`                         | string  | Opaque cursor from the previous page.                                |
| `datasource`                     | string  | Restrict to one datasource by its slug (up to 200 characters).       |
| `status`                         | enum    | `indexed`, `processing`, or `failed`.                                |
| `sourceProvider`                 | string  | Restrict to one connector provider, for example `google` or `slack`. |
| `contentType`                    | string  | Restrict to one knowledge type, for example `policy` or `runbook`.   |
| `createdAfter` / `createdBefore` | string  | ISO 8601 timestamps with an offset, both inclusive.                  |
| `sort`                           | enum    | `createdAt` (default), `updatedAt`, or `title`.                      |
| `direction`                      | enum    | `desc` (default) or `asc`.                                           |

```bash
curl "https://nordvec.com/api/v1/documents/list?status=indexed&datasource=finance&sort=updatedAt&direction=desc&limit=100" \
  -H "Authorization: Bearer $NORDVEC_API_KEY"
```

### Which sort to walk with [#which-sort-to-walk-with]

* **A complete, one-shot enumeration**: `sort=createdAt`. The creation time
  never changes, so every document appears exactly once.
* **Incremental catch-up from a watermark**: `sort=updatedAt&direction=asc`. A
  document updated while you walk can appear twice, so upsert by `id`.
* **Display order**: `updatedAt` or `title` descending. A document updated
  between two pages can move past the cursor and be skipped, so do not use it
  to enumerate.

## Fetch a single document [#fetch-a-single-document]

`/documents/{id}` returns one document's metadata, processing state and text.
A long text can be read in windows: `contentOffset` and `contentMaxChars`
(counted in UTF-16 code units) select one window, and `content_range` reports
the window and the full length, so keep reading until `offset + length` reaches
`total`.

```bash
curl "https://nordvec.com/api/v1/documents/DOCUMENT_ID?contentMaxChars=100000" \
  -H "Authorization: Bearer $NORDVEC_API_KEY"
```

## Fetch many at once [#fetch-many-at-once]

To resolve up to 200 ids in one call, POST them to `/documents/batch` instead of
making one request per id. The batch returns metadata and processing state;
`content` is always `null`, so read the text with the single-document call.

```bash
curl https://nordvec.com/api/v1/documents/batch \
  -H "Authorization: Bearer $NORDVEC_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "ids": ["DOCUMENT_ID_1", "DOCUMENT_ID_2"] }'
```

<Callout>
  Listing, like search, only ever returns documents the caller is allowed to see.
  `status` tells you where a document is in processing: `processing` while it is
  still being indexed, `failed` when it could not be. See
  [Documents & search](/docs/guides/concepts/documents) for the lifecycle.
</Callout>

## Next steps [#next-steps]

<Cards>
  <Card title="Filter and refine search" href="/docs/guides/how-to/filter-search" />

  <Card title="Documents & search" href="/docs/guides/concepts/documents" />

  <Card title="API reference" href="/docs/api" />
</Cards>
