# Push documents from your own systems
Source: https://nordvec.com/docs/guides/how-to/push-documents

Create a datasource, push documents into it with an indexing API key, choose who can read them, and pause or delete it when the source changes.



The push API indexes documents from systems Nordvec has no connector for: an
internal wiki export, a ticket archive, a database of notes. You send the text
and who may read it; Nordvec stores it in the EU, indexes it, and makes it
searchable and citable like any other document. Every pushed document lands in
a **datasource**, a named container in your workspace that a workspace admin
creates first. A push that names a datasource which does not exist, or one that
is paused, is refused.

## Create a datasource [#create-a-datasource]

Open **Workspace settings > Datasources** and choose **Create datasource**.
Workspace admins and owners can do this; in a personal workspace, that is you.

| Field | Notes                                                                                                                                             |
| ----- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
| Name  | What people see in the settings list. Up to 200 characters.                                                                                       |
| Slug  | What every push names. Lowercase letters, digits, `-` and `_`, starting with a letter or digit, up to 200 characters. It cannot be changed later. |

The slug `confluence-export` is used in the examples below.

## Create an indexing API key [#create-an-indexing-api-key]

Pushes authenticate with an API key of the **Indexing** class that carries the
`index:write` scope; add `index:status` to track ingestion and `index:delete`
to remove documents or replace a whole datasource. Create one under **Workspace settings > API keys**; the
raw key starts with `nv_eu_idx_` and is shown once. See
[Authentication](/docs/guides/authentication). Each request also names your
workspace id as `tenantId`, the id in your workspace's address in the app
(`/w/<workspace id>/...`), and it must be the workspace the key belongs to.

## Push one document [#push-one-document]

`/documents/push` creates the document, or updates it when a document with the
same `id` already exists in the datasource.

```bash
curl https://nordvec.com/api/v1/documents/push \
  -H "Authorization: Bearer $NORDVEC_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: page-4711-2026-09-28" \
  -d '{
    "tenantId": "YOUR_WORKSPACE_ID",
    "document": {
      "id": "page-4711",
      "title": "Travel expense policy",
      "datasource": "confluence-export",
      "body": { "mimeType": "text/markdown", "content": "# Travel expenses\n..." },
      "permissions": {},
      "sourceUrl": "https://wiki.example.com/pages/4711",
      "type": "policy"
    }
  }'
```

```json
{ "documentId": "page-4711", "status": "queued", "updated": false }
```

* `id` is your stable id for the document within the datasource. Pushing the
  same `id` again updates it; unchanged content is recognised by its hash and
  not indexed twice.
* `body.mimeType` is one of `text/plain`, `text/markdown`, `text/html`,
  `application/pdf`, or the Word, Excel and PowerPoint (`.docx`, `.xlsx`,
  `.pptx`) types. Binary content is sent base64-encoded.
* `sourceUrl` becomes the "jump to source" link on every citation of the
  document. Omit it on a re-push to keep the stored one, or send `null` to
  clear it.
* `type` sets the document's `content_type`, which search and list filter on.

The whole request body is capped at 1 MB, so a large file or a big batch
answers `413`; split it.

## Choose who can read it [#choose-who-can-read-it]

`permissions` is required on every push, so a sharing decision is never made
by leaving a field out. In a datasource visible to the workspace:

| `permissions`                             | Who can read the document                                  |
| ----------------------------------------- | ---------------------------------------------------------- |
| `{}`                                      | Every member of the workspace                              |
| `{ "allowedUsers": ["ana@example.com"] }` | Only the people listed                                     |
| `{ "allowedGroups": ["GROUP_ID"] }`       | Members of those workspace groups, including nested groups |
| `{ "allowAllTenantMembers": false }`      | Refused: a document nobody can read is a delete            |

To change who can read a document without sending its content again, use
`POST /documents/push/permissions`. Making an already-restricted document
visible to the whole workspace additionally needs the `index:acl-widen` scope,
so a routine sync cannot quietly undo a restriction someone set by hand.

## Push in batches [#push-in-batches]

`/documents/push/bulk` takes up to 100 documents for one datasource per call.
The answer counts `accepted` and `rejected` and gives a result per document, so
one bad document does not fail the batch. The 1 MB body cap applies per call,
so split large uploads into several calls under the same `uploadId`.

```bash
curl https://nordvec.com/api/v1/documents/push/bulk \
  -H "Authorization: Bearer $NORDVEC_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "tenantId": "YOUR_WORKSPACE_ID",
    "uploadId": "nightly-2026-09-28",
    "datasource": "confluence-export",
    "documents": [
      { "id": "page-4711", "title": "Travel expense policy", "datasource": "confluence-export",
        "body": { "mimeType": "text/plain", "content": "..." }, "permissions": {} }
    ]
  }'
```

## Replace a whole datasource [#replace-a-whole-datasource]

When your system can list everything a datasource should hold, send the full
listing as one **upload session**, and the documents it no longer contains are
moved to the trash when the session closes. Sessions need an indexing API key
holding `index:delete` as well as `index:write`, because the close removes
documents; the key that opens one is the only key that can continue it.

1. Send the first page with `"isFirstPage": true`. It is page `0`.
2. Send every further page with its `pageIndex` (`1`, `2`, ...), in any order.
   A page sent twice counts once, so a retry is always safe.
3. Send the last page with `"isLastPage": true` and its `pageIndex`. A listing
   that fits in one page sends `isFirstPage` and `isLastPage` together. The
   last page may carry no documents.

Every page uses the same `uploadId`, and each answer carries the session's
progress under `upload`. The session closes only when every page from `0` to
the last has arrived. Closing it moves to the trash each document in the
datasource that no page of the session named and that existed before the
session opened. Any other push to the datasource while the session runs keeps
the document it names: a single push, a batch without session fields, a
permissions update, and a re-push of unchanged content alike. The trash keeps
what the close moved there for 30 days; pushing a document again brings it
back, and so does restoring the whole session (see below).

A session that receives no page for 24 hours expires and closes without
removing anything. A refused page is answered with `409 Conflict`, writes
nothing, and its `data.reason` says why:

| `reason`                                               | What to do                                                                                                                                                                                                                                                                                                  |
| ------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `upload_incomplete`                                    | Send the pages listed in `missingPageIndexes`, then the last page again                                                                                                                                                                                                                                     |
| `deletion_confirmation_required`                       | The close would trash more than 20% of the datasource. If that is right, send the last page again with `"confirmDeletions"` set to `wouldTombstone`                                                                                                                                                         |
| `deletion_confirmation_too_large`                      | `confirmDeletions` is larger than the number of documents the datasource held when the session opened. Send the count you expect to remove                                                                                                                                                                  |
| `upload_in_progress`                                   | A session is open on this datasource. If it is your key's, finish it, wait for it to expire, or start over with `"forceRestartUpload": true` on your first page. If another key opened it, `forceRestartUpload` replaces it only once it has received no page for an hour, from the time in `restartableAt` |
| `upload_expired`, `upload_missing`, `upload_restarted` | The session is gone; start a new one with a new `uploadId`                                                                                                                                                                                                                                                  |
| `upload_closed`, `upload_id_reused`                    | The `uploadId` is spent; use a new one                                                                                                                                                                                                                                                                      |
| `page_index_required`                                  | Your key has a session open on this datasource; send `pageIndex` with the page                                                                                                                                                                                                                              |

To resume after a crash, read the session with
`GET /documents/push/upload?tenantId=...&datasource=...&uploadId=...` (scope
`index:status`). Its `missingPageIndexes` lists the pages still to send.

### Undo a session's close [#undo-a-sessions-close]

If a session removed documents it should not have, for example because the
listing it sent was cut short, restore them in one call:

```bash
curl https://nordvec.com/api/v1/documents/push/upload/restore \
  -H "Authorization: Bearer $NORDVEC_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "tenantId": "YOUR_WORKSPACE_ID", "datasource": "confluence-export", "uploadId": "nightly-2026-09-28" }'
```

The key that opened the session can restore it, and so can a workspace admin
signed in to Nordvec, for a session opened by any key. Every document the close
moved to the trash comes back with the content it had, and the answer counts
them: `restored` are live again, `purged` had already been deleted for good by
the trash, and `skipped` had changed since the close (pushed again, or removed
again) and were left as they are. Restoring a session twice answers with the
first restore's counts and `"replayed": true`, and queues any restored document
still waiting to be indexed, so repeating a restore that failed to answer is
safe. A session can be restored for 35 days after it closed, and for as long
as the trash still holds any document it removed. A refused restore is answered
with `409 Conflict` and its `data.reason`:

| `reason`                 | What it means                                                                                                                                                                                      |
| ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `upload_not_closed`      | The session never closed, so it removed nothing                                                                                                                                                    |
| `upload_in_progress`     | A session is open on the datasource. Restore once it has closed or expired                                                                                                                         |
| `restore_purged`         | More than 30 days have passed, and the trash has deleted every one of the documents. Push them again                                                                                               |
| `workspace_not_entitled` | The workspace's plan does not currently allow restoring from the trash                                                                                                                             |
| `corpus_cap_exceeded`    | Bringing the documents back would pass the workspace's document limit, so none came back. `data.wouldRestore` is how many it needs and `data.headroom` how many fit. Free room, then restore again |

## Track ingestion [#track-ingestion]

A push answers as soon as the document is queued. Ask for its progress with
`GET /documents/push/status` (scope `index:status`), filtered by datasource or
document id. A document moves from `queued` through `processing` to
`completed`, or to `failed` with an `error`.

## When a push is refused [#when-a-push-is-refused]

A push that names an unknown or paused datasource is answered with
`422 Unprocessable Content`. The message names the slug and links to
**Workspace settings > Datasources** in your workspace, and the error's `data`
says why and what to do:

```json
{
  "defined": true,
  "code": "UNPROCESSABLE_CONTENT",
  "status": 422,
  "message": "Datasource \"confluence-export\" is paused and accepts no documents. A workspace admin resumes it under Workspace settings > Datasources: https://nordvec.com/w/YOUR_WORKSPACE_ID/workspace/settings?tab=datasources",
  "data": {
    "why": "The datasource \"confluence-export\" is paused",
    "fix": "Resume it at https://nordvec.com/w/YOUR_WORKSPACE_ID/workspace/settings?tab=datasources, then retry the push",
    "link": "https://nordvec.com/docs/guides/how-to/push-documents"
  }
}
```

Do not retry these automatically: they succeed only after an admin creates or
resumes the datasource.

## Remove a document [#remove-a-document]

`POST /documents/push/delete` (scope `index:delete`) removes a pushed document
by its `datasource` and `id`. Documents you stop pushing are not removed on
their own: delete each one you retire, or send the datasource's full listing as
an upload session, described above.

## Pause, resume and delete [#pause-resume-and-delete]

* **Pause** refuses every further push into the datasource. Its documents stay
  searchable. A push already being written when you pause completes.
* **Resume** accepts pushes again.
* **Delete** removes the datasource and every document pushed to it, with
  their search index. Your own system keeps its copy, so pushing again after
  you recreate the datasource restores them. A delete cannot be undone.

If another admin changed the datasource after your list loaded, the action is
refused and the list reloads, so you decide again against what is there now.
Every create, pause, resume and delete is recorded in the workspace audit log.

<Callout>
  The settings list shows who each datasource is visible to. Who can read a
  pushed document is decided by the `permissions` sent with it; creating,
  pausing or deleting a datasource never widens access to anything.
</Callout>

<Callout>
  Push writes are idempotent: repeat the same `Idempotency-Key` on every retry of
  one write, and a duplicate is answered from the first attempt instead of being
  applied twice. See [Errors and rate limits](/docs/guides/errors-and-rate-limits).
</Callout>

## Next steps [#next-steps]

<Cards>
  <Card title="List and retrieve documents" href="/docs/guides/how-to/list-documents" />

  <Card title="Filter and refine search" href="/docs/guides/how-to/filter-search" />

  <Card title="API reference" href="/docs/api" />
</Cards>
