---
title: "Files and dictation"
description: "Attach documents to an agent turn and transcribe audio, using the shared upload and dictation routes."
canonical_url: "https://scholarxiv.com/developers/docs/agent-api/files"
markdown_url: "https://scholarxiv.com/developers/docs/agent-api/files.md"
---

> For the complete documentation index, see [llms.txt](/llms.txt).

# Files and dictation
URL: /developers/docs/agent-api/files
LLM index: /llms.txt
Description: Attach documents to an agent turn and transcribe audio, using the shared upload and dictation routes.
Related: agent-api, agent-api/quickstart, agent-api/conversations, health-api

# Files and dictation

These routes use the same `sxv_` key and cover the rest of the `/chat` composer. The [Health Chat API](/developers/docs/health-api) uses them with `surface` or `scope` set to `health`.

## Attach a file

Presign, `PUT` the bytes, then complete:

```bash
curl -X POST https://scholarxiv.com/api/v1/uploads/presign \
  -H "Authorization: Bearer sxv_your_key_here" \
  -H "Content-Type: application/json" \
  -d '{"filename":"review.pdf","mime":"application/pdf","size":184320,"scope":"chat"}'
```

`PUT` the bytes to the returned `uploadUrl`, then call `POST /api/v1/uploads/complete` with `{ "id" }`. Pass the completed id in `attachment_ids` on the next agent call — up to five per turn, subject to per-chat plan limits.

`GET /api/v1/uploads/{id}` returns a short-lived `download_url`. That is also how you download any file the agent produced, such as a figure or a Python output, when it comes back with an attachment id.

## Dictate a question

```bash
curl -X POST https://scholarxiv.com/api/v1/transcribe \
  -H "Authorization: Bearer sxv_your_key_here" \
  -F "audio=@memo.m4a" \
  -F "provider=addis"
```

Send it as `multipart/form-data` with `audio` and `provider` (`addis` or `elevenlabs`). Put the returned `text` in `message`.

## Context on the message

Selected papers and highlighted text are `selected_papers` and `selected_texts` on the message. Science viewer context is the `context` object. The server runs the tools; a client does not implement them. Voice memos, figures, and Python outputs come back inside `message.parts` or as stream events.

`GET /api/v1/models?surface=research` lists the models the current plan can run.

## Sitemap

See the full [sitemap](/sitemap.md) for all pages.
Docs-scoped sitemap: [/docs/sitemap.md](/docs/sitemap.md).
Well-known sitemap: [/.well-known/sitemap.md](/.well-known/sitemap.md).
