---
title: "Preprint Phases"
canonical_url: "https://scholarxiv.com/developers/docs/preprint-phases"
markdown_url: "https://scholarxiv.com/developers/docs/preprint-phases.md"
---

> For the complete documentation index, see [llms.txt](/llms.txt).

# Preprint Phases
URL: /developers/docs/preprint-phases
LLM index: /llms.txt

# Preprint Server — Build Phases

Phased plan. Each phase ships something a real user can use, and each has an explicit **not in this
phase** list so scope doesn't creep.

Design detail: [preprint-design-ingest-citations-profiles.md](preprint-design-ingest-citations-profiles.md) ·
Research: [journal-platform-research.md](journal-platform-research.md) ·
[scholar-profiles-and-indexing-research.md](scholar-profiles-and-indexing-research.md)

Sizes are rough relative estimates, not commitments.

---

## Phase 1 — Publish (LaTeX only) · size L

**Goal:** a stranger can post a paper and other people can read it.

Ships:
- **Submit from `/write`** — freeze the project into an immutable version snapshot (files + compiled
  PDF + thumbnail). Reuses the existing Tectonic/E2B pipeline as-is.
- **Metadata form** — title, authors (+ affiliations), abstract, categories, comments, license
  picker. Modelled on arXiv's field rules (abstract ≤1920 chars, author formatting, ASCII-safe).
- **Public paper page** — `/abs/<id>` and `/abs/<id>v2`, SSR'd, abstract visible with no gate,
  PDF download, version history.
- **Browse + search** — by date, by category, by author. Plain server-rendered links.
- **Submission dashboard** — drafts, posted papers, "post a new version".
- **Google Scholar meta tags + sitemap** — nearly free once the pages exist, and skipping them costs
  6–9 months later when Scholar has cached wrong metadata.
- **Basic abuse controls** — account email verification, per-account rate limit (arXiv caps
  back-catalogue posting at 3/day), compile-must-succeed as the accept gate.

**Not in this phase:** DOIs · PDF/DOCX upload · citations · profiles · HTML rendering · OAI-PMH.

**Done when:** you can post a LaTeX paper from `/write` and send someone a public link to it.

---

## Phase 2 — Permanent and indexed · size M

**Goal:** stop being a website with PDFs on it and become a real preprint server.

Ships:
- **Crossref membership + DOI minting** — one DOI per paper, version-specific URLs, the
  `is-preprint-of` relationship so aggregators cluster our copy with the eventual journal version.
- **OAI-PMH endpoint** — the standard feed that OpenAlex, CORE, BASE, Semantic Scholar and OpenAIRE
  harvest. Dublin Core + a native format, hierarchical category sets, nightly refresh. This is what
  actually gets us into the citation databases.
- **PDF hygiene gate** — enforce <5 MB, text layer present, Type 1 fonts (all three are hard Google
  Scholar requirements; failing any means the paper can never be indexed or cited).
- **Withdrawal flow** — arXiv-style: a withdrawal posts a new version marked withdrawn with a
  reason; earlier versions stay reachable. Plus the grace-period decision (see open questions).
- **Robots/canonical/301 discipline** — permanent URLs that never break.

**Not in this phase:** citation counts (we're only making ourselves *findable* here).

**Done when:** search Google Scholar for one of our papers and it's there. Look it up in OpenAlex and
it exists.

---

## Phase 3 — Open the doors (PDF + DOCX) · size M

**Goal:** researchers who don't write LaTeX can use the product. This is what makes
"all disciplines" true rather than aspirational.

Ships:
- **PDF lane** — text extraction with `pdfjs-dist`/`unpdf` (already installed), then an **LLM pass**
  over the first page and reference section to pull title, authors, affiliations, abstract, and
  individual references. Uses `@ai-sdk/google` / `openai` / `ai`, all already installed.
  **Zero new infrastructure.** See the GROBID note below.
- **DOCX lane** — a `.docx` is a ZIP of XML; unzip with `fflate` and parse with `fast-xml-parser`,
  both already installed. Add `mammoth` for clean HTML, or run `pandoc` in the existing E2B sandbox
  for the LaTeX→PDF leg.
- **Confirmation screen** — mandatory for both lanes. Show every extracted field with a confidence
  indicator and require the author to confirm before posting. Non-negotiable: Scholar takes 6–9
  months to correct a bad author name.
- **TeX-derived-PDF nudge** — detect `pdfTeX`/`XeTeX`/`LuaTeX` in the PDF's `/Producer` and offer
  "upload your source instead and get X, Y, Z". Offer, don't block.
- **Tier badges** on paper pages (`source` / `converted` / `PDF only`) and a capability comparison
  shown *before* the author picks a lane.

**Not in this phase:** GROBID.

**Done when:** someone with only a Word file can post a paper end to end.

---

## Phase 4 — Researcher profiles · size M

**Goal:** a researcher's profile is full and impressive *before* they've posted anything here.

Ships:
- **Import + claim** — pull their publication history from OpenAlex, show it as a checklist, they
  confirm what's theirs. Unconfirmed matches never count toward public metrics.
- **Public profile page** — name, affiliations (with ROR), interests, links, works list (papers
  hosted here + claimed external ones), co-author list.
- **Metrics** — citations, h-index, i10-index, All and Recent-5y, plus a per-year chart. Fetched
  from OpenAlex/Semantic Scholar and cached. **Every number labelled with its source and age.**
- **Re-sync** button and scheduled refresh.

**Not in this phase:** our own citation graph — this phase borrows everything from OpenAlex, which
is why it's cheap and why it works on day one.

**Done when:** a researcher completes their profile and sees their ScholarXIV papers and
an h-index they didn't have to type in.

---

## Phase 5 — Citations · size L

**Goal:** papers show who cites them; profiles compute metrics from data we own.

Ships:
- **Reference extraction and resolution** — from `.bbl`/`.bib` (LaTeX), the bibliography section
  (DOCX), or the LLM pass (PDF). Resolve each reference via DOI regex → arXiv ID → Crossref
  `query.bibliographic` (score-and-margin threshold) → OpenAlex → internal corpus → unresolved,
  retried nightly.
- **Internal citation graph** — instant, exact, and the only citation data a brand-new preprint has.
  Surfaces as "Cited by N papers on ScholarXIV".
- **External citation polling** — the `nextCheckAt` cron described in the design doc. Batched (50
  DOIs per OpenAlex call, 500 per Semantic Scholar batch), tiered freshness, snapshots-not-overwrites
  so the history chart is free.
- **Crossref reference deposit** — turns on Cited-by and stops us under-reporting our own counts by
  the ≥20% Crossref warns about.
- **Citation displays** — counts with source attribution, "cited by" lists, citations-over-time.
- **Profile metrics recomputed** from our own confirmed-works graph, shown alongside the fetched
  numbers.

**Done when:** a paper page shows a real, sourced citation count that matches what a researcher sees
elsewhere.

---

## Phase 6 — Depth and defence · size L, mostly optional

Pick from this menu based on what's actually hurting by then:

- **HTML full text via LaTeXML** — the converter arXiv itself uses. Big win for accessibility,
  mobile, and search ranking. *Promote to Phase 3 if SEO turns out to matter more than expected.*
- **GROBID** — swap in for the LLM extraction pass if quality data says we need it (see below).
- **Version diffs** — semantic diffs between v1 and v2, which only LaTeX submissions can support.
- **Co-author graph visualisation** and topic maps over a researcher's corpus.
  Google Scholar has never done it.
- **Full anti-abuse suite** — similarity checking, scope classifier, figure-integrity checks,
  retracted-reference warnings, paper-mill signals.
- **JATS XML export**, public profile API, JSON-LD.
- **Withdrawal/correction record types** if the corpus starts needing them.

---

## Appendix — GROBID, Docker, and E2B

**Does GROBID need Docker?** In practice yes. It's a Java service and the supported distribution is
Docker images. Two of them:

| Image | Size | Notes |
|---|---|---|
| `grobid/grobid:0.9.0-crf` | **~500 MB** | CRF models only. Best runtime and memory. 2–4 F1 points lower on reference parsing, 2–5 lower on citation contexts. |
| `grobid/grobid:0.9.0-full` | **~8 GB** | Adds TensorFlow (3 GB) + embeddings (5 GB). GPU recommended; on CPU the DL models are "much slower with significantly higher memory usage." |

Both run as an HTTP service on port 8070. **amd64 only** — they do not run on linux/arm64 and are
only emulated (slowly) on Apple Silicon.

*(The ~0.87–0.90 F1 reference-extraction figures quoted earlier are the **full** image's deep-learning
model. The 500 MB CRF image is a few points below that.)*

**Can E2B run it?** Technically yes. E2B templates are built in a full sandbox environment that can
run Docker containers during setup, and `setStartCmd` captures a running process into the snapshot —
so a sandbox could boot with GROBID already listening. The CRF image at 500 MB fits comfortably
inside the 10 GB Hobby disk limit; the 8 GB full image would not, and its 8 GB memory need is at the
plan ceiling.

**But it's the wrong shape.** E2B sandboxes are ephemeral and per-session (1 h on Base, 24 h on Pro).
GROBID is a long-running stateless HTTP service whose expensive step is loading models into memory at
startup. Using E2B means paying that startup cost repeatedly for a job that takes a fraction of a
second once warm. E2B is the right tool for LaTeX compiles — untrusted, per-user, isolated. GROBID is
none of those things.

**If we ever want GROBID**, the right home is a small always-on container on Fly.io / Railway /
Render / a cheap VPS running the CRF image — roughly 2 GB RAM, on the order of $5–10/month — called
over HTTP from Vercel.

**Recommendation: don't build any of that in Phase 3.** Text extraction is already solved in our
stack (`pdfjs-dist`, `unpdf`) and the "which line is the title, where do the references start" problem
is something an LLM handles well — and we already have three LLM SDKs installed. Ship the PDF lane
with zero new infrastructure, measure how often authors correct the confirmation screen, and add
GROBID in Phase 6 only if that number is bad.

---

## Open questions before Phase 1

1. **Withdrawal grace period** — arXiv/bioRxiv say posted is permanent forever. Copy that, or allow
   deletion within a short window (e.g. 24 h) before the DOI is minted? *(Recommendation: 24 h grace,
   permanent after. Mint the DOI at the end of the window.)*
2. **Category taxonomy** — adopt arXiv's list, or extend it? *(Recommendation: arXiv's as the base,
   plus categories for biology, medicine, and humanities, which arXiv covers poorly.)*
3. **Phase 1 scope check** — is "LaTeX only, no DOIs" an acceptable first release, or does Phase 2
   need to merge into Phase 1 so the first public papers are citable from day one?

## Sitemap

See the full [sitemap](/sitemap.md) for all pages.
Docs-scoped sitemap: [/docs/sitemap.md](/docs/sitemap.md).
Well-known sitemap: [/.well-known/sitemap.md](/.well-known/sitemap.md).
