---
title: "Preprint Design Ingest Citations Profiles"
canonical_url: "https://scholarxiv.com/developers/docs/preprint-design-ingest-citations-profiles"
markdown_url: "https://scholarxiv.com/developers/docs/preprint-design-ingest-citations-profiles.md"
---

> For the complete documentation index, see [llms.txt](/llms.txt).

# Preprint Design Ingest Citations Profiles
URL: /developers/docs/preprint-design-ingest-citations-profiles
LLM index: /llms.txt

# Design: Three-Format Ingest, Citations, and Researcher Profiles

Design doc for the preprint server. Decisions it implements: arXiv-style, no review, auto-listed,
all disciplines, free, **accepting LaTeX + PDF + DOCX**.

Research backing: [journal-platform-research.md](journal-platform-research.md) ·
[scholar-profiles-and-indexing-research.md](scholar-profiles-and-indexing-research.md)

---

## 1. Ingest — three lanes, one normalized object

### 1.1 The canonical artifact set

Whatever comes in, every accepted submission version produces the same set of artifacts. The rest of
the system only ever talks to these — nothing downstream knows or cares which lane produced them.

| Artifact | Purpose | Always present? |
|---|---|---|
| `original/` | The exact bytes the author uploaded, archived forever | ✅ |
| `paper.pdf` | Canonical rendering. **<5 MB, text layer, Type 1 fonts** | ✅ |
| `paper.html` | Accessible/mobile/indexable full text + MathML | Lanes A, B |
| `document.json` | Structured AST: sections, figures, tables, equations | Lanes A, B |
| `metadata.json` | Title, authors, affiliations, abstract, categories, license | ✅ |
| `references.json` | Parsed reference list, one entry per citation | ✅ (quality varies) |
| `fulltext.txt` | Plain text for search + similarity + classification | ✅ |
| `thumb.png` | Page-one thumbnail (already produced by `/write`) | ✅ |

### 1.2 Lane A — LaTeX source (`tier: source`)

```
upload (.tex/.bib/figures, or .zip/.tar.gz)
  → Tectonic in E2B                    → paper.pdf          [already built]
  → LaTeXML                            → paper.html + MathML + document.json
  → parse \author/\title/\begin{abstract} + .bbl/.bib
                                       → metadata.json, references.json
  → pdftotext                          → fulltext.txt
```

**LaTeXML is the right converter** — it's what arXiv itself uses for HTML papers. arXiv's own caveat
applies to us too: it is "a very extensible language used in myriad unique ways," so a small
percentage of papers won't convert. **HTML conversion failure must never block posting** — the paper
goes live with PDF only and a "HTML unavailable" note, exactly as arXiv does.

References come from the `.bbl` (or `.bib`) — this is structured data, not guesswork, so Lane A gives
us the best citation graph by a wide margin.

### 1.3 Lane B — DOCX (`tier: converted`)

```
upload (.docx)
  → pandoc → structured AST            → document.json, paper.html (OMML → MathML)
  → pandoc → .tex → Tectonic           → paper.pdf
  → AST heading/style walk             → metadata.json (title, authors, abstract)
  → bibliography section → GROBID       → references.json
  → AST text                           → fulltext.txt
```

DOCX carries real structure — styles, heading levels, footnotes, equations as OMML — so this
conversion is genuinely good. It is nothing like trying to reconstruct a PDF. Keep the original
`.docx` archived; if pandoc improves, we can re-run the whole corpus.

**Author confirms the extracted metadata** on a review screen before posting (§1.6).

### 1.4 Lane C — PDF (`tier: unstructured`)

```
upload (.pdf)
  → validate: text layer present, ≤5 MB, not encrypted
  → GROBID /processFulltextDocument    → TEI XML
        ├─ header model                → metadata.json (title, authors, affiliations, abstract)
        ├─ reference model             → references.json (+ DOI/PMID where present)
        └─ body segmentation           → fulltext.txt, partial document.json
  → original PDF passes through as paper.pdf
```

**GROBID is the tool.** Purpose-built, open source, Docker + REST, and production-proven at
ResearchGate, Academia.edu, HAL, Mendeley, CERN and the Internet Archive. Published accuracy:

- Reference extraction/parsing: **~0.87 F1** on a PubMed Central set (90k references), **~0.90 F1**
  on a bioRxiv PDF set, using the deep-learning citation model.
- Isolated reference parsing: **>0.90 F1** instance-level, **0.95 F1** field-level.
- DOI resolution from extracted references: **>0.95 F1** (via Crossref REST or biblio-glutton).
- Throughput: ~2.5 PDF/sec full processing on 8 threads; ~4 GB RAM for the full pipeline.

That accuracy is good enough to be useful and **not** good enough to trust blindly — hence §1.6.

No HTML view, no version diffs, no JATS. The paper page shows a `PDF only` tier badge.

### 1.5 The TeX-derived-PDF nudge (not a block)

arXiv hard-rejects PDFs whose `/Producer` or `/Creator` is a TeX engine (`pdfTeX`, `XeTeX`, `LuaTeX`).
You chose to accept PDF, so we **detect and nudge** instead:

> *"This PDF was produced by pdfTeX. Uploading your LaTeX source instead gives your paper an HTML
> version, better search ranking, accurate citation extraction, and version diffs — and takes about
> 30 seconds."* — with buttons **[Upload source]** and **[Continue with PDF]**.

Cheap to implement, preserves author choice, and reclaims most of the value of the hard rule.

### 1.6 The confirmation screen — the single most important UI in the ingest

Lanes B and C extract metadata by machine. **Never post machine-extracted metadata unconfirmed.**
Google Scholar takes 6–9 months to correct an indexed record, so a wrong `citation_author` is
effectively permanent.

After processing, show the author a form pre-filled with everything we extracted — title, each
author + affiliation, abstract, and the parsed reference list with a per-entry confidence indicator —
and require an explicit confirm. Lane A shows the same screen but pre-filled from source, so it's
usually a one-click pass.

### 1.7 Capability matrix (surface this to authors *before* they choose a lane)

| | LaTeX | DOCX | PDF |
|---|---|---|---|
| Posts publicly | ✅ | ✅ | ✅ |
| DOI, OAI-PMH, Scholar meta tags | ✅ | ✅ | ✅ |
| HTML full text + MathML | ✅ | ✅ | ❌ |
| Metadata accuracy | exact | good | extracted, ~0.87–0.9 F1 |
| Reference extraction | exact (`.bbl`) | good | extracted |
| Version diffs | semantic | textual | page images only |
| JATS export | ✅ | ✅ | ❌ |
| Compile-time validation gate | ✅ | ✅ | ⚠️ format checks only |

---

## 2. Citations

Two independent directions. Conflating them is the classic mistake.

```
        references.json                    external indexes
     (what THIS paper cites)          (what cites THIS paper)
              │                                  │
        OUTBOUND edges                     INBOUND counts
              └──────────► citation_edges ◄──────┘
                                │
                          profile metrics
```

### 2.1 Outbound — resolving our reference lists

Every reference string goes through a resolution ladder, stopping at the first confident hit:

1. **Explicit DOI** in the string → `10.\d{4,9}/\S+` regex → validate against Crossref.
2. **Explicit arXiv ID** → `arXiv:\d{4}\.\d{4,5}` or the old `archive/YYMMNNN` form.
3. **Other IDs** — PMID, ISBN, ACL/DBLP handles.
4. **Crossref bibliographic query** — `GET /works?query.bibliographic=<raw string>&rows=2`.
   Verified live: querying *"LeCun Bengio Hinton Deep learning Nature 2015 521 436"* returns
   `10.1038/nature14539` at **score 81.1** with the runner-up at **46.8**. The score gap is the
   confidence signal — accept above a tuned absolute threshold *and* a minimum margin over #2.
   Always send `mailto=` to stay in Crossref's polite pool.
5. **OpenAlex title search** as a second opinion for things Crossref doesn't have (preprints,
   dissertations, non-DOI works).
6. **Internal match** against our own corpus by DOI / arXiv id / normalized title.
7. **Unresolved** — store the raw string anyway. It still renders in the reference list, and a
   nightly job retries unresolved edges as external indexes catch up.

```jsonc
// citation_edge
{
  "fromWorkId": "sx:2607.00123",         // always one of ours
  "fromVersion": 1,
  "rawString": "LeCun Y, Bengio Y, Hinton G. Deep learning. Nature 521, 436–444 (2015).",
  "position": 12,
  "toDoi": "10.1038/nature14539",
  "toOpenAlexId": "W2117539524",
  "toWorkId": null,                      // set when the target is also hosted here
  "method": "crossref_bibliographic",
  "confidence": 0.93,
  "resolvedAt": "2026-07-26T…"
}
```

### 2.2 Inbound — who cites us

Three sources, unioned and deduplicated by DOI:

**(a) Internal, instant.** Any ScholarXIV paper citing another ScholarXIV paper is an edge we already
own. Zero latency, zero cost, 100% accurate — and it's the *only* citation data that exists for a
brand-new preprint. Surface it as **"Cited by N papers on ScholarXIV."**

**(b) External indexes, keyed by our DOI.** Nightly/weekly polling:
- OpenAlex `cited_by_count` + the `cites:` filter for the actual citing list.
- Semantic Scholar `/paper/DOI:…/citations` (also accepts `ARXIV:` ids), plus
  `influentialCitationCount`.
- Crossref `is-referenced-by-count`.
- OpenCitations for a fully-open third opinion.

**(c) Crossref Cited-by.** Free, but it only returns useful data **if we deposit our own reference
lists**. Crossref's own warning: not depositing under-reports counts **by ≥20%**. Since §2.1 already
produces resolved reference lists, depositing them is nearly free — do it.

Store per-source counts, never a single blended number:

```jsonc
// work.citationCounts
[ { "source": "internal",         "count": 3,   "fetchedAt": "…" },
  { "source": "openalex",         "count": 47,  "fetchedAt": "…" },
  { "source": "semanticscholar",  "count": 51,  "influential": 7, "fetchedAt": "…" },
  { "source": "crossref",         "count": 44,  "fetchedAt": "…" } ]
```

Display the max as the headline with the source named, and the breakdown on hover. **Never show a
number without its source and age.** This is the fastest way to lose credibility with researchers,
who will absolutely cross-check against Scholar.

### 2.3 The cold-start reality

A brand-new preprint has **zero** citations everywhere, and stays at zero until Scholar and the
aggregators index it — weeks at best. Nothing we build changes that. The product answer is §3.3:
profiles are seeded from the researcher's **existing** work, not from what they post here.

---

## 3. Researcher profiles

### 3.1 Model

```
ResearcherProfile
  userId, handle, visibility: private | public
  displayName, nameVariants[]          ← seeded from OpenAlex raw_author_names
  photoUrl, bio, homepageUrl, socials[]
  affiliations[]  { name, rorId, department, role, startYear, endYear, isCurrent }
  verifiedEmails[]{ address, isInstitutional, verifiedAt, public: false }
  interests[], topics[]
  identifiers    { openAlexId, semanticScholarId, arxivAuthorId, dblp, googleScholarUrl }
  settings       { autoAddWorks, publicMetrics, showEmail: false }

ProfileWork
  profileId, workId
  origin: native | claimed_external
  claimState: suggested | confirmed | rejected     ← only `confirmed` counts toward metrics
  identifiers { doi, arxivId, openAlexId, s2PaperId }
  authorPosition, isCorresponding
  mergedInto?                                       ← Scholar-style duplicate consolidation

ProfileMetrics (cached, recomputed by job)
  source, fetchedAt
  worksCount, citations, hIndex, i10Index
  citationsRecent5y, hIndexRecent5y, i10IndexRecent5y
  perYear: [{ year, publications, citations }]
```

### 3.2 Claiming — the disambiguation ladder

Author name disambiguation is the unsolved problem of this domain. Strongest signal first, and
**nothing below rung 3 ever auto-confirms**:

1. **Native works** — anything submitted through ScholarXIV by that account.
2. **Verified email match** against corresponding-author metadata.
3. **Candidate suggestions** — OpenAlex author entities whose `raw_author_names` overlap, weighted by
   affiliation and topic overlap. Presented as a checklist the human confirms. This is precisely
   Scholar's "select your articles from clusters" step, and it works because the human does the last
   mile.
4. **Manual add** by DOI / arXiv ID / pasted citation.
5. **Merge / split**, with the rule that merging must not double-count a paper citing both versions.

### 3.3 Solving day one

**The insight that makes profiles worth building:** a new user's ScholarXIV paper count is zero and
their citation count is zero. If that's what the profile shows, nobody returns.

The onboarding should let users build a useful profile from their ScholarXIV work before they post a
preprint.

The profile is the acquisition surface. The preprint server is what they stay for.

### 3.4 Metrics — compute *and* fetch

Do both, and show both:

- **Fetch** OpenAlex `summary_stats` (`h_index`, `i10_index`, `2yr_mean_citedness`) and Semantic
  Scholar `hIndex` — these are authoritative for the author's global record.
- **Compute** h-index and i10-index ourselves over the confirmed, deduplicated work list, so the
  numbers stay correct when a user claims a paper OpenAlex hasn't linked to them.
- Offer the **All / Recent (5y)** split that Scholar users expect.
- Definitions, matching Scholar exactly: **h-index** = largest *h* such that *h* works each have ≥*h*
  citations. **i10-index** = count of works with ≥10 citations.
- If our computed number differs from the fetched one, show ours and label it
  *"computed from N confirmed works"*.

### 3.5 Refresh strategy

Never on page load. A queue with tiered cadence:

| Job | Cadence |
|---|---|
| `resolve_references` | on post, retry unresolved nightly |
| `refresh_work_citations` (our papers) | daily for <90 days old, weekly after |
| `refresh_profile_metrics` | daily if profile viewed in last 7 days, else weekly |
| `discover_claim_candidates` | weekly per profile |
| `oai_pmh_reindex` | nightly, after the day's papers go live |

---

## 4. Build order

1. **Lane A end-to-end** (LaTeX → PDF + HTML + metadata + references) — extends what `/write` does.
2. **Paper landing pages** — SSR, unique URL per version, Highwire `citation_*` meta tags, sitemap,
   browse-by-date/category/author within 10 clicks.
3. **PDF hygiene gate** — <5 MB, text layer, Type 1 font lint.
4. **DOIs (DataCite)** + version and `is-preprint-of` relationships.
5. **OAI-PMH endpoint** — `oai_dc` + native format, hierarchical sets, nightly.
6. **Lane C (PDF + GROBID)** — biggest audience unlock per unit of work.
7. **Confirmation screen** — required by lanes B and C.
8. **Reference resolution + internal citation graph** → "Cited by on ScholarXIV".
9. **Profile v1** (identity and native works).
10. **Metrics v1** — OpenAlex ingest, cached, attributed, All/Recent split.
11. **Lane B (DOCX + pandoc)**.
12. **Claim & merge UI**, external citation polling, Crossref reference deposit.
13. **Differentiators** — co-author graph, topic map, cited-by reading list, public profile API.

---

## 5. Open decisions

1. **DOI registry — RESOLVED: Crossref.** (Reverses an earlier recommendation of DataCite; real
   pricing, fetched 2026-07-26, says otherwise at our volume.)

   | | Crossref | DataCite |
   |---|---|---|
   | Annual membership | **$275** (tier: $1k–$1M publishing revenue/expenses; $200 below that) | **€2,000** |
   | Per-DOI | **$0.25** per preprint (current); $0.15 back-year | **€0** up to 100k/year |
   | Extra for for-profits | none | **+€1,000–€2,500** shared infrastructure fee by revenue |
   | Realistic year-one total | **~$400** at 500 papers | **~$3,250** |

   **Crossover is ≈12,000 papers/year** — below that Crossref is cheaper, above it DataCite wins.
   That's why arXiv (≈200k/year) uses DataCite and why we shouldn't, yet.

   Crossref also gives us Cited-by, reference deposit, and `is-referenced-by-count` — the citation
   plumbing in §2 — on the same membership. Check eligibility for Crossref's **GEM programme**
   (fee waivers) before paying anything. Revisit DataCite if we pass ~10k papers/year; old DOIs keep
   resolving, we'd just run two prefixes.
2. **Permanence** — a minted DOI is a promise. arXiv/bioRxiv: posted = permanent, withdrawal marks a
   version rather than deleting. Adopt as-is, or allow deletion in an initial grace window?
3. **GROBID hosting** — self-host the Docker image (needs ~4 GB RAM, gives ~2.5 PDF/sec) or run it as
   a separate service alongside the E2B sandboxes?
4. **Category taxonomy** — adopt arXiv's (familiar, maps cleanly for cross-posting) or build our own
   for the non-physics disciplines arXiv covers poorly?

## Sitemap

See the full [sitemap](/sitemap.md) for all pages.
Docs-scoped sitemap: [/docs/sitemap.md](/docs/sitemap.md).
Well-known sitemap: [/.well-known/sitemap.md](/.well-known/sitemap.md).
