---
title: "Scholar Profiles And Indexing Research"
canonical_url: "https://scholarxiv.com/developers/docs/scholar-profiles-and-indexing-research"
markdown_url: "https://scholarxiv.com/developers/docs/scholar-profiles-and-indexing-research.md"
---

> For the complete documentation index, see [llms.txt](/llms.txt).

# Scholar Profiles And Indexing Research
URL: /developers/docs/scholar-profiles-and-indexing-research
LLM index: /llms.txt

# Discoverability, Citations & Researcher Profiles — Research Dossier

Companion to [journal-platform-research.md](journal-platform-research.md). Covers two questions:

1. **Should we accept PDF and DOCX alongside required LaTeX source?** (§1)
2. **How do Google Scholar indexing, citation counting, and author profiles actually work — and how
   do we build a researcher profile with real citation data?** (§2–§7)

Sources are primary and fetched July 2026. API field shapes in §4 were verified with live calls, not
recalled.

---

## 1. PDF and DOCX — the elaboration you asked for

### 1.1 What the incumbents actually do

| Server | LaTeX source | PDF | DOCX |
|---|---|---|---|
| **arXiv** | **preferred** | accepted **only if not TeX-derived** | **rejected outright** |
| bioRxiv / medRxiv | rejected (must convert to PDF first) | accepted | accepted → auto-converted to PDF |
| Nature | accepted at acceptance stage only, then converted to Word | initial submission only | **preferred** |
| PLOS ONE | rejected (submit the PDF) | accepted for LaTeX users | accepted (DOC/DOCX/RTF) |

arXiv's rule is the sharp one: *"We do not accept dvi, PS, or PDF created from TeX/LaTeX source."*
Their stated reason is portability and stability over time — a PDF is a dead end, source can be
re-rendered forever.

**That rule is mechanically enforceable.** Every TeX engine stamps the PDF's `/Producer` and
`/Creator` metadata (`pdfTeX-1.40.x`, `XeTeX`, `LuaTeX`). We can detect a TeX-derived PDF in
milliseconds and reject it with "you have the source — upload it." No judgement call needed.

### 1.2 The tension in your own decisions

You chose **all disciplines** *and* **LaTeX required**. Those pull against each other, hard:

- LaTeX is the norm in physics, math, CS, astronomy, quantitative economics, and statistics.
- Word is the norm in **biology, medicine, chemistry, psychology, most social sciences, and the
  humanities** — the majority of the world's researchers by headcount.

bioRxiv exists precisely because life scientists don't write TeX. If we require source-only, we
build an arXiv competitor, not an all-disciplines preprint server. That's a fine product — it's just
a different one than "all disciplines" implies, and it should be a deliberate choice rather than a
side effect.

### 1.3 What we gain by having the source (this is the real argument)

The compile pipeline you already have (`src/lib/write/` + Tectonic in E2B) makes source-in
disproportionately valuable:

| Capability | With LaTeX source | With DOCX | With PDF only |
|---|---|---|---|
| **Automated accept gate** (compiles = valid) | ✅ free, replaces the human moderator you're not hiring | ⚠️ via conversion | ❌ nothing to check |
| **HTML full text** (accessibility, mobile, SEO, Scholar likes it) | ✅ | ✅ | ❌ |
| **Structured metadata** (title/authors/abstract/refs read from source) | ✅ exact | ✅ good | ❌ guessed by a parser |
| **Reference list extraction** → citation graph | ✅ from `.bbl`/`.bib` | ✅ | ⚠️ heuristic |
| **MathML / accessible equations** | ✅ | ⚠️ | ❌ |
| **Semantic version diffs** ("what changed in v2") | ✅ | ⚠️ | ❌ |
| **JATS XML export** (indexers, PMC, preservation) | ✅ | ✅ | ❌ |
| **Format migration in 20 years** | ✅ | ✅ | ⚠️ |
| **Spam resistance** | ✅ high friction | ⚠️ medium | ❌ upload anything |

That last row matters more than usual for us. **With no human moderation, the compile gate is the
only gate we have.** A PDF-only lane has no gate at all.

### 1.4 Recommendation — a three-tier acceptance model

Don't make it a binary. Make the *tier* visible on the paper's page, and let it drive features:

**Tier 1 — `source` (LaTeX).** Required path, the default, and the only one available at launch.
Full feature set: HTML view, MathML, structured metadata, version diffs, JATS export, "Source
available" badge, higher default ranking in our own search.

**Tier 2 — `converted` (DOCX).** Ship right after launch. Pipeline: DOCX → pandoc → LaTeX → Tectonic
→ PDF + HTML. DOCX carries real structure (styles, headings, footnotes, equations as OMML), so
conversion is genuinely good — far better than anything you can do from a PDF. Keep the original
DOCX as an archived asset. This is what unlocks "all disciplines" honestly.

**Tier 3 — `unstructured` (PDF).** Accept only with these guardrails:
- **Reject if `/Producer` or `/Creator` matches a TeX engine** — arXiv's exact rule, with a message
  telling them to upload the source.
- **Reject if there is no extractable text layer** (scans). Google Scholar rejects these too, so
  accepting them would mean hosting papers that can never be indexed or cited.
- **Reject if >5 MB** (Scholar's hard cap — see §2.2). Or warn and offer compression.
- Metadata typed by hand by the submitter, flagged as unverified.
- No HTML view, no diffs, no JATS. Visibly a second-class object.

**My call:** launch Tier 1 only (it's what you already have and it's the strongest product), add
Tier 2 within the first few iterations because "all disciplines" demands it, and treat Tier 3 as a
grudging escape hatch that we may never need. If you want to keep it maximally simple, ship
**Tier 1 + Tier 3-with-guardrails** and add DOCX when the first biologist complains.

### 1.5 One hard constraint you should design around now

**Google Scholar will not index a PDF over 5 MB, and will not index a PDF without a text layer.**
Our compiler should surface the output PDF's size at compile time and warn above ~4 MB, with the
same advice arXiv gives: JPEG for photographs, PDF/PNG/PS for line art, and watch for images above
34 megapixels. A paper we host that Scholar can't index is a paper that will never accrue citations —
which makes the whole profile feature in §6 worthless for it.

---

## 2. Google Scholar — how content gets in

Scholar has **no public API and no submission form.** You do not submit to Scholar; Scholar crawls
you. Getting indexed is purely a matter of meeting their technical guidelines.

### 2.1 Site-level requirements

- The site must host **primarily scholarly articles** — journal papers, conference papers,
  dissertations, preprints, technical reports. (Scholar generally excludes book reviews, editorials,
  and **documents without authors**.)
- **Every article needs its own unique URL.** Not multiple abstracts on one page, not multiple papers
  in one PDF, not one paper split across several files.
- **Complete abstracts must be visible** without signing in, installing anything, dismissing a popup,
  or scrolling past other content.
- **Browse interface:** every article URL must be reachable from the homepage in **≤10 simple HTML
  links**. Recommended shapes: a single publications page (small collections); browse by publication
  date or entry date (thousands); plus a **"last two weeks" listing** for frequent recrawl once you
  pass ~100,000 papers.
- **No Flash, JavaScript, or form-based navigation** for those links — plain HTML `GET` links.
  (Relevant to us: SvelteKit must SSR the browse pages and the article pages.)
- **robots.txt** must not block article or browse URLs; *should* block search results, carts,
  comment forms and other dynamic junk.
- **Uptime and correct status codes:** `5xx` for temporary errors, `4xx` for permanent ones,
  **HTTP 301 when an article moves — never redirect to the homepage.**

### 2.2 File requirements

- **HTML or searchable PDF only. Maximum 5 MB.** (Over 5 MB → they tell you to use Google Books.)
- PDFs must have text extractable by Acrobat Reader. No scans.

### 2.3 Metadata — the Highwire Press tags

Scholar reads four schemes; **Highwire Press `citation_*` is preferred** (also supported: BE Press
`bepress_citation_*`, PRISM `prism.*`, Dublin Core `DC.*` — DC is explicitly the least reliable).

Minimum required: **title, first author, publication year.**

```html
<meta name="citation_title"            content="On the Electrodynamics of Moving Bodies">
<meta name="citation_author"           content="Einstein, Albert">   <!-- one tag per author -->
<meta name="citation_author"           content="Besso, Michele">
<meta name="citation_publication_date" content="1905/6/30">          <!-- YYYY/M/D -->
<meta name="citation_online_date"      content="2026/07/26">         <!-- repository entry date -->
<meta name="citation_journal_title"    content="ScholarXIV Preprints">
<meta name="citation_technical_report_institution" content="ScholarXIV">
<meta name="citation_pdf_url"          content="https://…/abs/2607.00123v1.pdf"> <!-- absolute -->
<meta name="citation_issn"             content="…">
<meta name="citation_doi"              content="10.xxxxx/…">
```

Also available: `citation_conference_title`, `citation_volume`, `citation_issue`,
`citation_firstpage`, `citation_lastpage`, `citation_isbn`,
`citation_dissertation_institution`, `citation_technical_report_number`.

Rules: author names as `Smith, John` or `John Smith`; **omit affiliations and degrees**; present
values as they'd appear in a citation (`Trans. Mag. Real.`, not `Magic Realism, Transactions on`);
escape HTML special characters properly.

### 2.4 If you have no meta tags, the PDF's *visual layout* becomes the metadata

Worth knowing because it constrains our LaTeX templates:

- **Title = the largest text on page 1, minimum 24pt, at the top.**
- **Authors = 16–23pt**, immediately before or after the title.
- **A formal bibliographic citation line alone on its own line** in the header or footer.
- **Use Type 1 fonts, not Type 3** — i.e. Times/Helvetica/Palatino in LaTeX. (Old `latex`+`dvips`
  bitmap-font setups produce Type 3 and break Scholar's text extraction. Tectonic/pdfTeX with
  standard font packages is fine; this is a template-lint rule we should enforce.)
- Journal/repository name must be in **smaller** type than the title and authors.

We will emit meta tags, so this is a belt-and-braces concern — but our default templates should
comply anyway, because the PDF travels beyond our site.

### 2.5 Reference formatting — this is how citations happen

Scholar **parses reference lists out of the full text.** Their requirements:

- Mark the section with a standard heading — **"References" or "Bibliography"** — on its own line.
- Format entries as **numbered**: `1.`/`2.`/`3.` or `[1]`/`[2]`/`[3]` in PDFs; an `<ol>` in HTML.
- Each entry must be a **formal bibliographic citation with no commentary mixed in.**
- *"Incorrect formatting may cause exclusion or low ranking."*

Practical consequence for us: our LaTeX templates should default to a numeric bibliography style
(`plainnat`/`unsrt`/`ieeetr`-like), and our HTML renderer must emit the bibliography as a real
`<ol>` under an `<h2>References</h2>`. Author-year styles are riskier for parsers.

### 2.6 Timeline, and the risk our no-moderation choice creates

- Scholar updates **several times a week**; new papers typically appear **within weeks**.
- **Corrections to already-indexed papers take 6–9 months, sometimes years.** Get the metadata right
  the first time — a bad `citation_author` is effectively permanent for the better part of a year.
- The `site:` operator's result count is explicitly unreliable for measuring coverage.

**The risk, stated once:** Scholar's first content requirement is that the site host *primarily
scholarly articles*. A free, all-disciplines, zero-moderation server is a spam magnet, and the
penalty for failing that test is de-indexing — which would destroy the citation and profile features
that are the whole point of §6. I'd still build it your way, but with an **automated** gate that
costs no staff time:

1. Compile must succeed (already have it).
2. Similarity check against the corpus + web (Crossref Similarity Check or an open equivalent).
3. Account verification + per-account rate limits (arXiv caps back-catalogue posting at 3/day).
4. Structural sanity: has an abstract, has a references section, above a minimum length, is in
   English or declares its language, isn't a duplicate of an existing entry.
5. LLM scope/type classifier that flags — not blocks — likely non-research uploads.

That is a machine, not a moderation team, and it keeps us inside Scholar's content requirement.

---

## 3. Google Scholar — how citations are actually counted

- Scholar crawls full text, **extracts each paper's reference list, and matches those references to
  records it has already indexed.** Citation counts are the size of the resulting inbound set.
- Their own caveat: counts *"reflect the state of the web as it is currently visible to our search
  robots."* They are an estimate, not a ledger.
- **Version clustering is the important mechanic.** Scholar automatically identifies different
  versions of the same article — preprint, publisher version, repository copies — and groups them
  into one record ("All versions"). **Citations aggregate across the cluster.** This is why an arXiv
  preprint and its published journal version share a citation count.

**Why this matters enormously for us:** if a ScholarXIV preprint is correctly recognised as a version
of the eventual journal article, our copy inherits the whole citation count and appears in "All
versions" — free credibility and traffic. To make that clustering work we need: identical/near
title and author strings, a DOI, and ideally an explicit `is-preprint-of` relationship registered
with the DOI (§5).

- An **asterisk (\*)** on a profile means the count includes citations that might not match that
  article — i.e. automated matching wasn't confident.

---

## 4. Citation data we can actually use in-product

**Google Scholar cannot be a data source for us.** No API, and scraping violates their terms. Every
"Scholar-like" product is built on the open stack below. Field shapes here were verified with live
API calls today.

### 4.1 OpenAlex — the best default

Free, CC0, ~2× the coverage of Scopus/WoS by their claim. A free API key now gives $1/day of usage;
heavier use needs a paid plan or the quarterly snapshot.

`GET https://api.openalex.org/authors?search=Yoshua%20Bengio` returns, per author:

```jsonc
{
  "id": "https://openalex.org/A5086198262",
  "display_name": "Yoshua Bengio",
  "raw_author_names": ["Bengio Y.", "Bengio, Yoshua", "Y. Bengio", "YOSHUA BENGIO", …],
  "works_count": 1290,
  "cited_by_count": 475760,
  "summary_stats": { "2yr_mean_citedness": 7.90, "h_index": 186, "i10_index": 723 },
  "affiliations": [
    { "institution": { "ror": "https://ror.org/01sdtdd95", "display_name": "CIFAR",
                       "country_code": "CA", "type": "facility" },
      "years": [2025, 2024, 2023, …] }
  ]
}
```

**`summary_stats` hands us h-index and i10-index precomputed** — the two headline numbers on a Google
Scholar profile — and `raw_author_names` is a ready-made name-variant list for disambiguation.

`GET https://api.openalex.org/works?filter=doi:…` returns, per work: `doi`, `title`,
`publication_date`, `type` (`"preprint"`), **`indexed_in: ["arxiv","datacite"]`**,
`primary_location { landing_page_url, pdf_url, version: "submittedVersion", source { type:
"repository", … } }`, `open_access { is_oa, oa_status: "green", oa_url }`, and `authorships[]` with
each author's OpenAlex id, institutions (with **ROR**), `is_corresponding`, and
`raw_affiliation_strings`.

Note the `primary_location.id` on that arXiv record: **`pmh:oai:arXiv.org:2205.01833`**. OpenAlex
ingested it **through OAI-PMH**. That is the doorway (§5.2).

### 4.2 Semantic Scholar Academic Graph (S2AG)

Free; API key optional but strongly recommended — unauthenticated calls get 429'd quickly (mine did).

- Paper: `citationCount`, **`influentialCitationCount`** (their distinctive signal), `referenceCount`,
  `externalIds` (ArXiv, DOI, DBLP, MAG, ACL, PubMed, CorpusId), `openAccessPdf`, `tldr`,
  `embedding.specter_v2`, plus `/citations` and `/references` endpoints.
- Author: `paperCount`, `citationCount`, `hIndex`. *(Verified live: Oren Etzioni → 258 / 43,671 / 88.)*
- Batch: `POST /paper/batch`, ≤500 ids, ≤10 MB, ≤9,999 citations per call. Accepts
  `ARXIV:2106.15928` style ids directly — useful for us.

### 4.3 Crossref

- Any DOI's inbound count is on the work record as **`is-referenced-by-count`** (verified:
  `10.1038/nature14539` → 74,490). Free, no key.
- **Cited-by** service is free but requires you to **deposit your reference lists**; Crossref warns
  that not depositing your own references under-reports counts **by at least 20%**. Crossref only
  matches Crossref-registered → Crossref-registered works.

### 4.4 OpenCitations

CC0 open citation index (COCI, OpenCitations Meta). Good as a third source and as the ideologically
consistent one for an open platform. REST APIs + full dumps.

### 4.5 Summary

| Source | Cost | Author metrics | Coverage | Use it for |
|---|---|---|---|---|
| **OpenAlex** | free (key, $1/day) | **h-index, i10-index, works, citations** | broadest | **primary** — profiles, metrics, affiliations, ROR |
| Semantic Scholar | free (key advised) | hIndex, paperCount, citationCount | strong in CS/bio | cross-check, influential citations, TLDRs, embeddings |
| Crossref | free | none | DOI-registered only | per-paper counts, reference deposit, our own DOIs |
| OpenCitations | free | none | DOI-based | open citation graph, third opinion |
| Google Scholar | ❌ no API | — | broadest | **link out only** |
| Scopus / WoS | paid | yes | curated | not worth it |

**Design rule:** merge sources by DOI → arXiv id → normalised title+year, and **always display which
source a number came from and when it was fetched.** Nothing destroys trust in an academic product
faster than an unattributed citation count that disagrees with Scholar.

---

## 5. Making *our* papers citable and indexable

Six things, in dependency order. All are cheap; none are optional if profiles are to mean anything.

### 5.1 Persistent identifiers

Mint a DOI per paper and per version. **Use Crossref, not DataCite** — see the pricing comparison in
[preprint-design-ingest-citations-profiles.md](preprint-design-ingest-citations-profiles.md) §5.1.
(arXiv uses DataCite — the OpenAlex record above shows `indexed_in: ["arxiv","datacite"]` — but that
only pays off above ~12,000 papers/year, and arXiv posts ~200k.) Register the **`is-preprint-of` /
`is-previous-version-of` relationships** so aggregators can cluster our copy with the published
version (§3).

Version policy to copy from bioRxiv: **all versions share one DOI**, with version-specific URLs
(`/abs/XXXX.NNNNNv2`) for citing a specific one.

### 5.2 An OAI-PMH endpoint

This is how the aggregators find us — it is the single highest-leverage item on this list.
arXiv's implementation is the template:

- OAI-PMH **v2.0**, one stable base URL.
- **Item = article**, latest version exposed (with version history inside the richer format).
- Metadata formats: **`oai_dc`** (Simple Dublin Core — the universal one) plus a native format with
  authors split out, categories, and license. arXiv offers `oai_dc`, `arXiv`, `arXivRaw`.
- **Sets** for selective harvesting, hierarchical: `group:archive:CATEGORY`.
- Datestamps = last modification, supporting incremental harvest via `from`.
- Update nightly, right after new papers go live.

Harvesters that will pick us up: OpenAlex, CORE, BASE, Semantic Scholar, OpenAIRE, Unpaywall.

### 5.3 Landing pages that satisfy §2

SSR'd HTML, unique URL per paper and per version, abstract visible with no interstitial, full
Highwire `citation_*` tags, absolute `citation_pdf_url`, and a browse interface within 10 clicks
(by date, by category, by author) plus a "recent" feed. Sitemap. Correct 301s forever.

### 5.4 Full text in both HTML and PDF

PDF ≤5 MB with a real text layer and Type 1 fonts; HTML for accessibility, mobile, and better
parsing. Tier-1 (LaTeX) submissions can produce both from one source — this is exactly why §1
recommends requiring source.

### 5.5 A conventionally formatted reference list

`<h2>References</h2>` + `<ol>` in HTML; numbered entries under a "References" heading in the PDF;
no commentary interleaved. See §2.5.

### 5.6 Deposit our references

If we register DOIs and deposit reference lists, we join the citation graph as a *citing* party, not
just a cited one — and Crossref Cited-by starts returning real data for our papers.

---

## 6. The researcher profile — design

### 6.1 What Google Scholar's profile actually is (the thing to match, then beat)

**Creation:** sign in with a Google account → four steps — confirm name spelling; enter affiliation
and research interests; **select your articles from clusters of similarly-named authors**; choose how
future updates are handled.

**Fields:** name, affiliation, research interests, photo, homepage URL, and a **verified
institutional email**. The email is *hidden from the public* but is **required for the profile to
appear in Scholar search results** — profile must be public **and** email-verified.

**Adding works, three ways:** search by title/keyword; enter bibliographic data manually if search
fails; or **"add article groups"** for people who publish under several name forms.

**Merging:** duplicates are consolidated with a **Merge** button that combines counts **without
double-counting** papers that cite both versions.

**Update mode:** *automatic* (matching articles are added without approval) or *notify me* (email
review first). This applies only to the article list — **citation counts always update automatically.**

**Metrics shown**, each in an **All** and a **Recent (last 5 years)** column:
- **Citations** — total.
- **h-index** — "the largest number *h* such that at least *h* articles were cited at least *h* times
  each." (Example from their own docs: citations of 17, 9, 6, 3, 2 → h = 3.)
- **i10-index** — number of articles with **≥10 citations**. Scholar-specific; nobody else uses it.
- An **asterisk** marks counts that include possibly-mismatched citations.

**Also:** public/private toggle; a stable public URL `scholar.google.com/citations?user=<ID>`;
co-author list; a per-article citation-over-time graph; "follow" alerts for new articles by an author
or new citations to a paper.

*(For venues rather than people, Scholar Metrics uses **h5-index** and **h5-median** — the h-index and
h-core median restricted to the last five complete calendar years. If we ever want to publish a
"ScholarXIV h5-index", that's the definition.)*

### 6.2 Proposed data model

```
ResearcherProfile
  userId, handle (public slug), visibility: private | public
  displayName, nameVariants[]            ← seed from OpenAlex raw_author_names
  photoUrl, pronouns?, bio, homepageUrl, socials[]
  affiliations[] { institutionName, rorId, department, role, startYear, endYear, isCurrent }
  verifiedEmails[] { address, domain, verifiedAt, isInstitutional, public: false }
  interests[] / topics[]                 ← controlled vocabulary + free text
  identifiers { openAlexId, semanticScholarId, arxivAuthorId, dblp, googleScholarUrl }
  metricsCache { source, fetchedAt, citations, hIndex, i10Index,
                 citationsRecent5y, hIndexRecent5y, i10IndexRecent5y,
                 worksCount, perYear: [{year, count}] }
  settings { autoAddWorks: bool, publicMetrics: bool, showEmail: false }

ProfileWork
  profileId, workId
  origin: native | claimed_external
  claimState: auto_matched | confirmed | rejected
  identifiers { doi, arxivId, openAlexId, s2PaperId }
  authorPosition, isCorresponding
  citationCounts[] { source, count, fetchedAt }
  mergedInto?  ← duplicate consolidation, Scholar-style
```

### 6.3 Claiming and disambiguation (the genuinely hard part)

Author name disambiguation is the unsolved problem of this domain. Ladder, strongest first:

1. **Native works.** Anything submitted through ScholarXIV by that account is auto-attached.
2. **Email match** against the corresponding-author email in the metadata.
3. **Candidate matching** — OpenAlex author entities whose `raw_author_names` overlap plus
   affiliation/topic overlap — presented as *suggestions the user confirms*, never auto-attached.
   This is exactly Scholar's step 3 (pick your articles from clusters) and it works because the human
   does the last mile.
4. **Manual add** by DOI / arXiv ID / free-form citation.
5. **Merge / split** controls, with a rule that merging never double-counts a citing paper.

Store `claimState` so an unconfirmed auto-match can never inflate public metrics.

### 6.5 Metrics: compute, cache, attribute

- Pull `summary_stats` from OpenAlex as the headline; cross-check totals against Semantic Scholar and
  Crossref; store all three with `fetchedAt`.
- Compute **h-index** and **i10-index** ourselves too, from the merged deduplicated work list, so a
  claimed-works profile stays correct when a user adds a paper OpenAlex hasn't linked to them.
- Offer the **All / Recent(5y)** split Scholar users expect.
- Show a **per-year citation histogram** and a **per-year publication count**.
- **Always label the source and the as-of date.** "Citations 1,204 · via OpenAlex, 2 hours ago."
- Refresh on a schedule (daily for active profiles, weekly otherwise), never on page load.

### 6.6 Where we can beat Scholar

Scholar's profile is a 2011 web page. Cheap differentiators, roughly in order of value/effort:

1. **Platform-native metrics Scholar can't have** — views, downloads, and version history for papers
   hosted here (arXiv-style, and honest: label them as engagement, not impact).
2. **Co-author graph** — an actual visualisation; OpenAlex `authorships` gives us the edges free.
3. **Topic map** over the author's corpus (we already do classification in this product).
4. **"Cited by" reading list** with our own summarisation on top — this is the natural bridge to the
   rest of ScholarXIV.
5. **One-click claim at onboarding**: paste an arXiv author ID, get a populated profile in
   seconds. First-run experience decides whether anyone ever comes back.
7. **Public API + machine-readable profile** (JSON-LD `ScholarlyArticle` / `Person`), because Scholar
   has none and researchers want their data out.
8. **DORA-consistent framing** — show distributions, not just a single h-index. It costs nothing and
   signals we understand the field's own critique of these metrics.

---

## 7. Build order for this half

1. **Landing pages + Highwire meta tags + sitemap + SSR browse** — nothing else works without these.
2. **PDF hygiene in the compile pipeline** — <5 MB warning, text layer assertion, Type 1 font lint.
3. **DOIs via DataCite**, with version + `is-preprint-of` relationships.
4. **OAI-PMH endpoint** (`oai_dc` + a native format, hierarchical sets, nightly).
5. **Profile v1** — identity, affiliations, and native works.
6. **Metrics v1** — OpenAlex ingest, cached, attributed, with All/Recent split.
7. **Claim & merge UI** — candidate matching, confirm/reject, merge.
8. **Differentiators** — co-author graph, topic map, cited-by reading list, public profile API.

---

## Sources

Google Scholar: [technical inclusion guidelines](https://scholar.google.com/intl/en/scholar/inclusion.html) ·
[Scholar Citations / profiles](https://scholar.google.com/intl/en/scholar/citations.html) ·
[how Scholar works](https://scholar.google.com/intl/en/scholar/help.html) ·
[Scholar Metrics](https://scholar.google.com/intl/en/scholar/metrics.html) —
[OpenAlex API docs](https://docs.openalex.org/) (author/work shapes verified live against
`api.openalex.org`) —
[Semantic Scholar Academic Graph API](https://api.semanticscholar.org/api-docs/graph) (verified live) —
[Crossref Cited-by](https://www.crossref.org/documentation/cited-by/) (verified live via
`api.crossref.org`) —
[OpenCitations](https://opencitations.net/) —
[arXiv OAI-PMH interface](https://info.arxiv.org/help/oa/index.html) —
[arXiv submission formats](https://info.arxiv.org/help/submit/index.html) —

## Sitemap

See the full [sitemap](/sitemap.md) for all pages.
Docs-scoped sitemap: [/docs/sitemap.md](/docs/sitemap.md).
Well-known sitemap: [/.well-known/sitemap.md](/.well-known/sitemap.md).
