nature-reader
yuan1z0825/nature-skills
Create bilingual Chinese-English paper readers with aligned text, figures, tables, and equations from PDFs or HTML.
What is nature-reader?
A skill for building source-grounded readers that present papers in parallel Chinese-English translation with preserved layout, figures, and tables. Use it for full-paper translation, bilingual side-by-side reading, detailed paper study, or answering source-linked questions without regenerating the entire reader.
- Extracts and aligns text, figures, tables, and equations from PDFs (selectable or scanned), HTML, or pasted content
- Generates bilingual Chinese-English readers with parallel translation for meaning, not word-for-word
- Creates source maps with glossaries and recurring-term tables for consistent terminology
- Handles multiple source formats (PDF, HTML, DOI, arXiv links) with format-specific extraction rules
- Answers source-linked questions or translates specific excerpts without requiring full reader regeneration
- Preserves original document layout and structure in the output
How to install nature-reader
npx skills add https://github.com/yuan1z0825/nature-skills --skill nature-reader- A source document in one of the supported formats: PDF (text or scanned), HTML, DOI, arXiv link, or pasted text
How to use nature-reader
- 1.Provide the paper source (PDF file, HTML link, DOI, arXiv URL, or pasted text)
- 2.Specify your task: full-paper reader, excerpt translation, or a source-linked question
- 3.The skill will detect the source format and confirm it with you before processing
- 4.For a full reader, it extracts text, figures, tables, and equations, then generates aligned bilingual output
- 5.Review the generated reader files (paper.md, source_map.json, translation_notes.md) and ask follow-up questions as needed
Use cases
- Translate a full academic paper into bilingual format for research teams working across languages
- Study a specific section of a paper in detail with aligned Chinese and English text
- Extract and translate figures, tables, and equations from a scanned or image-heavy PDF
- Answer questions about a paper's content with citations grounded in the source material
- Build a terminology glossary for recurring terms across a multi-page research document
- Researchers and academics reading papers across Chinese and English
- Translation teams working with technical or scientific documents
- Students conducting detailed paper reviews or literature surveys
- Teams needing consistent terminology mapping across bilingual documents
nature-reader FAQ
You can ask about a specific excerpt or section without requiring a full reader. The skill will extract and ground its answer in just that part of the source.
The skill detects scanned PDFs and applies OCR-aware extraction rules. It will note any low-confidence crops in the translation_notes.md file.
It produces a full translation for meaning, never degrading to a summary. The bilingual reader preserves the original structure and content.
Yes. The skill can resolve DOI and arXiv links directly and extract the underlying PDF or HTML automatically.
It builds a terminology ledger during translation to ensure consistent terminology throughout, and applies specialized handling for equations, chemical formulae, and mathematical expressions.
Full instructions (SKILL.md)
Source of truth, from yuan1z0825/nature-skills.
name: nature-reader description: "Create source-grounded Chinese-English paper readers with aligned text, figures, tables, and equations. Use for 全文翻译、中英文对照、论文精读 or source-linked questions about a paper; respect a requested excerpt or question without generating a full reader." metadata: version: "2.1.1" author: Community contribution, refactored into static/dynamic layers
Full-Paper Markdown Reader — Router
Routing protocol
First distinguish creating a reader from answering a question or translating an excerpt.
For a source-linked question, read references/grounding-rules.md and inspect only the relevant
source material; reuse existing source-map IDs when available. Do not regenerate the reader or
require a full source map before answering. For an explicit excerpt request, apply extraction,
translation, and grounding rules to that excerpt. The full-artifact workflow below applies when
the user requests a reader or full-paper translation.
For a new task, load the core and matching resources below. Reuse already loaded guidance on follow-ups; load more only when the task needs it.
1. Load the manifest and the core layer
Read manifest.yaml. It declares the source_format axis, the allowed values, and the file paths each value maps to.
Also read every file listed under always_load. These hold the core principles, the reading workflow, and the output contract that apply to every reading job, plus the shared Terminology Ledger used to build the recurring-term table.
2. Detect the source format
Decide the source_format value using the manifest's detect: hint and the user's input:
pdf-text— selectable-text PDF. Default.scanned-pdf— image-only or OCR-required PDF.html— publisher or preprint HTML page.doi-arxiv— a bare DOI or arXiv link that must be resolved first.pasted-text— pasted prose or notes with no retrievable original layout.
State the detected value in one short line to the user before processing, so they can correct you cheaply. A source may map to more than one value (for example a DOI that resolves to a PDF); load the resolution fragment first, then the fragment for the resolved artifact.
3. Load the matching fragment(s)
Read the file mapped for the detected source_format. Do not read every fragment in static/. Load only what step 2 selected.
4. Build the reader using the loaded material
Apply the loaded fragments in this priority order:
- Core principles (
core/principles.md) — bilingual reader by default, translate for meaning, never degrade to a summary, copyright caution. - Source-format fragment — how to extract text, figures, and tables for this input.
- Reading workflow (
core/workflow.md) — the six-step source-map-first process. - Output contract (
core/output-contract.md) — required files and the pre-response verification checklist.
Build the Terminology Ledger as you translate (../nature-shared/core/terminology-ledger.md); it becomes the paper.md recurring-term table and the source_map.json glossary.
If constraints prevent full processing, still create a draft reader and label missing pages, figures, or low-confidence crops in translation_notes.md. Do not switch to summary mode.
5. Reach for references only when needed
The files under references/ are deep references, not defaults. Open them on demand per the references.on_demand table in the manifest:
- detailed figure/table cropping and placement →
references/figure-extraction.md. - exact field schema for
paper.md/source_map.json→references/output-spec.md. - equations, mathematical expressions, chemical formulae, or image-only formulae →
references/equation-handling.md. - answering follow-up questions with source citations →
references/grounding-rules.md.
Related skills
More from yuan1z0825/nature-skills and the wider catalog.

nature-ref-verifier
Multi-source cross-verification of academic references with field-level comparison and structured error reporting.

nature-response
Draft, audit, and revise peer-review responses, rebuttals, and revision cover letters.

nature-reviewer
Simulate Nature-style peer review of manuscripts with evidence-grounded assessments from three independent reviewers.

nature-shared
Internal reference library for Nature Skills—load only as a dependency, not standalone.

nature-statistics
Audit and improve manuscript statistical reporting for transparency, reproducibility, and design clarity.

nature-writing
Draft and restructure scientific manuscripts and initial-submission materials using Nature-style guidance.