PluginBench
Skill
Review
Audit score 70

nature-reader

yuan1z0825/nature-skills

Create bilingual Chinese-English paper readers with aligned text, figures, tables, and equations from PDFs or HTML.

What is nature-reader?

A skill for building source-grounded readers that present papers in parallel Chinese-English translation with preserved layout, figures, and tables. Use it for full-paper translation, bilingual side-by-side reading, detailed paper study, or answering source-linked questions without regenerating the entire reader.

  • Extracts and aligns text, figures, tables, and equations from PDFs (selectable or scanned), HTML, or pasted content
  • Generates bilingual Chinese-English readers with parallel translation for meaning, not word-for-word
  • Creates source maps with glossaries and recurring-term tables for consistent terminology
  • Handles multiple source formats (PDF, HTML, DOI, arXiv links) with format-specific extraction rules
  • Answers source-linked questions or translates specific excerpts without requiring full reader regeneration
  • Preserves original document layout and structure in the output

How to install nature-reader

npx skills add https://github.com/yuan1z0825/nature-skills --skill nature-reader
Prerequisites
  • A source document in one of the supported formats: PDF (text or scanned), HTML, DOI, arXiv link, or pasted text
Claude Code
Cursor
Windsurf
Cline

How to use nature-reader

  1. 1.Provide the paper source (PDF file, HTML link, DOI, arXiv URL, or pasted text)
  2. 2.Specify your task: full-paper reader, excerpt translation, or a source-linked question
  3. 3.The skill will detect the source format and confirm it with you before processing
  4. 4.For a full reader, it extracts text, figures, tables, and equations, then generates aligned bilingual output
  5. 5.Review the generated reader files (paper.md, source_map.json, translation_notes.md) and ask follow-up questions as needed

Use cases

Good for
  • Translate a full academic paper into bilingual format for research teams working across languages
  • Study a specific section of a paper in detail with aligned Chinese and English text
  • Extract and translate figures, tables, and equations from a scanned or image-heavy PDF
  • Answer questions about a paper's content with citations grounded in the source material
  • Build a terminology glossary for recurring terms across a multi-page research document
Who it's for
  • Researchers and academics reading papers across Chinese and English
  • Translation teams working with technical or scientific documents
  • Students conducting detailed paper reviews or literature surveys
  • Teams needing consistent terminology mapping across bilingual documents

nature-reader FAQ

Do I need to provide the entire paper, or can I ask about a specific section?

You can ask about a specific excerpt or section without requiring a full reader. The skill will extract and ground its answer in just that part of the source.

What if my PDF is scanned or image-only?

The skill detects scanned PDFs and applies OCR-aware extraction rules. It will note any low-confidence crops in the translation_notes.md file.

Will this be a summary or a full translation?

It produces a full translation for meaning, never degrading to a summary. The bilingual reader preserves the original structure and content.

Can I use this with papers from arXiv or DOI links?

Yes. The skill can resolve DOI and arXiv links directly and extract the underlying PDF or HTML automatically.

How does it handle technical terms and equations?

It builds a terminology ledger during translation to ensure consistent terminology throughout, and applies specialized handling for equations, chemical formulae, and mathematical expressions.

Full instructions (SKILL.md)

Source of truth, from yuan1z0825/nature-skills.


name: nature-reader description: "Create source-grounded Chinese-English paper readers with aligned text, figures, tables, and equations. Use for 全文翻译、中英文对照、论文精读 or source-linked questions about a paper; respect a requested excerpt or question without generating a full reader." metadata: version: "2.1.1" author: Community contribution, refactored into static/dynamic layers

Full-Paper Markdown Reader — Router

Routing protocol

First distinguish creating a reader from answering a question or translating an excerpt. For a source-linked question, read references/grounding-rules.md and inspect only the relevant source material; reuse existing source-map IDs when available. Do not regenerate the reader or require a full source map before answering. For an explicit excerpt request, apply extraction, translation, and grounding rules to that excerpt. The full-artifact workflow below applies when the user requests a reader or full-paper translation.

For a new task, load the core and matching resources below. Reuse already loaded guidance on follow-ups; load more only when the task needs it.

1. Load the manifest and the core layer

Read manifest.yaml. It declares the source_format axis, the allowed values, and the file paths each value maps to.

Also read every file listed under always_load. These hold the core principles, the reading workflow, and the output contract that apply to every reading job, plus the shared Terminology Ledger used to build the recurring-term table.

2. Detect the source format

Decide the source_format value using the manifest's detect: hint and the user's input:

  • pdf-text — selectable-text PDF. Default.
  • scanned-pdf — image-only or OCR-required PDF.
  • html — publisher or preprint HTML page.
  • doi-arxiv — a bare DOI or arXiv link that must be resolved first.
  • pasted-text — pasted prose or notes with no retrievable original layout.

State the detected value in one short line to the user before processing, so they can correct you cheaply. A source may map to more than one value (for example a DOI that resolves to a PDF); load the resolution fragment first, then the fragment for the resolved artifact.

3. Load the matching fragment(s)

Read the file mapped for the detected source_format. Do not read every fragment in static/. Load only what step 2 selected.

4. Build the reader using the loaded material

Apply the loaded fragments in this priority order:

  1. Core principles (core/principles.md) — bilingual reader by default, translate for meaning, never degrade to a summary, copyright caution.
  2. Source-format fragment — how to extract text, figures, and tables for this input.
  3. Reading workflow (core/workflow.md) — the six-step source-map-first process.
  4. Output contract (core/output-contract.md) — required files and the pre-response verification checklist.

Build the Terminology Ledger as you translate (../nature-shared/core/terminology-ledger.md); it becomes the paper.md recurring-term table and the source_map.json glossary.

If constraints prevent full processing, still create a draft reader and label missing pages, figures, or low-confidence crops in translation_notes.md. Do not switch to summary mode.

5. Reach for references only when needed

The files under references/ are deep references, not defaults. Open them on demand per the references.on_demand table in the manifest:

  • detailed figure/table cropping and placement → references/figure-extraction.md.
  • exact field schema for paper.md / source_map.json → references/output-spec.md.
  • equations, mathematical expressions, chemical formulae, or image-only formulae → references/equation-handling.md.
  • answering follow-up questions with source citations → references/grounding-rules.md.