PluginBench
Skill
Review
Audit score 70

nature-reader

yuan1z0825/nature-skills

Full-paper Chinese-English bilingual Markdown reader for PDFs, DOI, arXiv, and HTML papers with figures and tables preserved.

What is nature-reader?

Converts academic papers from multiple sources (PDF, DOI, arXiv, HTML, pasted text) into side-by-side Chinese-English Markdown with figures and tables positioned near relevant text. Use this skill whenever you need to read, translate, or deeply understand a research paper while maintaining source grounding and document structure.

  • Extracts and converts papers from PDF, DOI, arXiv, publisher HTML, or pasted text into structured Markdown
  • Creates bilingual Chinese-English side-by-side translations focused on meaning, not word-for-word
  • Preserves figure and table placement near relevant prose sections
  • Maintains exact source anchors and location maps for every content block
  • Builds a recurring-term Terminology Ledger for consistent translation of domain-specific terms
  • Generates source maps and grounding references for follow-up questions

How to install nature-reader

npx skills add https://github.com/yuan1z0825/nature-skills --skill nature-reader
Prerequisites
  • Source material: PDF file, DOI, arXiv link, publisher HTML URL, or pasted paper text
Claude Code
Cursor
Windsurf
Cline

How to use nature-reader

  1. 1.Provide the paper source (PDF, DOI link, arXiv URL, HTML page, or pasted text)
  2. 2.The skill will detect the source format and confirm it with you
  3. 3.The skill loads the appropriate extraction fragment for your source type
  4. 4.The skill processes the paper following the six-step workflow: source mapping, text extraction, figure/table extraction, bilingual translation, terminology ledging, and output generation
  5. 5.Review the generated `paper.md` (bilingual reader), `source_map.json` (location anchors), and `translation_notes.md` (any processing notes)

Use cases

Good for
  • Reading a research paper with full Chinese-English translation while preserving layout and figures
  • Extracting and understanding specific figures or tables from an academic paper
  • Translating conference papers or journal articles for non-English readers
  • Creating a structured, searchable reference document from a PDF or arXiv preprint
  • Building a terminology ledger for consistent translation of technical terms across a paper
Who it's for
  • Researchers and academics reading papers in non-native languages
  • Students studying papers that require deep understanding of both English and Chinese content
  • Teams collaborating across language barriers on literature review
  • Anyone needing to extract and organize figures, tables, and text from academic PDFs

nature-reader FAQ

Will this skill summarize the paper instead of providing full translation?

No. This skill produces full-paper bilingual readers by default. It will not degrade to summary-only output unless you explicitly request a summary.

What source formats are supported?

PDF (selectable text or scanned/OCR), publisher HTML, DOI links, arXiv links, and pasted text. The skill auto-detects the format and confirms it with you before processing.

How are figures and tables handled?

Figures and tables are extracted and positioned near the relevant prose sections in the output Markdown, preserving the document's logical flow and structure.

What output files will I receive?

You'll receive `paper.md` (the bilingual reader), `source_map.json` (source anchors and glossary), and `translation_notes.md` (processing notes or missing sections if any).

Can I use this for papers in languages other than English and Chinese?

This skill is optimized for English-to-Chinese bilingual reading. For other language pairs, results may vary.

Full instructions (SKILL.md)

Source of truth, from yuan1z0825/nature-skills.


name: nature-reader description: Build full-paper Chinese-English side-by-side, figure/table-aware, source-grounded Markdown readers for journal or conference papers from PDF, DOI, arXiv, publisher HTML, or pasted text. Use whenever the user asks to translate or read a paper, make 中英文对照/原文对照/全文翻译解读, extract figures or tables into the right positions, preserve figure/table placement near relevant prose, or keep exact source anchors for every block. This skill must not degrade into a summary-only output unless the user explicitly asks for a summary. Also trigger on general paper-reading and translation requests even without the word "Nature", such as reading/translating an academic paper, literature reading, understanding a paper, and Chinese phrasings like 读论文、精读论文、论文翻译、文献翻译、文献阅读、学术阅读、帮我读这篇文章、翻译这篇paper. version: 2.0.0 author: Community contribution, refactored into static/dynamic layers

Full-Paper Markdown Reader — Router

This skill is split into two layers:

  • A static layer under static/ that holds versioned, reusable content fragments (core principles, the reading workflow, the output contract, and per-source-format extraction guidance).
  • A dynamic layer (this file plus manifest.yaml) that detects the request's source format and loads only the fragments needed for the current job.

Do not try to apply the reading logic from memory or from this router. Always load fragments from disk as described below.

Routing protocol

Follow these five steps every time the skill is invoked.

1. Load the manifest and the core layer

Read manifest.yaml. It declares the source_format axis, the allowed values, and the file paths each value maps to.

Also read every file listed under always_load. These hold the core principles, the reading workflow, and the output contract that apply to every reading job, plus the shared Terminology Ledger used to build the recurring-term table.

2. Detect the source format

Decide the source_format value using the manifest's detect: hint and the user's input:

  • pdf-text — selectable-text PDF. Default.
  • scanned-pdf — image-only or OCR-required PDF.
  • html — publisher or preprint HTML page.
  • doi-arxiv — a bare DOI or arXiv link that must be resolved first.
  • pasted-text — pasted prose or notes with no retrievable original layout.

State the detected value in one short line to the user before processing, so they can correct you cheaply. A source may map to more than one value (for example a DOI that resolves to a PDF); load the resolution fragment first, then the fragment for the resolved artifact.

3. Load the matching fragment(s)

Read the file mapped for the detected source_format. Do not read every fragment in static/. Load only what step 2 selected.

4. Build the reader using the loaded material

Apply the loaded fragments in this priority order:

  1. Core principles (core/principles.md) — bilingual reader by default, translate for meaning, never degrade to a summary, copyright caution.
  2. Source-format fragment — how to extract text, figures, and tables for this input.
  3. Reading workflow (core/workflow.md) — the six-step source-map-first process.
  4. Output contract (core/output-contract.md) — required files and the pre-response verification checklist.

Build the Terminology Ledger as you translate (../_shared/core/terminology-ledger.md); it becomes the paper.md recurring-term table and the source_map.json glossary.

If constraints prevent full processing, still create a draft reader and label missing pages, figures, or low-confidence crops in translation_notes.md. Do not switch to summary mode.

5. Reach for references only when needed

The files under references/ are deep references, not defaults. Open them on demand per the references.on_demand table in the manifest:

  • detailed figure/table cropping and placement → references/figure-extraction.md.
  • exact field schema for paper.md / source_map.jsonreferences/output-spec.md.
  • answering follow-up questions with source citations → references/grounding-rules.md.

Why this split

  • The static layer is versioned and reviewable. Adding a new source format is one new fragment plus one manifest line.
  • The dynamic layer keeps each invocation cheap: only the fragment relevant to this input enters context.
  • The router itself is short on purpose. Update fragments, not this file, when adding scope.
  • This structure mirrors nature-writing and nature-polishing so shared content lives in _shared/.