citation-management
lingzhi227/agent-research-skills
Manage BibTeX citations for LaTeX papers with Semantic Scholar integration and validation.
What is citation-management?
Automates the full lifecycle of citations in LaTeX documents: harvesting missing citations from drafts via Semantic Scholar, validating cite keys against .bib files, deduplicating entries, and formatting bibliography. Use when writing academic papers and need to manage references efficiently.
- Harvest missing citations automatically by scanning .tex files and searching Semantic Scholar
- Validate all cite keys to detect missing citations, unused entries, duplicates, and undefined references
- Add specific papers by searching Semantic Scholar and extracting clean BibTeX entries
- Format and standardize .bib files with consistent indentation, sorted keys, and protected proper nouns
- Auto-fix missing citation placeholders by generating entries for all unresolved keys
- Check figures and cross-references for completeness
How to install citation-management
npx skills add https://github.com/lingzhi227/agent-research-skills --skill citation-management- Python installed
- Access to Semantic Scholar API (optional; required for harvest and add actions)
- LaTeX project with .tex and .bib files
How to use citation-management
- 1.Choose an action: harvest (auto-find missing citations), validate (check all citations), add (search for specific paper), or format (standardize .bib)
- 2.For harvest: run the harvest script pointing to your .tex and .bib files; it will iteratively search and add citations up to 20 rounds
- 3.For validate: run the validation script to detect missing citations, unused entries, duplicates, and undefined references
- 4.For add: search Semantic Scholar for a specific paper, extract the BibTeX, and append to your .bib file with a clean key
- 5.For format: sort entries alphabetically, standardize indentation, remove empty fields, and protect proper nouns in titles
Use cases
- Automatically find and add citations for uncited claims in a research paper draft
- Validate a complete LaTeX project before compilation to catch citation errors
- Search for and add a specific paper to your bibliography with proper BibTeX formatting
- Deduplicate and standardize an existing .bib file across multiple papers
- Generate a bibliography from a structured paper database in JSONL format
- Academic researchers writing LaTeX papers
- PhD students managing large bibliographies
- Authors preparing papers for submission
- Anyone maintaining BibTeX reference files
citation-management FAQ
It reads your .tex draft, identifies important missing citations, searches Semantic Scholar for each one, selects the most relevant result, extracts BibTeX, generates a clean key, and appends to your .bib file. It repeats up to 20 rounds until no more gaps are found.
Use the format firstAuthorLastNameYearFirstContentWord, e.g., vaswani2017attention. The skill auto-generates keys in this format when harvesting or adding papers.
Yes, use the --dry-run flag with the harvest script to preview candidate citations without modifying your .bib file.
It reports missing citations, unused bib entries, duplicate keys, duplicate sections, duplicate labels, undefined references, and missing figures.
An API key is optional but recommended for harvest and add actions. Without it, you can still validate, format, and manually manage citations.
Full instructions (SKILL.md)
Source of truth, from lingzhi227/agent-research-skills.
name: citation-management description: Manage BibTeX citations for LaTeX papers. Harvest missing citations from a draft using Semantic Scholar, validate cite keys against .bib files, deduplicate entries, and format bibliography. Use when working with references, BibTeX, or citations. argument-hint: [tex-or-bib-file]
Citation Management
Manage the full lifecycle of citations in a LaTeX paper.
Input
$0— Action:harvest,validate,add,format$1— Path to.texor.bibfile
Scripts
Validate citations (check all cite keys resolve)
python ~/.claude/skills/citation-management/scripts/validate_citations.py \
--tex paper/main.tex --bib paper/references.bib --check-figures --figures-dir paper/figures/
Reports: missing citations, unused bib entries, duplicate keys, duplicate sections, duplicate labels, undefined references, missing figures.
Generate BibTeX from paper database
python ~/.claude/skills/deep-research/scripts/bibtex_manager.py \
--jsonl paper_db.jsonl --output references.bib
Search for a specific paper to add
python ~/.claude/skills/deep-research/scripts/search_semantic_scholar.py \
--query "attention is all you need" --max-results 5 \
--api-key "$(grep S2_API_Key /Users/lingzhi/Code/keys.md 2>/dev/null | cut -d: -f2 | tr -d ' ')"
Harvest missing citations automatically
python ~/.claude/skills/citation-management/scripts/harvest_citations.py \
--tex paper/main.tex --bib paper/references.bib --output candidates.bib --max-rounds 10
Scans .tex for uncited claims, searches Semantic Scholar, outputs candidate BibTeX entries.
Key flags: --dry-run (preview only), --verbose, --api-key
Auto-fix missing citation placeholders
python ~/.claude/skills/citation-management/scripts/validate_citations.py \
--tex paper/main.tex --bib paper/references.bib --fix
Generates references_fixed.bib with placeholder entries for all missing citation keys.
Action: harvest — Iterative Citation Harvesting
Based on AI-Scientist's 20-round citation harvesting loop. For each round:
- Read the current
.texdraft - Identify the most important missing citation
- Search Semantic Scholar via script
- Select the most relevant paper from results
- Extract BibTeX and generate a clean key (
lastNameYearWord) - Append to
.bib(skip if key exists) - Insert
\cite{key}at the appropriate location - Stop when no more gaps or 20 rounds reached
Key rules:
- DO NOT add a citation that already exists
- Only add citations found via API — never fabricate
- Cite broadly — not just popular papers
- Do not copy verbatim from prior literature
Action: validate — Pre-Compilation Check
Run validate_citations.py to catch all issues before compilation. Fix any reported problems.
Action: add — Add Specific Paper
Search Semantic Scholar for the paper, extract BibTeX, clean the key, append to .bib.
BibTeX key format: firstAuthorLastNameYearFirstContentWord (e.g., vaswani2017attention)
Action: format — Standardize .bib
- Sort entries alphabetically by key
- Ensure consistent indentation (2 spaces)
- Remove empty fields
- Protect proper nouns with
{Braces}in titles - Ensure required fields per entry type
Related Skills
- Upstream: literature-search, deep-research
- Downstream: paper-compilation, latex-formatting
- See also: related-work-writing
Related skills
More from lingzhi227/agent-research-skills and the wider catalog.

code-debugging
Debug experiment code with structured error analysis and retry logic.

data-analysis
Generate rigorous statistical analysis code with multi-round review and proper uncertainty reporting.

deep-research
Conduct systematic 6-phase academic literature reviews with structured output and synthesis.

excalidraw-skill
Programmatic canvas toolkit for creating and refining Excalidraw diagrams with real-time sync and element-level control.

experiment-code
Generate and iteratively improve ML experiment code for research papers with training pipelines, debugging, and result optimization.

experiment-design
Design structured, progressive experiment plans for research papers with staged baselines, hyperparameter sweeps, and ablation studies.