PluginBench
Skill
Pass
Audit score 90

gget

affaan-m/everything-claude-code

Quick genomic database queries, BLAST searches, and reproducible bioinformatics evidence logs via gget CLI.

What is gget?

gget is a unified CLI and Python interface for querying genomic reference databases including Ensembl, UniProt, and specialized modules for BLAST, protein structures, pathways, and disease associations. Use it for rapid first-pass bioinformatics lookups before moving to production pipelines.

  • Search and retrieve Ensembl IDs, gene metadata, and transcript details
  • Fetch nucleotide and amino-acid sequences from reference databases
  • Run BLAST and BLAT sequence similarity searches without local setup
  • Query protein structures, pathways, expression, cancer, and disease-association data
  • Generate reproducible evidence logs with explicit version, arguments, and database metadata
  • Support multiple output formats including JSON, CSV, and FASTA

How to install gget

npx skills add https://github.com/affaan-m/everything-claude-code --skill gget
Prerequisites
  • Python 3.7 or later
  • Virtual environment (venv or uv recommended)
  • pip or uv package manager
Claude Code
Cursor
Windsurf
Cline

How to use gget

  1. 1.Create and activate a clean Python virtual environment
  2. 2.Install gget using pip or uv: `pip install --upgrade gget`
  3. 3.Verify installation with `gget --help`
  4. 4.Identify the species, assembly, gene ID type, and database module needed
  5. 5.Run a small test query first (e.g., `gget search -s human BRCA1`)
  6. 6.Save output with explicit filename and date stamp
  7. 7.Record module name, version, arguments, and database assumptions in a reproducibility log
  8. 8.Consult upstream documentation for module-specific arguments before each query

Use cases

Good for
  • Finding gene IDs and metadata for a list of gene symbols before downstream analysis
  • Running a quick BLAST search on a protein sequence to identify homologs
  • Fetching reference genome download links and annotations for a specific assembly
  • Querying enrichment or disease-association databases to contextualize genomic findings
  • Creating a timestamped, reproducible log of database queries for scientific documentation
Who it's for
  • Bioinformaticians performing exploratory genomic analysis
  • Researchers needing quick reference lookups without building full pipelines
  • Scientists documenting evidence trails for publications or clinical contexts
  • Developers integrating genomic queries into Python workflows

gget FAQ

When should I use gget instead of a dedicated tool like BLAST+ or Biopython?

Use gget for rapid exploratory lookups and first-pass evidence gathering. Switch to dedicated tools when you need production-scale pipelines, fine-grained control over database versions, local indexes, or regulated clinical interpretation.

How do I ensure my gget queries are reproducible?

Record the gget version, module name, exact query arguments, species/assembly, output filename, and date in a metadata table. Save the command or Python snippet used. Upgrade gget regularly and re-check upstream docs before running queries.

What output formats does gget support?

gget outputs JSON, CSV, FASTA, and pandas DataFrames depending on the module. Specify output format with the `-o` flag and filename extension (e.g., `-o results.json`).

Do all gget modules work with all Python versions?

Most modules support Python 3.7+, but some optional scientific dependencies have narrower version constraints. Test in an isolated environment and check upstream docs for version-specific requirements.

Can I use gget for clinical or regulated genomic interpretation?

No. gget is designed for research and exploratory analysis. For clinical interpretation, use validated, regulated workflows and specialized clinical genomics platforms.

Full instructions (SKILL.md)

Source of truth, from affaan-m/everything-claude-code.


name: gget description: gget CLI and Python workflow for quick genomic database queries, sequence lookup, BLAST-style searches, enrichment checks, and reproducible bioinformatics evidence logs. metadata: origin: community

gget

Use this skill when a task needs quick bioinformatics lookup across genomic reference databases with the gget CLI or Python package.

When to Use

  • Finding Ensembl IDs, gene metadata, transcript details, or sequences.
  • Running quick BLAST or BLAT lookups without building a full local pipeline.
  • Fetching reference genome links and annotations from Ensembl.
  • Querying protein structure, pathway, cancer, expression, or disease-association modules through a single interface.
  • Creating a reproducible first-pass evidence log before moving to heavier tools such as Biopython, Snakemake, Nextflow, BLAST+, or database-specific clients.

Use a dedicated workflow instead of gget when the task requires regulated clinical interpretation, high-throughput production pipelines, or fine-grained control over database versions and local indexes.

Installation

Use a clean Python environment.

python -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install --upgrade gget
gget --help

If uv is available:

uv venv
. .venv/bin/activate
uv pip install gget

Before relying on an older environment, upgrade gget and re-check the module docs. The upstream databases queried by gget change over time.

Basic Patterns

CLI shape:

gget <module> [arguments] [options]

Python shape:

import gget

result = gget.search(["BRCA1"], species="human")
print(result)

Common workflow:

  1. Identify the species, assembly, gene ID type, and database needed.
  2. Check the current module documentation for arguments.
  3. Run a small query first.
  4. Save output with an explicit filename and date.
  5. Record module name, version, arguments, and database assumptions.

Common Modules

Use current upstream docs for exact arguments. These modules are common first choices:

  • gget search: find Ensembl IDs from search terms.
  • gget info: retrieve metadata for Ensembl, UniProt, or related IDs.
  • gget seq: fetch nucleotide or amino-acid sequences.
  • gget ref: retrieve reference genome download links.
  • gget blast: run a quick BLAST query.
  • gget blat: locate a sequence against supported genome assemblies.
  • gget muscle: run multiple sequence alignment.
  • gget diamond: run local sequence alignment against reference sequences.
  • gget alphafold and gget pdb: inspect protein-structure references.
  • gget enrichr, gget opentargets, gget archs4, gget bgee, gget cbio, and gget cosmic: explore enrichment, target, expression, cancer, and disease association data.

Do not assume every module supports every Python version or dependency set. Some optional scientific dependencies have narrower version support than the core package.

Quick Examples

Find genes:

gget search -s human brca1 dna repair -o brca1-search.json

Fetch gene metadata:

gget info ENSG00000012048 -o brca1-info.json

Fetch a sequence:

gget seq ENSG00000012048 -o brca1-seq.fa

Run a small BLAST query:

gget blast "MEEPQSDPSVEPPLSQETFSDLWKLLPEN" -l 10 -o blast-results.json

Python example:

import gget

genes = gget.search(["BRCA1", "DNA repair"], species="human")
info = gget.info(["ENSG00000012048"])
sequence = gget.seq("ENSG00000012048")

Reproducibility Log

For scientific outputs, include enough metadata to replay the query.

| Date | gget version | Module | Query | Species/assembly | Output | Notes |
| --- | --- | --- | --- | --- | --- | --- |
| 2026-05-11 | `gget --version` | search | `BRCA1 DNA repair` | human | `brca1-search.json` | Docs checked before run |

Also record:

  • Python version and environment manager.
  • Any optional dependency installed through gget setup.
  • Database-specific identifiers returned by the query.
  • Whether output is JSON, CSV, FASTA, or a DataFrame export.
  • Any failures that were resolved by upgrading gget.

Review Checklist

  • Did you upgrade or verify the installed gget version?
  • Did you check the current upstream module docs before using arguments?
  • Is the species or assembly explicit?
  • Are identifiers preserved exactly, including Ensembl/UniProt prefixes?
  • Is the result labeled as database output rather than clinical interpretation?
  • Is the query reproducible from the saved command or Python snippet?
  • Are optional dependencies installed in an isolated environment?

References