gget
affaan-m/ecc
Quick genomic database queries, BLAST searches, and sequence lookups via unified CLI and Python interface.
What is gget?
gget is a CLI and Python package for rapid bioinformatics lookups across Ensembl, UniProt, and other genomic reference databases. Use it for gene searches, sequence retrieval, BLAST/BLAT queries, and enrichment checks when you need fast first-pass evidence before moving to production pipelines.
- Search for Ensembl IDs and gene metadata from text queries
- Fetch nucleotide and amino-acid sequences from gene or protein identifiers
- Run BLAST and BLAT sequence alignment queries without local setup
- Retrieve reference genome download links and annotations
- Query protein structures, pathways, expression, cancer, and disease associations through unified modules
- Generate reproducible evidence logs with explicit version and parameter tracking
How to install gget
npx skills add null --skill gget- Python 3.7 or later
- Virtual environment (venv or uv recommended)
- pip or uv package manager
How to use gget
- 1.Create and activate a clean Python virtual environment
- 2.Install gget with `pip install --upgrade gget`
- 3.Verify installation with `gget --help`
- 4.Identify the species, assembly, and gene ID type for your query
- 5.Choose the appropriate module (search, info, seq, blast, ref, etc.)
- 6.Run a small test query first, then save output with explicit filename and date
- 7.Record module name, version, arguments, and database assumptions in a reproducibility log
Use cases
- Finding BRCA1 gene metadata and sequences for a literature review
- Running a quick BLAST search on a protein sequence before committing to a full alignment pipeline
- Fetching reference genome links for multiple species in a comparative genomics task
- Querying enrichment and disease-association data for a candidate gene list
- Creating a timestamped, reproducible log of database queries for a research report
- Bioinformaticians doing exploratory analysis
- Researchers needing quick genomic lookups without building full pipelines
- Scientists preparing evidence logs for reproducible research
- Developers integrating genomic queries into Python workflows
gget FAQ
Use gget for exploratory queries, first-pass evidence gathering, and quick lookups. Switch to dedicated tools like Biopython, Snakemake, or Nextflow when you need regulated clinical interpretation, high-throughput production runs, or fine-grained control over database versions.
gget supports Python 3.7+, but some optional scientific dependencies have narrower version support. Always verify compatibility in a fresh environment and upgrade gget before relying on an older setup.
Save output with explicit filenames and dates, record the gget version, module name, exact arguments, species/assembly, and any optional dependencies installed. Include this metadata in a reproducibility log table.
First, upgrade gget to the latest version—upstream databases change over time. Then check the current upstream documentation for the module and verify your arguments match the expected format.
Yes. Import gget and call functions like `gget.search()`, `gget.info()`, `gget.seq()`, and `gget.blast()` directly. Return values are typically DataFrames or JSON-serializable objects.
Full instructions (SKILL.md)
Source of truth, from affaan-m/ecc.
name: gget description: gget CLI and Python workflow for quick genomic database queries, sequence lookup, BLAST-style searches, enrichment checks, and reproducible bioinformatics evidence logs. metadata: origin: community
gget
Use this skill when a task needs quick bioinformatics lookup across genomic
reference databases with the gget CLI or Python package.
When to Use
- Finding Ensembl IDs, gene metadata, transcript details, or sequences.
- Running quick BLAST or BLAT lookups without building a full local pipeline.
- Fetching reference genome links and annotations from Ensembl.
- Querying protein structure, pathway, cancer, expression, or disease-association modules through a single interface.
- Creating a reproducible first-pass evidence log before moving to heavier tools such as Biopython, Snakemake, Nextflow, BLAST+, or database-specific clients.
Use a dedicated workflow instead of gget when the task requires regulated
clinical interpretation, high-throughput production pipelines, or fine-grained
control over database versions and local indexes.
Installation
Use a clean Python environment.
python -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install --upgrade gget
gget --help
If uv is available:
uv venv
. .venv/bin/activate
uv pip install gget
Before relying on an older environment, upgrade gget and re-check the module
docs. The upstream databases queried by gget change over time.
Basic Patterns
CLI shape:
gget <module> [arguments] [options]
Python shape:
import gget
result = gget.search(["BRCA1"], species="human")
print(result)
Common workflow:
- Identify the species, assembly, gene ID type, and database needed.
- Check the current module documentation for arguments.
- Run a small query first.
- Save output with an explicit filename and date.
- Record module name, version, arguments, and database assumptions.
Common Modules
Use current upstream docs for exact arguments. These modules are common first choices:
gget search: find Ensembl IDs from search terms.gget info: retrieve metadata for Ensembl, UniProt, or related IDs.gget seq: fetch nucleotide or amino-acid sequences.gget ref: retrieve reference genome download links.gget blast: run a quick BLAST query.gget blat: locate a sequence against supported genome assemblies.gget muscle: run multiple sequence alignment.gget diamond: run local sequence alignment against reference sequences.gget alphafoldandgget pdb: inspect protein-structure references.gget enrichr,gget opentargets,gget archs4,gget bgee,gget cbio, andgget cosmic: explore enrichment, target, expression, cancer, and disease association data.
Do not assume every module supports every Python version or dependency set. Some optional scientific dependencies have narrower version support than the core package.
Quick Examples
Find genes:
gget search -s human brca1 dna repair -o brca1-search.json
Fetch gene metadata:
gget info ENSG00000012048 -o brca1-info.json
Fetch a sequence:
gget seq ENSG00000012048 -o brca1-seq.fa
Run a small BLAST query:
gget blast "MEEPQSDPSVEPPLSQETFSDLWKLLPEN" -l 10 -o blast-results.json
Python example:
import gget
genes = gget.search(["BRCA1", "DNA repair"], species="human")
info = gget.info(["ENSG00000012048"])
sequence = gget.seq("ENSG00000012048")
Reproducibility Log
For scientific outputs, include enough metadata to replay the query.
| Date | gget version | Module | Query | Species/assembly | Output | Notes |
| --- | --- | --- | --- | --- | --- | --- |
| 2026-05-11 | `gget --version` | search | `BRCA1 DNA repair` | human | `brca1-search.json` | Docs checked before run |
Also record:
- Python version and environment manager.
- Any optional dependency installed through
gget setup. - Database-specific identifiers returned by the query.
- Whether output is JSON, CSV, FASTA, or a DataFrame export.
- Any failures that were resolved by upgrading
gget.
Review Checklist
- Did you upgrade or verify the installed
ggetversion? - Did you check the current upstream module docs before using arguments?
- Is the species or assembly explicit?
- Are identifiers preserved exactly, including Ensembl/UniProt prefixes?
- Is the result labeled as database output rather than clinical interpretation?
- Is the query reproducible from the saved command or Python snippet?
- Are optional dependencies installed in an isolated environment?
References
Related skills
More from affaan-m/ecc and the wider catalog.
git-workflow
Git workflow patterns, branching strategies, and collaborative development best practices.
github-ops
Automate GitHub issue triage, PR management, CI/CD debugging, and releases using gh CLI.
golang-patterns
Idiomatic Go patterns, best practices, and conventions for building robust, efficient, and maintainable applications.
golang-testing
Go testing patterns: table-driven tests, subtests, benchmarks, fuzzing, and TDD methodology.
google-workspace-ops
Operate Google Drive, Docs, Sheets, and Slides as unified workflows for plans, trackers, and shared documents.
growth-log
Agent skill from affaan-m/ecc.