PluginBench
Skill
Review
Audit score 70

nutrient-document-processing

affaan-m/ecc

Convert, extract, OCR, redact, sign, and fill documents via Nutrient DWS API.

What is nutrient-document-processing?

Process documents across PDFs, DOCX, XLSX, PPTX, HTML, and images using the Nutrient Document Web Services API. Use it to convert formats, extract text and tables, OCR scanned documents, redact sensitive data, add watermarks, digitally sign, and auto-fill PDF forms.

  • Convert between PDF, DOCX, XLSX, PPTX, HTML, and image formats
  • Extract plain text and structured tables from documents
  • OCR scanned documents and images in 100+ languages
  • Redact PII using preset patterns (SSN, email, credit card, phone, etc.) or custom regex
  • Add watermarks with customizable text, size, opacity, and rotation
  • Digitally sign PDFs with CMS signatures

How to install nutrient-document-processing

npx skills add null --skill nutrient-document-processing
Prerequisites
  • Free Nutrient API key from nutrient.io dashboard
  • NUTRIENT_API_KEY environment variable set
  • curl or HTTP client for API calls, or Node.js for MCP server integration
Claude Code
Cursor
Windsurf
Cline

How to use nutrient-document-processing

  1. 1.Sign up for a free API key at nutrient.io/dashboard
  2. 2.Set the NUTRIENT_API_KEY environment variable
  3. 3.Choose an operation (convert, extract, OCR, redact, watermark, sign, or fill)
  4. 4.Prepare your input document and craft the corresponding curl request or use the MCP server
  5. 5.Send the multipart POST request to https://api.nutrient.io/build with document and instructions JSON
  6. 6.Retrieve the processed output file

Use cases

Good for
  • Convert a DOCX contract to PDF for signing and archival
  • Extract tables from a scanned expense report into Excel for processing
  • OCR a batch of historical documents to make them searchable
  • Redact employee SSNs and email addresses before sharing HR documents
  • Watermark draft proposals with 'CONFIDENTIAL' before distribution
Who it's for
  • Document automation engineers
  • Compliance and legal teams handling sensitive documents
  • Data extraction specialists
  • Workflow automation developers
  • Organizations processing high-volume document conversions

nutrient-document-processing FAQ

What document formats are supported?

Input: PDF, DOCX, XLSX, PPTX, DOC, XLS, PPT, PPS, PPSX, ODT, RTF, HTML, JPG, PNG, TIFF, HEIC, GIF, WebP, SVG, TGA, EPS. Output: PDF, DOCX, XLSX, text, and more depending on the operation.

How many languages does OCR support?

Over 100 languages via ISO 639-2 codes (e.g., eng, deu, fra, jpn, chi_sim, ara) and full language names like 'english' or 'german'. See the complete language table in the Nutrient docs.

Can I redact multiple patterns in one request?

Yes, chain multiple redaction actions in the instructions JSON—e.g., redact SSNs and email addresses in the same operation.

Do I need to use curl, or are there other integration options?

You can use curl directly, or integrate via the MCP server (@nutrient-sdk/dws-mcp-server) for native tool support in Claude or Cursor.

Is there a cost for the API?

Nutrient offers a free tier; review their pricing and terms at nutrient.io before production use.

Full instructions (SKILL.md)

Source of truth, from affaan-m/ecc.


name: nutrient-document-processing description: Process, convert, OCR, extract, redact, sign, and fill documents using the Nutrient DWS API. Works with PDFs, DOCX, XLSX, PPTX, HTML, and images. metadata: origin: ECC

Nutrient Document Processing

Note: This skill integrates with the Nutrient commercial API. Review their terms before use.

Process documents with the Nutrient DWS Processor API. Convert formats, extract text and tables, OCR scanned documents, redact PII, add watermarks, digitally sign, and fill PDF forms.

Setup

Get a free API key at nutrient.io

export NUTRIENT_API_KEY="pdf_live_..."

All requests go to https://api.nutrient.io/build as multipart POST with an instructions JSON field.

Operations

Convert Documents

# DOCX to PDF
curl -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.docx=@document.docx" \
  -F 'instructions={"parts":[{"file":"document.docx"}]}' \
  -o output.pdf

# PDF to DOCX
curl -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"output":{"type":"docx"}}' \
  -o output.docx

# HTML to PDF
curl -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "index.html=@index.html" \
  -F 'instructions={"parts":[{"html":"index.html"}]}' \
  -o output.pdf

Supported inputs: PDF, DOCX, XLSX, PPTX, DOC, XLS, PPT, PPS, PPSX, ODT, RTF, HTML, JPG, PNG, TIFF, HEIC, GIF, WebP, SVG, TGA, EPS.

Extract Text and Data

# Extract plain text
curl -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"output":{"type":"text"}}' \
  -o output.txt

# Extract tables as Excel
curl -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"output":{"type":"xlsx"}}' \
  -o tables.xlsx

OCR Scanned Documents

# OCR to searchable PDF (supports 100+ languages)
curl -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "scanned.pdf=@scanned.pdf" \
  -F 'instructions={"parts":[{"file":"scanned.pdf"}],"actions":[{"type":"ocr","language":"english"}]}' \
  -o searchable.pdf

Languages: Supports 100+ languages via ISO 639-2 codes (e.g., eng, deu, fra, spa, jpn, kor, chi_sim, chi_tra, ara, hin, rus). Full language names like english or german also work. See the complete OCR language table for all supported codes.

Redact Sensitive Information

# Pattern-based (SSN, email)
curl -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"actions":[{"type":"redaction","strategy":"preset","strategyOptions":{"preset":"social-security-number"}},{"type":"redaction","strategy":"preset","strategyOptions":{"preset":"email-address"}}]}' \
  -o redacted.pdf

# Regex-based
curl -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"actions":[{"type":"redaction","strategy":"regex","strategyOptions":{"regex":"\\b[A-Z]{2}\\d{6}\\b"}}]}' \
  -o redacted.pdf

Presets: social-security-number, email-address, credit-card-number, international-phone-number, north-american-phone-number, date, time, url, ipv4, ipv6, mac-address, us-zip-code, vin.

Add Watermarks

curl -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"actions":[{"type":"watermark","text":"CONFIDENTIAL","fontSize":72,"opacity":0.3,"rotation":-45}]}' \
  -o watermarked.pdf

Digital Signatures

# Self-signed CMS signature
curl -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"actions":[{"type":"sign","signatureType":"cms"}]}' \
  -o signed.pdf

Fill PDF Forms

curl -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "form.pdf=@form.pdf" \
  -F 'instructions={"parts":[{"file":"form.pdf"}],"actions":[{"type":"fillForm","formFields":{"name":"Jane Smith","email":"jane@example.com","date":"2026-02-06"}}]}' \
  -o filled.pdf

MCP Server (Alternative)

For native tool integration, use the MCP server instead of curl:

{
  "mcpServers": {
    "nutrient-dws": {
      "command": "npx",
      "args": ["-y", "@nutrient-sdk/dws-mcp-server"],
      "env": {
        "NUTRIENT_DWS_API_KEY": "YOUR_API_KEY",
        "SANDBOX_PATH": "/path/to/working/directory"
      }
    }
  }
}

When to Use

  • Converting documents between formats (PDF, DOCX, XLSX, PPTX, HTML, images)
  • Extracting text, tables, or key-value pairs from PDFs
  • OCR on scanned documents or images
  • Redacting PII before sharing documents
  • Adding watermarks to drafts or confidential documents
  • Digitally signing contracts or agreements
  • Filling PDF forms programmatically

Links