clean-data-xls
anthropics/financial-services
Clean messy spreadsheet data: trim whitespace, fix casing, convert text-numbers, standardize dates, dedupe, and flag mixed types.
What is clean-data-xls?
Cleans up inconsistent or malformed data in Excel sheets by detecting and fixing common issues like whitespace, casing inconsistencies, numbers stored as text, mixed date formats, duplicates, and mixed-type columns. Use when preparing raw data for analysis or when spreadsheets have formatting problems.
- Trim leading/trailing whitespace and normalize spacing
- Standardize casing in categorical columns
- Convert numbers stored as text (with $, commas, %) to actual numbers
- Detect and standardize mixed date formats
- Identify and remove exact and near-duplicate rows
- Flag columns with mixed data types
How to install clean-data-xls
npx skills add https://github.com/anthropics/financial-services --skill clean-data-xlsHow to use clean-data-xls
- 1.Specify the range to clean (e.g., A1:F200) or leave blank to clean the entire used range
- 2.Review the proposed fixes summary showing detected issues, counts, and recommended actions
- 3.Confirm each category of fixes (whitespace → casing → number conversion → dates → deduplication) before applying
- 4.View before/after samples after each fix category to verify results
- 5.Accept the final cleaned data or undo specific categories if needed
Use cases
- Preparing raw financial or operational data for analysis before importing to BI tools
- Cleaning up customer or product lists with inconsistent formatting across entries
- Standardizing date columns that contain multiple formats from different data sources
- Removing duplicates and near-duplicates (differing only in whitespace or casing) from imported datasets
- Auditing data quality and identifying which columns need manual review before processing
- Data analysts preparing datasets for reporting
- Financial professionals cleaning transaction or account data
- Business users consolidating data from multiple sources
- Anyone working with messy Excel files before analysis or import
clean-data-xls FAQ
No by default. The skill uses helper columns with formulas to show cleaned results transparently. Only destructive operations (removing duplicates, overwriting originals) require explicit confirmation.
The skill detects mixed date formats in the same column and proposes standardization. It can convert common formats like 3/8/26, 2026-03-08, and March 8 2026 to a consistent format.
Yes. It detects numbers stored as text with $, commas, or % signs and converts them to actual numeric values using formulas like =VALUE(SUBSTITUTE(B2,"$","")).
It works in both. For Office Add-ins (Excel Online/desktop), it uses Office JS directly. For standalone .xlsx files, it uses Python/openpyxl.
Rows that are identical except for whitespace differences or casing variations (e.g., 'USA' vs 'usa'). The skill flags these for review before removal.
Full instructions (SKILL.md)
Source of truth, from anthropics/financial-services.
name: clean-data-xls description: Clean up messy spreadsheet data — trim whitespace, fix inconsistent casing, convert numbers-stored-as-text, standardize dates, remove duplicates, and flag mixed-type columns. Use when data is messy, inconsistent, or needs prep before analysis. Triggers on "clean this data", "clean up this sheet", "normalize this data", "fix formatting", "dedupe", "standardize this column", "this data is messy".
Clean Data
Clean messy data in the active sheet or a specified range.
Environment
- If running inside Excel (Office Add-in / Office JS): Use Office JS directly (
Excel.run(async (context) => {...})). Read viarange.values, write helper-column formulas viarange.formulas = [["=TRIM(A2)"]]. The in-place vs helper-column decision still applies. - If operating on a standalone .xlsx file: Use Python/openpyxl.
Workflow
Step 1: Scope
- If a range is given (e.g.
A1:F200), use it - Otherwise use the full used range of the active sheet
- Profile each column: detect its dominant type (text / number / date) and identify outliers
Step 2: Detect issues
| Issue | What to look for |
|---|---|
| Whitespace | leading/trailing spaces, double spaces |
| Casing | inconsistent casing in categorical columns (usa / USA / Usa) |
| Number-as-text | numeric values stored as text; stray $, ,, % in number cells |
| Dates | mixed formats in the same column (3/8/26, 2026-03-08, March 8 2026) |
| Duplicates | exact-duplicate rows and near-duplicates (case/whitespace differences) |
| Blanks | empty cells in otherwise-populated columns |
| Mixed types | a column that's 98% numbers but has 3 text entries |
| Encoding | mojibake (é, ’), non-printing characters |
| Errors | #REF!, #N/A, #VALUE!, #DIV/0! |
Step 3: Propose fixes
Show a summary table before changing anything:
| Column | Issue | Count | Proposed Fix |
|---|
Step 4: Apply
- Prefer formulas over hardcoded cleaned values — where the cleaned output can be expressed as a formula (e.g.
=TRIM(A2),=VALUE(SUBSTITUTE(B2,"$","")),=UPPER(C2),=DATEVALUE(D2)), write the formula in an adjacent helper column rather than computing the result in Python and overwriting the original. This keeps the transformation transparent and auditable. - Only overwrite in place with computed values when the user explicitly asks for it, or when no sensible formula equivalent exists (e.g. encoding/mojibake repair)
- For destructive operations (removing duplicates, filling blanks, overwriting originals), confirm with the user first
- After each category of fix (whitespace → casing → number conversion → dates → dedup), show the user a sample of what changed and get confirmation before moving to the next category
- Report a before/after summary of what changed
Related skills
More from anthropics/financial-services and the wider catalog.

competitive-analysis
Framework for building competitive landscape decks with market positioning, competitor deep-dives, and strategic synthesis.

comps-analysis
Build institutional-grade comparable company analyses with operating metrics, valuation multiples, and statistical benchmarking.

datapack-builder
Build investment-ready financial data packs from CIMs, SEC filings, and web sources into standardized Excel workbooks.

dcf-model
Build institutional-quality DCF equity valuation models with SEC data, cash flow projections, WACC calculations, and sensitivity analysis.

deck-refresh
Update presentation numbers across slides without rebuilding—quarterly refreshes, earnings swaps, and market data rebases.

earnings-analysis
Create professional 8-12 page equity research earnings update reports within 24-48 hours of quarterly results.