insforge-debug
insforge/insforge-skills
Diagnose InsForge project failures using logs, metrics, DB health, policies, and AI-assisted triage.
What is insforge-debug?
Insforge-debug provides observability primitives and symptom recipes for diagnosing reactive failures (SDK errors, HTTP 4xx/5xx, timeouts, auth issues, RLS denials) and proactive audits (security, performance, system health) in InsForge backends. Use it when you have an error or need a pre-launch readiness check.
- Query logs across 5 backend sources (InsForge, PostgREST, Postgres, edge functions, deploy logs)
- Check real-time DB health via 7 named checks (connections, slow queries, bloat, size, index usage, locks, cache hit)
- Review active RLS policies and metadata (auth config, tables, buckets, functions, realtime channels)
- Run static-scan advisor for security, performance, and health issues with specific recommendations
- Use AI-assisted triage to combine primitives and suggest diagnoses for concrete error descriptions
- Inspect deployment state and edge function deploy logs for Vercel and function failures
How to install insforge-debug
npx skills add https://github.com/insforge/insforge-skills --skill insforge-debug- Node.js and npm installed
- Access to an InsForge project with CLI credentials configured
- npx available (no global CLI installation needed)
How to use insforge-debug
- 1.Run `npx -y @insforge/cli diagnose --ai "<error description>"` for AI-assisted triage when you have a concrete error or symptom
- 2.Use the symptom recipes in the skill to identify which primitives to query (logs, db-health, policies, metadata, advisor)
- 3.Execute the primitive commands via `npx -y @insforge/cli <command>` (e.g., `logs <source>`, `diagnose db`, `db policies`)
- 4.Cross-reference findings against the per-primitive reference docs to interpret results and identify root cause
- 5.For RLS issues, verify policies against JWT claims and query as service role to distinguish filtering from missing data
- 6.For deploy failures, check deploy-state logs and run local `npm run build` to reproduce errors faster
Use cases
- Diagnose an SDK error with code/message by routing to the correct log source and checking DB health if needed
- Investigate HTTP 4xx/5xx responses on specific endpoints using error-object routing and metrics
- Debug RLS access issues (403 on write, empty results on read) by checking policies against JWT claims and actual data
- Troubleshoot login/OAuth failures and token expiration by verifying provider config and auth metadata
- Identify slow queries on one endpoint by checking Postgres logs, index usage, and RLS policy overhead
- Backend engineers debugging InsForge project failures
- DevOps/SRE teams performing system health checks and performance audits
- Full-stack developers investigating auth, RLS, or database issues
- Teams preparing for production launch needing pre-flight readiness checks
insforge-debug FAQ
Use `diagnose --ai` first when you have a concrete error message, failing URL, or HTTP status. It suggests which primitives to check. Skip it and go straight to primitives if you already know the issue category (e.g., you know it's an RLS problem) or if you need to verify the AI's diagnosis.
Reads that fail RLS checks return an empty array with no error logged. Query the table as service role using `npx -y @insforge/cli db query "SELECT id FROM <table>"` to confirm rows actually exist; if they do, RLS filtered them; if not, the data doesn't exist.
403 means the RLS policy blocked the request; check `postgREST.logs` for the policy violation event and verify policies against the JWT claim. 429 is rate-limited; check `metrics` for system-wide load and `logs` for the rate-limit event.
Use `diagnose advisor --category performance --json` to see `pg_stat_statements` text and mean execution time. Check `db-health slow-queries` only catches queries still running (>5s snapshot). Then review `index-usage` for missing indexes and `policies` for hidden RLS joins.
No. Always use `npx -y @insforge/cli` to run commands. Do not install the CLI globally; the skill enforces the npx pattern for consistency and version management.
Full instructions (SKILL.md)
Source of truth, from insforge/insforge-skills.
name: insforge-debug description: >- Use when diagnosing problems in an InsForge project — reactive failures (SDK error object, HTTP 4xx/5xx, gateway timeout 502/503/504, edge function failure or timeout, login/OAuth/auth errors, RLS denial, realtime channel issues, slow query on one endpoint, edge function or Vercel deploy failure), proactive audits (security/RLS review, performance/index review, system health check, pre-launch readiness), or when the user has an error but doesn't know where to start. license: Apache-2.0
InsForge Debug
Diagnose problems in InsForge projects by combining the backend's observability primitives — logs, metrics, db-health, advisor, policies, metadata, error objects, deploy state, and AI assist. This skill provides:
- A reference per debug primitive (one observability surface each — under
references/) - Symptom Recipes (below) that name the primitive sequence for known reactive symptoms and proactive audits
Always use npx -y @insforge/cli — never install the CLI globally.
Fastest Path: AI-Assisted Triage
When the user gives a concrete description (error message, failing URL, HTTP status), hand it to the InsForge debug agent. Unlike the other primitives, this one returns suggestions, not just observations — verify the diagnosis against the primitives it cites before acting on it.
npx -y @insforge/cli diagnose --ai "<issue description>"
See references/ai-assisted.md for when to use this first vs when to skip, and how to verify the output.
Debug Primitives
Each primitive is one independently-queryable observability surface backed by a distinct underlying data source. Real diagnoses are compositions of primitives.
All commands run via npx -y @insforge/cli .... The (command) shown next to each primitive is the actual CLI command — primitive names are concept labels, not CLI subcommand names (e.g., "DB health" is diagnose db, not diagnose db-health; "Policies" is db policies, not diagnose policies).
| Primitive (command) | What you see | Reference |
|---|---|---|
Logs (logs <source>; diagnose logs for cross-source aggregate) | Time-stream of events from 5 backend sources (insforge.logs / postgREST.logs / postgres.logs / function.logs / function-deploy.logs) | references/logs.md |
Metrics (diagnose metrics) | EC2 instance time-series (CPU / memory / disk / network) over 1h / 6h / 24h / 7d | references/metrics.md |
DB health (diagnose db) | Current Postgres state via 7 named checks (connections / slow-queries / bloat / size / index-usage / locks / cache-hit) | references/db-health.md |
Advisor (diagnose advisor --json) | Static-scan issues across 3 categories (security / performance / health) with ruleId / affectedObject / recommendation | references/advisor.md |
Policies (db policies) | Active RLS rules from pg_policies (USING / WITH CHECK per cmd per role) — returns all policies as a dump | references/policies.md |
Metadata (metadata --json) | Declarative backend state dump (auth config / tables / buckets / functions / AI models / realtime channels) | references/metadata.md |
| Error objects (no command — read SDK / HTTP response) | SDK error envelope + HTTP status — the routing table from a client-visible error to the right log source | references/error-objects.md |
Deploy state (deployments list + deployments status <id> --json + logs function-deploy.logs) | Frontend (Vercel) deployment history + per-deploy metadata, plus edge function deploy logs | references/deploy-state.md |
AI assist (diagnose --ai "<description>") | LLM agent that combines the other primitives — returns a diagnosis with suggestions | references/ai-assisted.md |
Symptom Recipes
Each recipe is a primitive call sequence with one-line "look for X" at each step. Command syntax, flags, and deep interpretation are in the per-primitive references above.
Recipe: SDK returned { data: null, error: { code, message } }
- error-objects — read code/message/details. If code starts with
PGRST*, route by prefix using the table in the reference. - logs (matching source per error-objects routing) — find the error timestamp, get the full backend-side context.
- db-health (
connections,locks,slow-queries) — only if the error suggests DB issue (PostgREST timeout, lock conflict).
Recipe: HTTP 4xx/5xx response on a specific request
- error-objects — use the HTTP status routing table to pick the log source (each status has a distinct path; 429 is special).
- logs (right source for that status) — find the failing request line and error.
- metrics — only for 5xx patterns spanning multiple endpoints, to confirm system-wide load issue.
Recipe: RLS access issue (403 on write, or empty result on read)
Same bug, two surfacings. Writes (INSERT / UPDATE / DELETE) fail loudly with 403. Reads (SELECT) fail silently with an empty array — PostgREST filters denied rows out instead of returning 403, so the request looks successful with zero rows. Diagnosis path is the same except step 1 only applies to the 403 variant.
- logs (
postgREST.logs) — 403 variant only: find the policy violation event with table and role context. Empty-result variant: skip — no error is logged for silently-filtered rows. - policies — list policies for that table; walk USING / WITH CHECK against the actual request and the JWT claim used.
- metadata — verify auth config (which claim feeds
auth.uid()/requesting_user_id(); for third-party auth like Clerk/Auth0, is the provider registered as a JWT issuer?). - db query (
db query "<sql>") — empty-result variant only: confirm rows that should be visible actually exist by querying as service role (not as the user):npx -y @insforge/cli db query "SELECT id, user_id FROM <table>". Distinguishes "RLS filtered everything" from "no matching data exists".
Recipe: Login fails / OAuth callback errors / token expired
- logs (
insforge.logs) — find auth errors with timestamp and provider context. - metadata — verify the provider is enabled, redirect URLs match the callback URL exactly (protocol + host + path).
Recipe: Edge function runtime error / timeout
- logs (
function.logs) — get the error stack and execution context. - metadata — confirm the function exists and
status: "active". - (If needed)
npx -y @insforge/cli functions code <slug>— inspect the source for obvious issues.
Recipe: functions deploy failed
- deploy-state (
function-deploy.logs) — find the build/push error. - metadata — confirm whether the function ended up in the active list (partial-deploy detection).
Recipe: deployments deploy failed (Vercel)
- deploy-state (
deployments list+status <id> --json) — readstatus,metadata.webhookEventType, andenvVarKeys. - Local
npm run build— reproduce the same error locally for faster iteration.
Recipe: Single slow query / one endpoint slow
- logs (
postgres.logs) — find the query text and timestamp. - db-health (
slow-queries,index-usage) —slow-queriesonly catches it while still running (>5s snapshot); checkindex-usagefor a missing index. Already finished? advisor (--category performance --json) has thepg_stat_statementstext + mean time; step 1 has the timestamp. - policies — if it's an RLS-gated table, verify the policy isn't adding hidden joins.
Recipe: "Memory is at ~80% but nothing is slow"
- Expected — say so first. A dedicated Postgres instance turns idle RAM into shared buffers and page cache; steady high memory with little traffic is its healthy state, not a leak (references/metrics.md, "Memory: high is normal").
- metrics (
--range 24h) — only a rising trend or OOM kills/restarts change the answer. OOM evidence lives inpostgres.logsas the crash-recovery aftermath ("terminating connection because of crash of another server process" / "automatic recovery in progress"). - With OOM evidence, the fix is headroom: upgrade to a paid plan and pick a larger instance size (dashboard → Project Settings → Compute & Disk). OOM on the smallest instances under real load is common and expected — never "restart to free memory".
Recipe: All responses slow / high CPU/memory (active incident)
- metrics (
--range 1h) — confirm system-wide pressure (CPU / memory / disk). - db-health — DB is the most common bottleneck; check
connections,locks,slow-queries. - logs (
diagnose logsaggregate) — error patterns across sources at the spike timestamp. - advisor (
--severity critical) — pre-existing known issues that may explain the degradation.
Recipe: Realtime channel won't connect / messages missing
- logs (
insforge.logs) — WebSocket errors and subscription failures. - metadata — verify the channel pattern matches what the client subscribes to,
enabled: true. - policies — RLS on the underlying table (realtime delivers row changes; RLS gates which rows the subscriber sees).
Recipe: 429 rate limit
- error-objects — confirm 429 status. No logs are recorded for 429s; no
Retry-Afterheader is returned. Don't waste time grepping logs. - metrics (
--range 1h) — overall backend load context. - Fix is always client-side: debounce, batch, exponential backoff, eliminate retry loops.
Recipe: Gateway timeout (502 / 503 / 504) on a specific URL
Route by URL subsystem before drilling:
| URL pattern | Drill into |
|---|---|
/api/database/records/... | logs (postgREST.logs → postgres.logs) + db-health (locks, slow-queries) |
/functions/<slug> | logs (function.logs) — function may be crash-looping |
/api/auth/... | logs (insforge.logs) |
| Any path during system-wide spike | metrics (--range 1h) |
504s across unrelated paths on a small instance: suspect OOM first. Intermittent gateway timeouts hitting database, auth, and functions alike are the classic out-of-memory signature on the smallest instance sizes: the kernel kills Postgres, every in-flight request times out at the gateway while crash recovery runs, and it repeats on the next load spike.
Fast path: npx -y @insforge/cli diagnose incident (Platform login required). The report is
built entirely on the cloud side — Prometheus scrape history, platform records, an outbound
database probe — so it works even while the instance is down or wedged, exactly when
diagnose logs stops answering. It returns a verdict (oom_likely,
platform_operation_in_progress, paused_or_suspended, metrics_stopped, down_unknown,
no_incident_detected) with the evidence and the recommended action; oom_likely already means
the restart/memory correlation checks below passed on the platform side.
If the command is unavailable (older CLI/backend, --api-key link mode), confirm manually in
logs (postgres.logs) via the crash-recovery aftermath — "terminating connection because of
crash of another server process" / "automatic recovery in progress" — time-correlated with the
5xx burst: recovery evidence alone only proves an unclean Postgres restart, so the timestamps
must line up before OOM becomes the leading diagnosis
(references/metrics.md). With that evidence the fix is headroom, not a
retry loop:
- Upgrade the instance —
npx -y @insforge/cli projects upgrade-instance <type>(nano→micro→small→medium→large→xl), or dashboard → Project Settings → Compute & Disk. On the free plan, upgrade to a paid plan first, then pick the size. The resize changes the bill and the CLI asks for interactive confirmation — get the user's go-ahead first, then run unattended with the CLI-level--yes(the-yinnpx -yis npm's install flag, not the confirm-skip). The resize is async — pollprojects getuntiloperation_statusclears before declaring the incident resolved. - The resize restarts the project as part of the change, which also clears any wedged state — there is no separate user-facing restart, and a bare restart would only buy minutes before the next spike OOMs again. OOM under real load on the smallest sizes is common and expected, not a bug.
Recipe: Pre-launch / proactive audit
Requires Platform login (
npx -y @insforge/cli login). Not available when the project is linked via--api-key— fall back todb-health+policies+metadatafor a manual audit in that case.
- advisor — full scan, then
--severity criticalfirst, then warnings. - advisor (
--category security) — focus on security issues; cross-verify with policies (RLS coverage) and metadata (auth config, public buckets, secret presence). - advisor (
--category performance) — cross-verify with db-health (slow-queries,index-usage,bloat). - advisor (
--category health) — cross-verify with metrics (resource trends over7d). - After fixes, re-run advisor and confirm
isResolved: truefor each addressedruleId.
Recipe: Don't know where to start
- ai-assisted (
diagnose --ai "<error or URL>") — get a starting hypothesis. - Verify by re-checking the primitives the diagnosis names. Trust the primitive observations over the suggestion.
When the Root Cause Is InsForge Itself
Some diagnoses end at an InsForge-side defect, not a project misconfiguration: a platform bug or regression, an SDK call that misbehaves, docs or a skill that contradict observed behavior, or a missing capability. A debug session is exactly where these get confirmed — report them while the evidence is in hand:
npx -y @insforge/cli feedback --json \
--type bug --component backend --area db \
--title "<one-line summary>" \
--detail "<what happened vs expected, minimal repro>" \
--command "<the failing call>" \
--error "<verbatim error from logs>" \
--workaround "<what you did instead>"
No login required; common PII patterns (emails, credential/key formats, public IPs, home-directory usernames) are redacted locally — pattern-based, so still keep user data out. Use --component sdk --language <lang> for SDK defects; --component docs or --component skills with --doc and --expected when documentation contradicts reality; --type feature-request when the finding is "not supported". Then continue the user's task with the workaround — never block on the report, and never file feedback for problems in the user's own app code or config. Full flag reference: the insforge-cli skill's Feedback section.
Related skills
More from insforge/insforge-skills and the wider catalog.

insforge-integrations
Wire external auth providers and x402 payment facilitators into InsForge for JWT-based RLS and onchain billing.

insforge
SDK integration for InsForge app backends: database, auth, storage, functions, AI, realtime, email, and payments.

insforge-cli
Manage InsForge backend and cloud infrastructure via CLI for projects, databases, deployments, and integrations.

insta
Operate InstaCloud infrastructure: deploy apps, manage databases, storage, compute, and branch environments via CLI.

instantdb
Build complete, functional apps with InstantDB as the backend for React, vanilla JS, and Expo.

anki-connect
This skill is for interacting with Anki through AnkiConnect, and should be used whenever a user asks to interact with Anki, including to read or modify decks, notes, cards, models, media, or sync operations.