audio-transcription
postplusai/postplus-skills
Transcribe audio to text with timestamps, or convert existing transcripts to SRT/ASS subtitles.
What is audio-transcription?
Transcribe local or remote audio files into text with precise timestamps suitable for subtitles. If you already have a timed transcript, convert it locally to SRT or ASS format without re-transcribing. Use this for speech-to-text, subtitle generation, multilingual transcription, and creating durable transcript artifacts.
- Transcribe audio (local files or HTTPS URLs) to text with timestamps
- Generate subtitle-ready output with precise timing information
- Convert existing timed transcripts to SRT or ASS subtitle formats locally without re-transcribing
- Support multilingual transcription
- Provide rough speech search and transcript artifacts
How to install audio-transcription
npx skills add https://github.com/postplusai/postplus-skills --skill audio-transcription- PostPlus CLI installed and configured
- Audio file (local path, HTTPS URL, or PostPlus media reference) or existing timed transcript
- Media duration in seconds for validation before submission
How to use audio-transcription
- 1.Prepare your audio source: local file path, HTTPS URL, or existing PostPlus media reference
- 2.Determine the audio duration in seconds
- 3.Run the transcription command with `postplus media transcribe transcription --audio <source> --duration-seconds <duration> --wait --output <result.json>`
- 4.If converting an existing transcript to subtitles, use the local subtitle conversion reference (subtitles.md) to convert the timed transcript to SRT or ASS format without re-transcribing
- 5.Check the CLI output for status; if pending, preserve the result path and use the returned resume command rather than resubmitting
Use cases
- Generate subtitles for video content by transcribing audio and converting to SRT/ASS format
- Create searchable transcripts of interviews, podcasts, or meetings with timestamps
- Convert an existing timed transcript to multiple subtitle formats for different platforms
- Transcribe multilingual audio content for international audiences
- Produce durable transcript artifacts for archival or accessibility purposes
- Video producers and editors
- Content creators working with podcasts or interviews
- Accessibility specialists creating subtitles
- Researchers and documentarians managing audio archives
- Media production teams
audio-transcription FAQ
Use audio-transcription when your input is audio and the main task is speech-to-text or subtitle generation. Use video-transcription for video inputs and media-analysis for semantic video understanding.
Yes. If you already have a timed transcript, skip transcription and use the local subtitle conversion reference to convert it to SRT or ASS format directly without submitting another hosted job.
Preserve the result path and follow the CLI-returned action or resume command for the same operation. Do not submit another job; wait for the first one to complete or reach the recovery boundary.
Yes. A higher-quality default model is available for subtitle-quality output, and a faster, cheaper variant is available for rough transcription passes.
The skill generates JSON transcripts with timestamps. You can then convert these locally to SRT or ASS subtitle formats using the provided subtitle conversion reference.
Full instructions (SKILL.md)
Source of truth, from postplusai/postplus-skills.
name: audio-transcription description: Transcribe local or remote audio into text and timestamps. Convert an existing timed transcript locally into SRT or ASS without another transcription job. metadata: postplus: familyId: media-production familyName: Media and Creative Production
Audio Transcription
Use When
- The input is audio and the main job is speech-to-text, subtitle-ready timing, rough speech search, multilingual transcription, or durable transcript artifacts.
- Use
video-transcriptionfor video inputs andmedia-analysisfor semantic video understanding.
If a timed transcript already exists and only subtitle output is requested, skip transcription and read the local subtitle conversion reference.
Do Not Use When
- The task needs new creative generation or visual analysis rather than speech or subtitles.
- Required inputs are missing and guessing would change the result.
Execution Boundary
- Hosted transcription runs through the public
postplus media transcribeverb and is async. A submit records the run handle, current status, and completed artifacts when available. - Pass a local path, HTTPS URL, existing PostPlus media reference, or data URI
directly to
--audio. The CLI validates and prepares local media before the single hosted submit. - A higher-quality default model and a faster, cheaper variant are available; prefer the default when subtitle quality matters and use the cheaper variant for an explicit rough pass. The generated example below shows the default endpoint key.
Source And Path
- Supply the media duration so PostPlus can validate the request before it runs; a missing duration fails before submission.
- Request timestamps when the output will feed subtitles or edit decisions.
- Start with one source file or audio URL before larger batches.
- Keep internal requests, responses, manifests, normalized transcripts, and
downloaded artifacts under
.postplus/audio-transcription; keep final user-facing transcript exports outside.postplus.
Handoff
- If status is pending, preserve the result path and follow the CLI-returned action or resume command for the same operation. Do not submit another job. Stop and report when the CLI wait/recovery boundary is reached.
- When SRT/ASS is requested, use the actual timed transcript and read local subtitle conversion. Convert locally without another hosted request; do not invent a CLI export command.
Stop Conditions
- Stop when required user intent, source evidence, or owned input artifacts are missing and guessing would change the result.
Public Command Boundary
-
Choose the smallest matching command or workflow from the user input and run it directly.
-
Readiness diagnostics:
postplus doctor --skill audio-transcription. -
Use
postplus media schema --jsononly when you need the full endpoint, flag, and enum contract or are repairing an unknown request shape. -
Run the hosted transcription job with the generated command below; do not use another execution interface.
-
Pass the source directly through
--audio; do not pre-upload it or construct a manual request object.
postplus media transcribe transcription \
--audio ./reference.wav \
--duration-seconds 1 \
--wait \
--output ./result.json
Follow the CLI's structured result and reported next action; do not infer recovery from free-text messages. Wait for explicit user approval when requested; an action does not authorize spending, publishing, or overwriting. Resume the same operation through its returned checkpoint or action; never resubmit uncertain work, repeat exhausted recovery, or switch providers to bypass failure.
<!-- END GENERATED EXECUTION EXAMPLE -->- If the CLI returns a quote-confirmation challenge, obtain user approval for its scope and cost before running
postplus quote confirm --json --challenge-file <challenge.json>and retry with the returned token.
Related skills
More from postplusai/postplus-skills and the wider catalog.

instagram-account-research
Research Instagram accounts for creator discovery, competitor profiling, and account health snapshots using hosted profile and post collection.

adversarial-review
Cross-model adversarial code review that spawns opposing AI reviewers to challenge work from distinct critical lenses.

powersync
Guided onboarding and best practices for building applications with PowerSync — Cloud and self-hosted setup, sync configuration, client SDK usage, backend integration (Supabase, custom Postgres, MongoDB, MySQL, MSSQL), and debugging. Use this skill whenever the user mentions PowerSync, offline-first sync, local-first architecture, sync rules, sync streams, uploadData, fetchCredentials, real-time data replication, or wants to add offline-capable sync to a mobile or web app — even if they don't explicitly name PowerSync.

chrome-extension
Chrome Extensions (Manifest V3) performance and code quality guidelines. Use when writing, reviewing, or refactoring Chrome extension code including service workers, content scripts, message passing, storage APIs, TypeScript patterns, and testing.

clean-architecture
Clean Architecture principles and best practices for designing maintainable, testable software systems.

clean-code
Use when writing, reviewing, or refactoring code for maintainability and readability. Triggers on code reviews, naming discussions, function design, error handling, and test writing. Based on Robert C. Martin's Clean Code handbook with modern corrections.