PluginBench
Skill
Review
Audit score 70

byted-mediakit-audio

volcengine/mediakit-cli

Process audio files and tracks: voice detection, metadata probing, transcoding, and voice-background separation.

What is byted-mediakit-audio?

Audio processing skill for voice activity detection, audio metadata extraction, transcoding, and voice/background separation in audio files or video tracks. Use when you need to understand audio content, convert formats, manage audio codecs, or isolate vocal elements.

  • Detect voice activity and speech boundaries in audio or video files
  • Extract standardized audio metadata (codec, bitrate, duration, etc.)
  • Separate human voice from background music and noise into independent files
  • Transcode audio between formats, codecs, and container types for different playback environments

How to install byted-mediakit-audio

npx skills add https://github.com/volcengine/mediakit-cli --skill byted-mediakit-audio
Prerequisites
  • mediakit-cli binary installed and available in PATH
  • Shared MediaKit skill (byted-mediakit-shared) must be installed first
  • Cloud API credentials configured for cloud-based operations (detect-voice-activity, separate-voice, transcode-audio)
Claude Code
Cursor
Windsurf
Cline

How to use byted-mediakit-audio

  1. 1.Load the shared MediaKit skill (byted-mediakit-shared) and run its pre-flight checks
  2. 2.Set environment variables: MEDIAKIT_SURFACE=skill and MEDIAKIT_RUNTIME to your agent name
  3. 3.Choose the appropriate tool based on your goal: detect-voice-activity for speech boundaries, probe-audio-metadata for file inspection, separate-voice for vocal isolation, or transcode-audio for format conversion
  4. 4.Provide the input audio/video file path or URL and required parameters (output format, codec, bitrate, etc.)
  5. 5.Execute the mediakit-cli command and retrieve the output (timestamps, metadata JSON, or processed audio files)

Use cases

Good for
  • Automatically identify and extract speech segments from recordings, filtering out silence and background noise
  • Convert audio files to different formats (MP3, AAC, WAV) and bitrates for platform compatibility
  • Split podcast or music tracks into isolated vocal and instrumental components for remixing or analysis
  • Inspect audio properties before processing to validate format and quality requirements
Who it's for
  • Audio engineers and producers
  • Content creators working with podcasts or music
  • Developers building audio processing pipelines
  • Media asset managers handling format standardization

byted-mediakit-audio FAQ

What's the difference between this skill and video editing?

This skill handles audio-specific tasks like voice detection, metadata extraction, transcoding, and voice separation. Audio editing (trimming, mixing, extracting tracks from video) belongs in the editing skill; video understanding and subtitle generation belong in the video skill.

Can I use these tools locally or do I need cloud APIs?

probe-audio-metadata supports both local and cloud modes. detect-voice-activity, separate-voice, and transcode-audio require cloud APIs.

What audio formats are supported?

Refer to the tool-specific reference documentation (reference/transcode-audio.md, etc.) for supported input and output formats, codecs, and containers.

How do I handle authentication?

Set up cloud API credentials as environment variables before running. The skill will use MEDIAKIT_SURFACE=skill and MEDIAKIT_RUNTIME to identify the calling agent.

What should I do if I'm unsure whether to use audio, editing, or video skills?

Clarify your goal: audio processing (transcoding, voice detection, separation) → audio skill; audio editing (trim, mix, merge) → editing skill; video understanding or subtitle generation → video skill.

Full instructions (SKILL.md)

Source of truth, from volcengine/mediakit-cli.


name: byted-mediakit-audio version: "0.2.1" license: "MIT" description: "面向音频文件或视频中的音轨,处理语音边界定位、音频媒资信息探测、音频转码与码流封装适配、人声与背景声分离等目标。若对象和目标族已明确属于音频内容理解、音频转码、音频格式治理或音轨分离,但具体做法不确定,可先加载本 Skill 探索。" permissions:

  • shell metadata: requires: bins: ["mediakit-cli"] cliHelp: "mediakit-cli audio --help" product: mediakit-cli/skills domain: audio capability_count: 4

audio MediaKit Skill

使用规则

  1. 先读取 ../byted-mediakit-shared/SKILL.md,执行统一前置检查;该 Skill 缺失时停止并提示安装。
  2. 只从下表选择 audio 域工具;相似能力按各工具“能力描述”和参数边界区分。
  3. 执行前按需读取对应 reference;参数与结果说明来自同一份已审核文案,完整机器合同以当前 CLI --schema 为准。
  4. 缺少必填参数、鉴权环境变量或真实输入资源时,向用户索取;通用可选字段只能透传用户明确提供的值,其他可选字段可由明确意图准确确定,但不得伪造。
  5. 执行时设置 MEDIAKIT_SURFACE=skill,指定调用来源是 Skill。
  6. 执行时把 MEDIAKIT_RUNTIME 设置为当前 Agent 宿主,避免 CLI 无法可靠识别父级 Agent。

澄清与跨域路由

若只说有音频而未说明业务目标,应先澄清。音频裁剪、拼接、调速、淡入淡出、混音、从视频抽取音轨或音视频合流等编辑合成诉求应路由到 editing;字幕生成、提取字幕、语音转字幕、视频理解、视频增强等应路由到 video。

工具列表

工具说明支持模式命令参考
detect-voice-activity用于语音端点识别。自动定位音频或视频文件中有效语音的起止时间。将人声和静音、背景噪声等无效片段区分开来。返回包含所有有效人声片段起止时间戳的列表。Cloudmediakit-cli audio detect-voice-activityreference/detect-voice-activity.md
probe-audio-metadata探测输入音频 URL,输出标准化媒资元信息,用于获取音频元信息。Cloud / Localmediakit-cli audio probe-audio-metadatareference/probe-audio-metadata.md
separate-voice用于人声背景声分离,可将音频或视频文件中的人声与背景音精准分离,输出为两个独立的音频文件。Cloudmediakit-cli audio separate-voicereference/separate-voice.md
transcode-audio音频转码将一个音频码流转换为另一个音频码流,通常涉及编码格式、编码参数和封装格式的转换,用于适应不同业务场景、播放终端和网络环境。Cloudmediakit-cli audio transcode-audioreference/transcode-audio.md