PluginBench
Skill
Fail
Audit score 45

metaclaw-evolving-agent

aradotso/trending-skills

Deploy an evolving agent that learns from conversations via skills injection, RL training, and smart scheduling.

What is metaclaw-evolving-agent?

MetaClaw is an OpenAI-compatible proxy agent that intercepts conversations, injects learned skills, and continuously improves through real-world interactions. Use it when you want an agent that adapts and evolves without manual retraining, with flexible modes for skills-only, immediate RL, or deferred training during idle windows.

  • Intercepts conversations and injects relevant learned skills into system prompts via vector search
  • Trains continuously using GRPO reinforcement learning when batches fill or during scheduled windows
  • Summarizes conversation patterns into reusable skills automatically at session end
  • Schedules weight updates to idle/sleep windows and calendar meetings to avoid interrupting active use
  • Supports multiple LLM providers (Kimi, Qwen, Claude, Minimax, OpenAI, Gemini) via OpenAI-compatible proxy
  • Separates experience into support and query sets to prevent reward staleness

How to install metaclaw-evolving-agent

npx skills add https://github.com/aradotso/trending-skills --skill metaclaw-evolving-agent
Prerequisites
  • Python 3.9+
  • LLM API key (Kimi, Qwen, Claude, Minimax, OpenAI, or Gemini)
  • Optional: GPU/Tinker API key for RL training mode
  • Optional: Google Calendar credentials for scheduler mode
Claude Code
Cursor
Windsurf
Cline

How to use metaclaw-evolving-agent

  1. 1.Run `metaclaw setup` to create interactive config at ~/.metaclaw/config.yaml
  2. 2.Set METACLAW_LLM_API_KEY environment variable with your LLM provider key
  3. 3.Run `metaclaw start` to launch in default madmax mode (skills + RL + scheduler)
  4. 4.Point your OpenAI SDK client to http://localhost:<port>/v1 instead of upstream endpoint
  5. 5.Skills are injected transparently; no client code changes needed
  6. 6.Monitor ~/.metaclaw/logs for training and skill injection activity

Use cases

Good for
  • Run a coding assistant that learns your team's code review patterns and injects them into future reviews
  • Deploy a customer support agent that evolves responses based on successful resolution patterns
  • Set up a research agent that discovers and applies effective search strategies from past sessions
  • Configure an agent that trains only during off-hours to avoid latency during peak usage
  • Build an agent that learns domain-specific terminology and best practices from live conversations
Who it's for
  • ML engineers building adaptive agent systems
  • Teams wanting agents that improve without manual retraining
  • Developers needing low-latency inference with background learning
  • Organizations with off-peak windows (sleep hours, meetings) for training
  • Anyone using OpenAI SDK clients who wants transparent skill injection

metaclaw-evolving-agent FAQ

Do I need a GPU to run MetaClaw?

No. Skills-only mode requires no GPU. RL training uses Tinker or Mint APIs, so you don't need local GPU hardware—just an API key.

What's the difference between skills_only, rl, and madmax modes?

skills_only: proxy + skill injection, no training. rl: trains immediately when batch fills. madmax: trains only during sleep hours, idle timeouts, or calendar meetings.

How are skills created and updated?

Skills are auto-summarized from conversation patterns at session end via LLM. You can also manually add skills via SkillStore.add() or the Python API.

Can I use MetaClaw with my existing OpenAI SDK code?

Yes. Just change the base_url to point to the MetaClaw proxy (http://localhost:8080/v1). Skills inject transparently without code changes.

What happens to my conversation data?

Conversations are stored in ExperienceBuffer for training and skill extraction. They're not sent to external services unless you enable Google Calendar integration.

Full instructions (SKILL.md)

Source of truth, from aradotso/trending-skills.


name: metaclaw-evolving-agent description: Deploy and configure MetaClaw — an agent that meta-learns and evolves from live conversations using skills injection, RL training, and smart scheduling. triggers:

  • set up metaclaw agent
  • configure evolving agent
  • metaclaw skills mode
  • metaclaw rl training
  • metaclaw madmax scheduler
  • agent meta-learning setup
  • tinker rl backend configuration
  • metaclaw proxy deployment

MetaClaw Evolving Agent

Skill by ara.so — Daily 2026 Skills collection

MetaClaw is an OpenAI-compatible proxy agent that intercepts conversations, injects learned skills, and continuously improves itself through real-world interactions. It supports three modes: lightweight skills injection, immediate RL training, and a smart "madmax" scheduler that defers weight updates to idle/sleep windows.


Installation

# Minimal — skills injection only, no GPU required
pip install -e .

# Full RL training support (torch, transformers, tinker)
pip install -e ".[rl]"

# Skill evolution via LLM summarization
pip install -e ".[evolve]"

# Google Calendar scheduler for madmax mode
pip install -e ".[scheduler]"

# Recommended: everything
pip install -e ".[rl,evolve,scheduler]"

Quick Start

# One-time interactive config wizard
metaclaw setup

# Start in default madmax mode (skills + RL + smart scheduler)
metaclaw start

# Skills only — no GPU, no Tinker needed
metaclaw start --mode skills_only

# RL mode — trains immediately when batch is full
metaclaw start --mode rl

# RL without scheduler (same as above, explicit)
metaclaw start --mode rl

After metaclaw start, a local OpenAI-compatible proxy is running. Point your client (OpenClaw or any OpenAI SDK consumer) at http://localhost:<port> instead of the upstream LLM endpoint.


Configuration

metaclaw setup writes a config file (default: ~/.metaclaw/config.yaml). You can also edit it directly:

# ~/.metaclaw/config.yaml

proxy:
  host: 0.0.0.0
  port: 8080

llm:
  provider: kimi          # kimi | qwen | claude | minimax | openai | gemini
  base_url: https://api.moonshot.cn/v1
  model: moonshot-v1-8k
  # api_key loaded from env: METACLAW_LLM_API_KEY

skills:
  enabled: true
  max_injected: 5         # max skills injected per turn
  summarize_after_session: true

rl:
  enabled: true
  backend: auto           # auto | tinker | mint
  batch_size: 32
  algorithm: grpo
  opd_teacher: false      # optional teacher distillation

scheduler:                # madmax mode only
  enabled: true
  sleep_hours: [22, 7]    # local 22:00–07:00
  idle_timeout_minutes: 15
  google_calendar: false  # set true + configure OAuth for meeting detection

logging:
  level: info
  log_dir: ~/.metaclaw/logs

Environment Variables

export METACLAW_LLM_API_KEY="your-llm-api-key"
export METACLAW_TINKER_API_KEY="your-tinker-api-key"   # rl mode
export METACLAW_MINT_API_KEY="your-mint-api-key"        # if backend=mint
export GOOGLE_CALENDAR_CREDENTIALS_PATH="path/to/creds.json"  # scheduler

Operating Modes

ModeCommandGPU RequiredDescription
skills_onlymetaclaw start --mode skills_onlyNoProxy + skills injection + auto-summarization
rlmetaclaw start --mode rlVia APISkills + GRPO training when batch fills
madmaxmetaclaw startVia APISkills + RL + scheduler (trains only during idle/sleep/meetings)

Python API

Programmatic startup

import asyncio
from metaclaw import MetaClawAgent, AgentConfig, Mode

async def main():
    config = AgentConfig.from_yaml("~/.metaclaw/config.yaml")
    agent = MetaClawAgent(config, mode=Mode.MADMAX)
    await agent.start()

asyncio.run(main())

Manual skill injection

from metaclaw.skills import SkillStore, SkillInjector

store = SkillStore(path="~/.metaclaw/skills")

# Add a skill manually
store.add(
    name="code-review-checklist",
    content="Always check for: 1) error handling, 2) type hints, 3) docstrings.",
    tags=["code", "review"]
)

# Retrieve top-k relevant skills for a query
injector = SkillInjector(store)
relevant = injector.retrieve(query="review my Python function", top_k=3)
for skill in relevant:
    print(skill.name, skill.score)

Intercepting and recording conversations

from metaclaw.proxy import ConversationInterceptor
from metaclaw.memory import ExperienceBuffer

buffer = ExperienceBuffer(max_size=1000)

interceptor = ConversationInterceptor(
    upstream_url="https://api.moonshot.cn/v1",
    on_complete=buffer.record   # called after each turn with (messages, response)
)

# buffer.record signature:
async def on_complete(messages: list[dict], response: dict) -> None:
    ...

Triggering RL training manually

from metaclaw.training import RLTrainer, TrainingConfig

trainer = RLTrainer(
    config=TrainingConfig(
        backend="tinker",       # or "mint"
        algorithm="grpo",
        batch_size=32,
        lora_rank=16,
    )
)

# Collect a batch from the experience buffer and train
async def run_training(buffer):
    batch = buffer.sample(n=32, split="support")   # support/query separation
    result = await trainer.train(batch)
    print(f"Training complete. Loss: {result.loss:.4f}, Steps: {result.steps}")

Reward modeling

from metaclaw.rewards import RewardModel

reward_model = RewardModel(provider="llm")  # uses configured LLM for scoring

async def score_turn(prompt: str, response: str) -> float:
    score = await reward_model.score(prompt=prompt, response=response)
    return score  # float in [-1.0, 1.0]

Skills Lifecycle

Conversation turn
       │
       ▼
 SkillInjector.retrieve()   ← vector search over SkillStore
       │  injects top-k skills into system prompt
       ▼
 LLM responds
       │
       ▼
 ExperienceBuffer.record()  ← stores (context, response, metadata)
       │
       ▼ (end of session)
 SkillSummarizer.run()      ← LLM extracts reusable patterns
       │
       ▼
 SkillStore.upsert()        ← new/updated skills persisted to disk

Integration: OpenAI SDK as Client

Point any OpenAI SDK client at the MetaClaw proxy:

from openai import OpenAI

# MetaClaw proxy is running on localhost:8080
client = OpenAI(
    base_url="http://localhost:8080/v1",
    api_key="not-used-but-required-by-sdk"
)

response = client.chat.completions.create(
    model="moonshot-v1-8k",   # passed through to upstream
    messages=[
        {"role": "user", "content": "Review my pull request strategy."}
    ]
)
print(response.choices[0].message.content)

Skills are injected transparently — the client code does not change.


Scheduler (MadMax Mode)

The scheduler ensures RL weight updates never interrupt active use:

from metaclaw.scheduler import MadMaxScheduler, SchedulerConfig

scheduler = MadMaxScheduler(
    config=SchedulerConfig(
        sleep_hours=(22, 7),          # train between 22:00–07:00 local time
        idle_timeout_minutes=15,      # train after 15 min of no conversations
        google_calendar=True,         # also train during calendar meetings
        credentials_path="creds.json"
    )
)

# Check if it's safe to train right now
if await scheduler.is_training_window():
    await trainer.train(batch)

Google Calendar Setup

# 1. Enable Google Calendar API in Google Cloud Console
# 2. Download OAuth2 credentials as creds.json
# 3. Set path in config or env
export GOOGLE_CALENDAR_CREDENTIALS_PATH="/path/to/creds.json"

# 4. First run will open browser for OAuth consent
metaclaw start

Support/Query Set Separation

MetaClaw separates experience into support and query sets to prevent stale rewards from polluting updates:

from metaclaw.memory import ExperienceBuffer

buffer = ExperienceBuffer(
    max_size=2000,
    support_ratio=0.5   # 50% support, 50% query
)

# During training:
support_batch = buffer.sample(n=16, split="support")  # used to compute reward signal
query_batch   = buffer.sample(n=16, split="query")    # used for gradient update

await trainer.train_meta(support=support_batch, query=query_batch)

RL Backends

Tinker (default)

rl:
  backend: tinker
  tinker_project: my-metaclaw-project
  lora_rank: 16
  learning_rate: 1e-4

MinT

# Install MinT compatibility layer separately
pip install metaclaw-mint
rl:
  backend: mint
  mint_endpoint: https://your-mint-endpoint

Auto-detection

rl:
  backend: auto   # tries tinker first, falls back to mint, errors if neither available

Troubleshooting

Proxy not reachable after metaclaw start

  • Check port conflicts: lsof -i :8080
  • Change proxy.port in config and restart

rl mode: "No training backend available"

  • Ensure pip install -e ".[rl]" completed successfully
  • Verify METACLAW_TINKER_API_KEY or METACLAW_MINT_API_KEY is set
  • Try rl.backend: tinker explicitly instead of auto

Skills not persisting between sessions

  • Confirm skills.summarize_after_session: true in config
  • Check write permissions on ~/.metaclaw/skills/
  • Run metaclaw skills list to inspect stored skills

Madmax mode never trains

  • Verify scheduler.sleep_hours covers your timezone's night
  • Lower scheduler.idle_timeout_minutes for testing (e.g., 1)
  • Check scheduler logs: ~/.metaclaw/logs/scheduler.log

Google Calendar integration fails

  • Re-run OAuth flow: delete ~/.metaclaw/token.json and restart
  • Ensure Calendar API is enabled in your Google Cloud project

OPD teacher distillation errors

  • Only supported with rl.backend: tinker
  • Requires a separate teacher model endpoint in config:
    rl:
      opd_teacher: true
      teacher_base_url: https://api.openai.com/v1
      teacher_model: gpt-4o
    

CLI Reference

metaclaw setup                   # interactive config wizard
metaclaw start                   # start in madmax mode
metaclaw start --mode skills_only
metaclaw start --mode rl
metaclaw start --config path/to/config.yaml

metaclaw skills list             # show all stored skills
metaclaw skills delete <name>    # remove a skill
metaclaw skills export skills.json

metaclaw status                  # show proxy, scheduler, training status
metaclaw logs                    # tail all logs
metaclaw logs --component scheduler