跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

agent-lightning

Train and optimize AI agents using Microsoft's Agent Lightning framework with reinforcement learning. Use when setting up agent training, instrumenting agents with tracing, configuring LightningStore, implementing reward functions, or optimizing prompts with RL/APO algorithms.

扩展智能体511skills/agent-lightning/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/coco-research/coco/agent-lightning/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Agent Lightning

Microsoft's framework for training AI agents with reinforcement learning, automatic prompt optimization, and supervised fine-tuning.

Quick Start

Installation

pip install agentlightning

For nightly builds:

pip install --upgrade --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ --pre agentlightning

Minimal Integration (Zero Code Change)

Add agl.emit_xxx() helpers to your existing agent:

import agentlightning as agl

# Your existing agent code
def my_agent(task):
    agl.emit_input(task)  # Track input
    
    response = llm.generate(task)
    agl.emit_output(response)  # Track output
    
    reward = evaluate(response)
    agl.emit_reward(reward)  # Track reward
    
    return response

Core Concepts

Architecture Flow

Agent (your code) → agl.emit_xxx() → Spans → LightningStore → Algorithm → Updated Resources

Key Components

ComponentPurpose
LightningStoreCentral hub for traces, tasks, and resources
TracerCollects spans from agent execution
AlgorithmConsumes traces, produces improvements
TrainerOrchestrates training loop

Instrumentation

Emit Functions

import agentlightning as agl

# Basic emissions
agl.emit_input(prompt)           # Track input to agent
agl.emit_output(response)        # Track agent output
agl.emit_reward(score)           # Track reward signal
agl.emit_tool_call(name, args)   # Track tool usage
agl.emit_tool_result(result)     # Track tool results

Tracer Context

from agentlightning import Tracer

tracer = Tracer(store=store)

with tracer.trace_context(task_id="task-123"):
    # All emissions within this context are grouped
    result = agent.run(task)
    
# Retrieve trace after execution
trace = tracer.get_last_trace()

OpenTelemetry Integration

Agent Lightning integrates with OpenTelemetry:

from agentlightning.utils.otel import get_tracer

tracer = get_tracer()  # Returns OTel tracer for "agentlightning"

LightningStore

In-Memory Store (Development)

from agentlightning.store.memory import InMemoryLightningStore

store = InMemoryLightningStore()

Client-Server Store (Production)

from agentlightning.store.client_server import (
    LightningStoreServer,
    LightningStoreClient
)

# Server side
server = LightningStoreServer(store, host="0.0.0.0", port=8080)
await server.start()

# Client side
client = LightningStoreClient("http://localhost:8080")

Store Operations

# Add rollouts (tasks for the agent)
await store.enqueue_rollout(task=task, config=RolloutConfig())

# Query rollouts
rollouts = await store.query_rollouts(status_in=["completed"])

# Add resources (updated prompts, weights)
await store.add_resources(resources)

# Get latest resources
resources = await store.get_latest_resources()

Training

Basic Trainer Setup

import agentlightning as agl

trainer = agl.Trainer(
    n_runners=8,           # Parallel rollout workers
    algorithm=algorithm,   # Your chosen algorithm
    store=store           # Optional, creates InMemory if not provided
)

trainer.run()

Custom Algorithm

from agentlightning import LightningStore
from agentlightning.types import ExecutionEvent

async def my_algorithm(store: LightningStore, event: ExecutionEvent):
    # Fetch completed rollouts
    rollouts = await store.query_rollouts(status_in=["completed"])
    
    # Process traces, compute gradients, etc.
    new_resources = optimize(rollouts)
    
    # Push updated resources
    await store.add_resources(new_resources)

Runner Function

async def my_runner(store: LightningStore, worker_id: int, event: ExecutionEvent):
    while not event.is_set():
        rollout = await store.dequeue_rollout()
        if rollout:
            result = execute_task(rollout.task)
            await store.update_rollout(
                rollout_id=rollout.id,
                status="completed",
                result=result
            )

Algorithms

Reinforcement Learning (GRPO/PPO)

For RL training with vLLM backend:

from agentlightning.algorithm.verl import VeRLAlgorithm

algorithm = VeRLAlgorithm(
    model="your-model",
    learning_rate=1e-5,
    batch_size=32
)

Automatic Prompt Optimization (APO)

from agentlightning.algorithm.apo import APOAlgorithm

algorithm = APOAlgorithm(
    optimizer_model="gpt-4",
    target_model="gpt-3.5-turbo"
)

Framework Adapters

LangChain

from agentlightning.instrumentation.langchain import instrument_langchain

instrument_langchain()  # Auto-traces all LangChain calls

OpenAI SDK

from agentlightning.instrumentation.openai import instrument_openai

instrument_openai()  # Auto-traces OpenAI API calls

vLLM

from agentlightning.instrumentation.vllm import instrument_vllm

instrument_vllm()  # Instrument vLLM for token-level tracing

Logging & Debugging

Configure Logging

from agentlightning import setup_logging

setup_logging(
    level="DEBUG",
    submodule_levels={
        "agentlightning.store": "INFO",
        "agentlightning.tracer": "DEBUG"
    }
)

Metrics

Agent Lightning emits Prometheus-compatible metrics:

  • agl.store.total - Store operation counts
  • agl.store.latency - Store operation latencies
  • agl.rollouts.total - Rollout counts by status
  • agl.rollouts.duration - Rollout execution times

Common Patterns

Reward Function Design

def compute_reward(task, response):
    """Good rewards are: normalized, dense when possible, aligned with goals."""
    
    correctness = check_correctness(task, response)  # 0-1
    efficiency = measure_efficiency(response)         # 0-1
    
    return 0.7 * correctness + 0.3 * efficiency

Multi-Agent Training

Train specific agents in a multi-agent system:

with tracer.trace_context(agent_id="planner"):
    plan = planner.run(task)

with tracer.trace_context(agent_id="executor"):
    result = executor.run(plan)
    
# Only the executor's traces are used for training

Checkpoint & Resume

# Save checkpoint
await store.add_resources(
    checkpoint=True,
    resources=current_resources
)

# Load latest
resources = await store.get_latest_resources()

Integration with JavaScript Agents

For JavaScript/TypeScript agents (like Claude-based apps), you have two options:

Option 1: Python Training Service

Create a Python microservice that:

  1. Receives trace events from your JS app via HTTP
  2. Stores them in LightningStore
  3. Runs training algorithms
  4. Returns optimized prompts

Option 2: REST API Integration

Use LightningStoreServer as a REST backend:

// JavaScript client
const response = await fetch('http://localhost:8080/rollouts', {
    method: 'POST',
    body: JSON.stringify({
        task: { prompt: userMessage },
        config: { max_retries: 3 }
    })
});

Resources

Troubleshooting

IssueSolution
Import errorsEnsure pip install agentlightning succeeded
Store connection failedCheck server is running, verify endpoint URL
No traces collectedVerify emit_xxx() calls are within trace context
Training not convergingCheck reward function normalization, increase rollouts

相似的 Skill

skill-creator
anthropics/skills180k

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

扩展智能体

template-skill
anthropics/skills180k

template-skill

Replace with description of the skill and when Claude should use it.

扩展智能体

mcp-builder
anthropics/skills180k

mcp-builder

Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).

扩展智能体

claude-api
anthropics/skills180k

claude-api

Reference for the Claude API / Anthropic SDK — model ids, pricing, params, streaming, tool use, MCP, agents, caching, token counting, model migration. TRIGGER — read BEFORE opening the target file; don't skip because it "looks like a one-liner" — whenever: the prompt names Claude/Anthropic in any form (Claude, Anthropic, Fable, Opus, Sonnet, Haiku, `anthropic`, `@anthropic-ai`, `claude-*`, `us.anthropic.*`, `[1m]`); the user asks about an LLM (pricing/model choice/limits/caching) — never answer from memory; OR the task is LLM-shaped with provider unstated (agent/MCP/tool-definition/multi-agent/RAG/LLM-judge/computer-use; generate/summarize/extract/classify/rewrite/converse over NL; debugging refusals/cutoffs/streaming/tool-calls/tokens). SKIP only when another provider is being worked on (overrides all triggers): OpenAI/GPT/Gemini/Llama/Mistral/Cohere/Ollama named in the query; OR `grep -rE 'openai|langchain_openai|google.generativeai|genai|mistralai|cohere|ollama'` over the project hits (run this grep FIRST if no provider named — don't Read the file).

扩展智能体

academy-guide
anthropics/skills180k

academy-guide

Stop and check this skill before finishing any reply to a question about how to use Claude or a Claude product — it recommends matching courses, tutorials, and use cases from Claude Academy (academy.claude.com), Anthropic's learning hub. Trigger on: "how do I", "how can I", "getting started with", "what can Claude do", "teach me", "learn to use"; questions about artifacts, projects, skills, plugins, connectors, MCP; requests about rolling Claude out to a team, class, or organization; and any ask for training materials, onboarding content, or learning resources. Use it when the user is learning how to use a feature or product — not when they are mid-task and just want the task done. This skill composes with other skills: after consulting product documentation to answer how a Claude feature works, also check here for a matching course or tutorial — a docs-grounded answer and an Academy recommendation belong together. Only recommend on a strong match; never invent Academy content.

扩展智能体

ponytail-help
DietrichGebert/ponytail160k

ponytail-help

Quick reference for ponytail levels, skills and commands. One-shot display. Use for /ponytail-help, "ponytail help", "how do I use ponytail".

扩展智能体