Multi-agent LLM orchestration for Python — Planner → Executor → Critic, with tool-calling, evals, observability, and a keyless LLM backend.
A task goes in; a Planner decomposes it into steps, an Executor runs each one (calling tools where needed), and a Critic judges the result. Ships an LLM-as-judge eval harness, OpenTelemetry spans, and a keyless provider chain — build and benchmark agents with no API key.
@tool decoratorllm_judge assertionsgit clone https://github.com/chirag127/agent-forge
cd agent-forge
pip install -e .
agent-forge run "What is the square root of 2025?"
agent-forge eval
Requires Python 3.13. PyPI publish is on the roadmap.