Platform tour
The complete agentic
engineering platform.
Point tools give you an agent. AnyForge gives you the whole engineering organization around the agent: a planning layer that turns demand into scored, spec-driven work; a delivery layer that ships it through governed crews; a governance layer with cryptographic audit and human approval gates; and an insight layer that tells leadership what actually happened: all on any model, from frontier clouds to private LLMs hosted on your own metal.
From raw customer signal to merged, audited pull request: one platform, one ledger, one flat fee.
Plan
From raw demand to a runnable spec.
Most AI coding tools start at the prompt. AnyForge starts where your business does: it takes the customer ask, the strategy bet, or the production incident and shapes it into prioritized, spec-verified work an AI crew can actually run.
Signals: capture demand where it lands
Client asks, sales requests, stakeholder ideas and strategy bets land as Signals, the cheapest thing to write down. Nothing lives in an inbox; every signal keeps its source so you always know what is customer-driven vs strategy-driven.
Opportunities with RICE scoring
Signals consolidate into scored Opportunities: Reach × Impact × Confidence ÷ Effort, server-computed from anchored buckets with mandatory rationale on every input. AI Suggest drafts scores grounded in your actual signals. No gut-feel prioritization.
Revenue at Risk
The backlog translated into money: monthly MRR blocked by unshipped work, annual run-rate exposure, concentration by region, account and reason. The exec-facing lens that turns a prioritization debate into a business decision.
Roadmap themes with investment guardrails
Now / Next / Later horizons plus a quarter-gridded timeline. Set an investment mix: say 60% features, 25% tech debt, 15% keep-the-lights-on, and the roadmap tells you when reality drifts from the strategy.
Capacity forecasting (Monte Carlo)
AnyForge models the real bottleneck of agentic delivery: operator review attention, not engineer hours. Little’s-Law lead times, P50/P85/P95 Monte-Carlo forecasts over your trailing eight weeks, and a what-if simulator for adding operators or raising autonomy.
Spec-driven development (ICE)
Every initiative carries a structured spec with Intent, Context, and Expectations. Every program carries a structured, ICE-scored brief. Both are hard-gated at plan approval: the architect's plan can't be approved until ICE scores 100% (genuinely simple work waives optional sections per-field; an owner-only override is audit-logged). The spec becomes the contract every agent reads and the checklist the PR is verified against, including "verify NOT introduced" non-goals, and briefs and specs take Google-Docs-style margin comments as you refine them. At the ship gate the spec decides refine-vs-ship: unmet acceptance criteria earn another engineer round, an open critical finding never rides through, and residual in-scope findings file as a single follow-up work item instead of burning refine rounds or blocking the merge.
Programs → Initiatives → Work items
A four-level hierarchy that matches how software actually ships: cross-codebase programs, per-repo initiatives, PR-sized work items with dependency waves. Initiatives can be bounded: delivered once and validated to a terminal close, or long-lived, staying open to absorb an ongoing stream of work. Mark one Blocked on an external dependency and it pauses without losing delivery state, rolling up into a blocked-by-program report; a cross-initiative dependency map shows what blocks what and lets you dispatch a crew straight from a node. Program cockpits roll up progress, spend, risks and blockers: exportable to PDF or Markdown for the board.
Business Capability Map & Value Chain
A BIZBOK-style capability tree synthesized from your own codebases, showing which repositories realize Payments, Risk, and Onboarding, with heatmap badges exposing coverage gaps and duplication. One click projects it onto Porter's Value Chain: primary activities inbound→operations→outbound→marketing/sales→service, with support activities beneath. Both views carry the same coverage badges, so leadership reads your estate on a frame it already knows. Both export to Markdown, PNG, or PDF.
Community board: demand from the people who use it
A public request board where users file bugs, features, or enhancements and attach screenshots as evidence. Oracle-guided ICE checks keep incomplete requests from being submitted. Users can browse the anonymized Explore feed, upvote requests, and track them from Received to Shipped. The delivery pulse shows open requests, active crews, recent shipments, and median time to ship. Each request becomes a work item for a crew.
Build
Governed delivery that ships real PRs.
One capable Engineer, specialist advisors on demand, deterministic edits, and your test suite as ground truth. That is the delivery architecture the failure data says actually works, running in your browser or as unattended background crews.
One Engineer, on-demand specialists
Crew keeps execution with one Engineer and uses Architect, Security & Compliance, and QA specialists as advisors during the run. This avoids handoffs between multiple executing agents.
Deterministic edits, real tests
Patches apply through git. Each patch is checked before it is applied, never fuzzy-matched, so a stale diff is rejected loudly instead of corrupting code silently. Your own test suite is the ground-truth oracle: the Engineer detects your runner and runs your tests, not a proxy metric.
Parallel specialist review on every PR
When the PR opens, Architect, Security and QA specialists fan out in parallel and review the actual diff: OWASP-class security findings, coverage analysis, architectural conformance, and attach their reports to the approval card, with overlapping findings consolidated into one evidence-backed item per underlying issue. Amend a PR and your feedback rides through to the specialists on re-review, so the second look answers what you asked.
A fifty-tool agent arsenal
AnyForge agents don’t just edit files. They run your terminal, search the live web, spawn read-only research sub-agents with isolated budgets, enter plan mode, work in isolated git worktrees, search structurally by AST as well as semantically, submit line-anchored PR reviews, and verify UI changes by starting a worktree-scoped dev server, driving localhost, screenshotting the page and reading its console and network errors, with the proof attached to the PR.
Integrations verified against the real thing
Crews build and verify real third-party API, payments, and webhook integrations in a live sandbox without ever holding your provider credentials or egress. A typed integration test returns real status, artifacts, and webhook confirmation; the crew can trigger the flow itself under an egress allow-list with credentials injected server-side, and confirm the async webhook through a connected observability tool like Datadog. The provider’s real response and webhook evidence land on the PR-review card, and a failed check holds Approve behind an explicit override.
Cross-repo context, single-repo blast radius
Dispatch a crew with sibling repositories mounted read-only, including a shared library, the upstream API producer, or your design system. The Engineer and specialists can read real cross-repo code, contracts, and patterns while every edit and the PR stay confined to the one writable target repo, enforced at the tool layer. Need to move code between repos? A crew migrates source paths into a target monorepo by copying them, preserving history, or creating a subtree as governed work like any other.
Code Studio: governed coding in the browser
A full agentic workspace on any device: Files, Diff, Terminal and Metrics tabs, 200+ models switchable per request, repo-grounded plans, surgical edits, main-branch locking. Commands run in an in-browser container or fail over to our cloud runtime: npm test works from your phone. The agent can connect GitHub repos and start their analysis from the conversation, and builds UI to your organization’s design system.
One-click escalation to Crew
When a task outgrows a single agent, Code Studio proposes a Crew inline. Accept once and the initiative, work items and crew are created atomically, and you stay on the same conversation, approving gates in-thread.
Automations & routines
Saved crew templates fire on cron schedules or signed webhooks (HMAC-SHA256) for nightly security audits, weekly tech-debt refactors, and dependency sweeps. Read-only audit crews reach a real finish line with a committed findings document and a PR instead of trailing off, and every routine card links straight to the crews it dispatched. Recurring engineering runs itself; you review the PRs.
Playbooks: an operator workflow in one click
A library of bundled, ready-to-run Operator Workflows, including the managed Production Watchdog, Cost Anomaly Watch, Release Health Report, and Security Findings Triage. Browse the gallery, enable one, and it clones into your workspace for customization. A live Enabled badge appears once it runs. Playbooks reach real production state through an ask-the-Oracle step, so a workflow can find out what’s actually happening before it acts. Operator Workflows themselves add loop steps and resumable budgets you can top up mid-run, with every spec edit versioned and rollback-able.
Self-healing releases
Your pipeline reports a release failure, AnyForge opens a P0 hotfix crew automatically and files the incident. If you opt in, it can also auto-merge hotfix PRs that fit your safe-diff envelope (small, safe-path-only changes). Change-failure rate and lead time are tracked honestly, DORA-style.
Automatic merge-conflict resolution
A sweep every 30 minutes detects cross-work-item conflicts before they block a release, attributes them to the sibling branches that caused them, and resolves them with a crew that merges (never rebases), runs your tests and updates the PR in place.
Mission control for every run
Watch a crew work in a live activity feed with a phase tracker: Engineer, Specialist Review, PR Review, and typed interaction cards for everything that needs you: plan approvals, agent questions, task proposals, backlog refinements, ADR revisions. Mention @architect in the thread, get desktop notifications when a gate needs a decision.
A review room, not a diff dump
Every crew PR opens in a three-pane review cockpit: an AI-generated walkthrough that tours the diff in the order it should be read, syntax-highlighted diffs with click-to-comment on any line (posted to GitHub and in-app, @-mentions included), and the agent chat with approve, amend and auto-merge right where you’re reading.
Workspace bootstrap that respects your repo
The agent runtime ships every major toolchain: Node, Python, Go, Rust, Java, Kotlin, .NET, Ruby, PHP: plus cloud CLIs. Pin a setup command, a test command, and your AGENTS.md conventions; every crew and Studio session honors them.
Code intelligence
The graph index: your codebase as a queryable brain.
Underneath every AnyForge surface sits one code-intelligence layer: a continuously refreshed call graph plus semantic embeddings of every connected repository. It's why AnyForge agents act like engineers who've read the codebase, and why your questions get answers with line numbers.
A graph of everything you ever shipped
AnyForge builds a call-and-import graph of every connected codebase using the TypeScript compiler API for TS/JS and tree-sitter parsing for Go, Kotlin, Python, Java, and Rust. Agents and humans can navigate your systems by structure, not by grep.
Semantic search over symbols, not strings
Vector embeddings of symbol-aligned code chunks power natural-language search across the fleet. Ask "where do we validate webhook signatures?" and find the function, its callers, and the docs that mention it across repositories.
Always fresh, never a batch job
Indexing is SHA-gated and incremental: it refreshes on every PR merge via webhook, on operator demand, and on a rolling sweep. The graph you query reflects the code you merged this morning, not last quarter.
Edges that admit what they know
Every call edge carries a confidence label: extracted (import-verified), inferred (unique name match), or ambiguous. Agents know when to double-check before acting on it. Honest metadata beats confident guessing.
Analysis on every pass
Each index run detects god nodes, module communities, surprising cross-boundary connections, and doc-to-code drift, then generates suggested questions worth asking about your own architecture. Incremental runs report the graph delta: what structurally changed since last time.
One index, every consumer
The same graph powers the Engineer mid-build, the PR-review specialists, the Oracle, semantic ⌘K search, and the interactive Code Graph views: per-repo and cross-repo. Ask, click, or let an agent traverse it.
Documentation engine
Documentation that regenerates itself.
Every codebase you connect gets living architecture documentation: generated from the code, refreshed on merge, and organized the way architects and boards actually think. The wiki that rotted the day it was written is over.
Deep analysis on connect, fresh on every merge
Point AnyForge at a repository and it reverse-engineers the architecture: components, boundaries, data flows, build tooling, test posture. Then the same merged-PR event that reindexes the code incrementally patches the affected docs: changed files mapped to the doc types they touch, each edit versioned and audited, so the documentation tracks HEAD instead of drifting between scheduled reanalyses. Freshness is a policy, not a hope.
C4 diagrams & ADRs, generated
Context, container and component views drawn from the code itself, plus Architectural Decision Records drafted by the architect agents as they work: the documentation trail writes itself while the work happens.
Deployment topology & dependency maps
A live L1 deployment topology with component-to-component dependency arrows and relationship overlays, plus cross-repo dependency maps: see what actually talks to what, across the whole estate.
Opinionated architecture lenses
Conway analysis (does your org chart match your architecture?), an AI-readiness score per codebase (how well can agents work here?), and a security-posture view: the assessments consultants charge six figures for, regenerated on merge.
The org narrative
A background worker synthesizes a business-readable overview of your entire estate: every analyzed system, woven into one wiki-home narrative your board can actually read, with live generation progress and ETA.
Searchable, shareable, exportable
Full-text search across all generated documentation, external share links for auditors and diligence teams, and Markdown/PNG/PDF exports. Due diligence prep becomes a link, not a fire drill.
AnyForge Oracle
The one teammate who has read everything.
The Oracle is a 360° intelligence layer over your entire engineering reality: every codebase, every architecture doc, every release, every crew: correlated live with production observability data from Datadog, Google Cloud Logging, AWS CloudWatch and Sentry. It's the difference between asking 'what broke?' and being told 'this release, this code path, this fix: shall I dispatch a crew?'
A 360° view of every codebase
The Oracle sits on top of the full code-intelligence layer, including the cross-repo call graph, semantic code search, generated architecture docs, ADRs, the business capability map, releases, and dependency maps. Ask "what talks to the payments service and who changed it last?" and get an answer with clickable deep links, not a guess.
Correlated with production reality
It reads your observability stack live: Datadog, Sentry, Google Cloud Logging and Crashlytics, AWS CloudWatch, Aikido, and correlates a production error with the code path that throws it, the release that shipped it, and the PR that introduced it. From stack trace to root cause to a dispatched fix crew, in one conversation.
Real computation, not vibes
A sandboxed, credential-isolated run_code interpreter (Python/Node) lets the Oracle parse, aggregate, calculate and chart inside the conversation, so the answer to "what did this initiative really cost per merged PR?" is computed, not estimated.
For every member, on every page
Summonable anywhere in the console (⌘-J) by every member including viewers, the Oracle is read-only by construction and never budget-gated in the console. It keeps answering even when a token budget pauses delivery. Runs on your own AI key.
Answers that leave the chat
Every answer exports to Markdown, JSON, CSV, HTML, or DOCX. Per-user memory carries your context between sessions, and ⌘K global search matches your generated architecture docs by title and type instantly, including runbooks, data models, and threat models on top of a semantic deep tier over the indexed code.
Decision support right at the gate
Hit a HIL approval or a halted run and "Consult the Oracle" opens it pinned to that crew, hydrated server-side with the run, the pending gate, the reviewer verdicts and the budget, so you get merge-safety, blast-radius and halt-diagnosis advice exactly where the decision lives. It advises; it never rubber-stamps.
A confirm-first action tier for operators
Read-only by default for everyone, but operators can also ask the Oracle to act on what it finds. Each step is confirmed first and role-checked server-side, whether it dispatches, amends, halts, or resumes a crew, approves a gate, dispatches the architect, closes a finished run as delivered, or connects GitHub repos and kicks off their analysis across a whole org. The read-only assistant becomes a doer only when you say yes.
Ask about a person, not just a project
"What did I ship last week?" "What’s blocking Priya?" The Oracle now answers per-operator with what someone launched, delivered, and which gate is holding them up, grounded in the record. Self-reporting is open to everyone; colleague queries are gated to owners and operators. It reports what happened, and says plainly that intent has to come from the person.
Everywhere your team already is
The same Oracle answers from Slack via the /oracle slash command, from Claude Code or Cursor via the anyforge_ask_oracle MCP tool, and over the headless Control API: one brain, every doorway.
The agents
Not a bot. A staff.
Every role a scaling engineering org hires for, AnyForge ships as a specialist agent: each one grounded in your workspace, metered on your keys, and accountable to the same audit chain. The Oracle answers; these agents act.
Opportunity Genesis: from signal to PRD
Feed it raw demand and it drafts the product case: versioned PRDs grounded in your actual signals, complete with browsable HTML UI mockups your stakeholders can click through before a line of code exists.
Program Genesis: the portfolio planner
Analyzes your repositories, proposes program scope, creates the initiatives, writes versioned program briefs, maintains the risk register, and generates status reports. It automates the program-management overhead.
Spec Genesis: the spec author
Drafts and refines ICE specs in conversation, dispatches the architect for decomposition, and requests backlog refinement, so every initiative reaches the crew with a contract worth building against.
Report Genesis: dashboards you talk into existence
The reporting module lets you describe what you want to see and assembles the dashboard: stat tiles, timeseries, bars, and tables over crews, work items, initiatives, token usage, and codebases. Queries are validated against a safe catalog and executed live at view time. Finished reports live under Insights → Reports for the whole org, share via a company-scoped deep link, and export to PDF or Markdown.
The delivery crew & Retro Analyst
The One Engineer and its Architect, Security & Compliance, QA and Sage specialists do the building, and when a crew completes, an operator-triggered Retrospective Analyst reads the step trail and audit log, writes the retro, and proposes new memory facts for humans to promote.
Govern
Autonomy you can hand an auditor.
Speed without governance is liability. Every AnyForge surface shares the same enforcement machinery: approval gates, policy engine, constraint checks and a tamper-evident audit chain, so moving faster never means explaining less.
One human gate. Cryptographic proof.
Every normal crew run ends at a mandatory human approval gate at PR review: unbypassable by the model. Each decision is recorded with a SHA-256 proof of approver, timestamp and content. The gate is enforced server-side on every path: the chat card, the control API and the Approvals queue all re-read the PR’s live GitHub checks and refuse to land an approval while CI is red, pending or the branch is in conflict; the CI verdict at approval is written to the audit trail, and shipping a failing check takes a conscious, recorded override.
Coverage as a gate, not a vibe
Turn on the enforced coverage gate per codebase and a crew whose measured PR-head coverage falls below your threshold holds the ship instead of presenting an approval card: a "> N% coverage" acceptance criterion can no longer pass on a self-attested number. Fail-open and off by default. When a merge is blocked by branch protection or required reviewers, the work item and initiative stay flagged until the change actually lands.
Hash-chained audit trail
Every governance event, including approvals, constraint checks, crew lifecycle, and access, is SHA-256 hashed and chained to the previous event in an append-only ledger that application code cannot write. Tampering with any historical event breaks verification of everything after it. Regulator-ready, exportable.
Enforced constraints (Atomic Facts)
Your engineering rules, including TLS requirements, API contracts, schema invariants, and compliance mandates, are injected into every agent’s prompt and verified per-constraint at PR time. PCI-DSS and SOC 2 profiles seed a working posture on day one. Facts are never deleted, only superseded.
Policy as code (OPA)
Budget caps, circuit breakers and per-role tool allow-lists run as Rego policies distributed via OPAL: the same policy engine your platform team already trusts, applied to AI agents. Redaction, PII blocking and rate limits toggle per rule.
Portable governance memory
Architecture decisions, API contracts and declined proposals persist across runs, sessions, and provider switches: policy-checked, fail-closed, tenant-isolated, every write signed into the audit chain. Your agents stop re-litigating decisions your team already made.
Autonomy is a dial, not a switch
Turn up unattended operation one envelope at a time: CI-failure triage, then auto-fix with hard caps (bounded rounds and tokens), then auto-merge for small fixes only (line and file limits you set), then Release Health auto-heal inside a safe-path allow-list with capped attempts. Auto-merge defaults cascade org → program → initiative → individual gate, with an explicit opt-out at every level. Everything above the envelope waits for a human.
Config Packs: compliance in one click
Installable bundles of agent configs, enforced constraints, routines and skills that stand up a working governance posture instantly: PCI-DSS, SOC 2, HIPAA, HIPAA + SOC 2 cross-mapped for healthcare SaaS, NIST AI-RMF, NIST CSF, NIST SSDF, EU CCD2, source-code escrow, and a startup pack. Your whole company configuration also exports and imports as a manifest: template one company, stamp out the next.
Team & access management
Teams, invitations, multi-domain email auto-join (claim several email domains and employees on any of them join the same org), and a role model that goes down to surface-locked board viewers. Every membership change and permission decision lands in the same audit chain as the code.
A deliberate trust boundary
AnyForge produces and merges pull requests. It never deploys and never holds your infrastructure keys. Secrets live in Google Secret Manager, while the database stores only pointers. Role-based access includes surface-locked board viewers, invited-email allowlisting, and pending-approval onboarding.
Config Packs
A governance posture, installed in one click.
A Config Pack is an installable bundle of platform configuration: agent configs, enforced constraints (Atomic Facts), routines, and skills. It stands up a working posture in one action instead of an afternoon of manual setup. Install one from the marketplace and the rules are live immediately: injected into every agent's prompt, verified per-constraint at PR time, and audited like everything else. These are the packs we ship today.
Compliance
PCI-DSS
Cardholder-data guardrails composed into every agent prompt — encryption, TLS minimums, access logging — plus a weekly routine that points the Compliance/Security agent at recent commits to catch violations before an assessor does.
Compliance
SOC 2 Type II
The trust-services common criteria as enforced constraints: logical access, change management, monitoring, and backup-recovery testing. A quarterly routine audits access-change logs and surfaces anything that would fail the audit.
Compliance
HIPAA Security & Privacy
Security Rule technical and administrative safeguards for anything touching ePHI — encryption at rest and in transit, unique user identification, audit controls, automatic logoff, minimum necessary, contingency-plan testing — plus the business-associate requirement and the Breach Notification Rule’s 60-day clock. Monthly §-cited safeguard report.
Compliance
HIPAA + SOC 2 (Healthcare SaaS)
One cross-mapped control set for a healthcare SaaS carrying both obligations: every constraint cites both its CFR section and the SOC 2 criterion it satisfies, so a single piece of evidence answers both auditors. Install this instead of the two standalone packs.
Compliance
NIST Cybersecurity Framework 2.0
The six core functions — Govern, Identify, Protect, Detect, Respond, Recover — expressed as engineering constraints: asset inventory and SBOM, least privilege with MFA and encryption, security logging and detection, incident-response hooks, tested backup and recovery. Quarterly posture review.
Compliance
NIST SSDF (SP 800-218)
The secure-development practice set behind US Executive Order 14028 and the federal secure-software attestation: code integrity and provenance, pinned well-secured dependencies, static review before merge, security testing, threat modelling, timely remediation — with review prompts that cite practice IDs.
Compliance
NIST AI RMF 1.0
For crews building AI features: model and system cards, context and risk mapping, accuracy / safety / bias evaluation with real metrics, documented human oversight, and hardening against prompt injection, data poisoning and model exfiltration. Quarterly AI-risk review.
Compliance
EU Consumer Credit Directive 2
Directive (EU) 2023/2225, fully applicable 20 November 2026, for BNPL and consumer-credit software: the expanded scope, the creditworthiness assessment, the right to human intervention on automated credit decisions, SECCI pre-contractual disclosure, the 14-day withdrawal right, and the tying and unsolicited-credit bans. Engineering guidance, not legal advice.
Compliance
Software Escrow
For licensors under a source-escrow agreement: deposits must be buildable and operable by a skilled third party, a current third-party dependency register must exist, and quarterly refresh plus annual audit routines mirror a standard obligations schedule. Pairs with the Escrow Bundle generator on the codebase page.
Velocity
Startup
The other direction. Loosened rails for teams optimising for velocity: feature flags preferred over branch protection, tests required on shipped paid paths, a daily PR digest instead of a heavy compliance audit, simpler architectures and cheaper default models.
A pack is a starting posture, not a lock. Everything it installs lands as normal, editable, operator-owned records, and packs stack: install several frameworks and their constraints accumulate, each tagged with the pack it came from. Your whole company configuration exports and imports as a manifest too, so a proven setup templates across companies: configure one portfolio company, stamp out the next.
See
Every question answered, from board room to terminal.
What shipped? What did it cost? What's at risk? Who approved it? AnyForge answers in whatever language the asker speaks: a board pack for directors, a trace waterfall for engineers, or a conversation with the Oracle for everyone in between.
Management Overview & board packs
A deterministic leadership lens: what shipped, what was decided, committed vs forecast dates with variance history, cost, and an honest change-failure rate. One click prints a board pack. A weekly brief email keeps execs current without another meeting.
Usage, traces & spend
Month-to-date token spend and projected month-end by user, initiative, provider, model, repo, and product surface. Full LLM trace waterfalls with three lenses: recent, highest-spend, and errors, filterable by source (Crew, Code Studio, or external agent), plus opt-in run-tree tracing that nests every LLM round under its graph node.
Operator effectiveness scores
A composite 0–100 score per operator covers outcomes, cost efficiency, spec quality, responsiveness, and context reuse, with coaching notes. The management question "who runs agents well?" finally has data behind it.
Two-ledger spend reconciliation
AnyForge’s metered ledger is reconciled daily against each provider’s own usage API: Anthropic, OpenAI, Google, and OpenRouter. Leakage is surfaced per provider, model, and day. Trust the meter because it audits itself.
Platform health & incidents
A live health dashboard with 30-second refresh, smoke-test heatmaps and halt taxonomies. Production errors from any provider are fingerprint-grouped into signals; accept one and it becomes a tracked KTLO incident a crew can fix.
Production Watchdog
An opt-in managed agent stands watch over production. On a cadence and reactively on new signals, it diffs your ingested observability signals to surface what is newly wrong: new errors, occurrence spikes, and worsening severity. The Oracle explains each one against your connected observability tools, with numbers copied verbatim. It delivers a digest to your inbox, Slack, and email, writes a run record every pass so you can prove it ran, and, if you allow it, auto-opens KTLO incidents for the critical ones. Turn it on from the Playbooks gallery.
Release notes that write themselves
Push a version tag and AnyForge generates redacted, customer-safe release notes covering what is new, improved, fixed, or scrubbed of secrets. It publishes them to your release timeline and changelog, and announces them to Discord. Shipped user-reported items announce themselves back to the community that asked.
Delivery analytics & DORA
Cost per completed crew as the headline number, plus deployment frequency, lead time (median and p90), change-failure rate, MTTR, first-pass acceptance-criteria rate, amend rounds per initiative, and token forecast-vs-actual: pivotable by operator, initiative, codebase, program or team. A correlation view empirically ties review-override decisions to later failures. Cloud spend syncs in too, with per-service and per-SKU GCP drill-downs.
Notifications that find you
Four channels: in-app inbox, email, Slack DM, and Discord, with an @mention inbox, per-event preferences, and approval alerts that reach the right operator with severity built in. Slack slash commands query the Oracle or file a signal without opening the console.
Extend
Plugged into everything you already run.
AnyForge meets your stack where it is: your observability, your ticketing, your CI, your editor, your compliance obligations.
Every MCP server. Not a curated few.
AnyForge speaks the Model Context Protocol natively, so any MCP server works. Connect it over Streamable HTTP or SSE with OAuth (auto-refreshing), Bearer tokens, API keys, or no auth at all, scoped per-user or company-wide. A one-click marketplace covers the popular ones: GitHub, Datadog, Sentry, Jira, Confluence, Linear, Notion, Salesforce, Stripe, Snowflake, Cloudflare, CloudZero, AWS Cost & Billing, Google Drive & Calendar, Microsoft 365, Firebase, Aikido, and everything else plugs in the same way. A mid-session credential lapse self-heals on the next call, or hands you a one-click reconnect link. Connected tools flow to the Oracle, the Genesis agents, and, with permission, your crews. Read tools are available to everyone; mutating tools are operator-gated.
Agents that build to your design system
AnyForge reads a repository’s root DESIGN.md as the visual source of truth for every agent, using machine-readable tokens and guardrails. It falls back to an org-wide default you manage in the console and pulls in managed brand assets like logos and imagery. Code Studio, the Genesis UI-mockup generators, and crews all build UI to your real identity, validated on write, so agent output looks like your product instead of a generic template.
Skills marketplace & curated registries
Reusable prompt packs capture your conventions, governance rules, and output contracts. They are composed into every agent and scoped by toolchain or codebase. Install from curated registries pinned to immutable commits, import SKILL.md files, or let AnyForge auto-detect skills committed in your repos. Every crew arrives pre-loaded with a bundled library of AnyForge engineering-discipline skills (planning, code review, security and compliance scanning, deploy, and QA) advertised by role and language.
Drive AnyForge from your editor
The AnyForge MCP server puts ~50 tools in Claude Code, Cursor, or any MCP client. Approve HIL gates, dispatch crews, query the Oracle, manage constraints, sync coverage, and drive escrow without leaving the terminal. Scoped read / write / approve keys enforce least privilege.
A full HTTP control surface
Everything the console does, an API does: repos, crews, initiatives, programs, work items, approvals, coverage, escrow, observability signals. SDK-compatible LLM endpoints mean your existing OpenAI or Anthropic client code works by changing one base URL.
Coverage & quality ingestion
Pipe the coverage your CI already computes from Jest, pytest, Go, LCOV, SonarQube, Codecov, or Coveralls into per-merge program and initiative metrics. Coverage guides scope; your test suite remains the gate.
Source-code escrow, verified
Escrow-grade deposits include SHA-256 manifests, build-and-deploy runbooks, dependency registers, and masked secret scans. They are delivered over pinned SFTP to agents like Escode/NCC during quarterly sweeps across your fleet. A Replicate rehearsal proves each deposit actually rebuilds.
Connectors & integrations
Any model
Frontier clouds, open weights, or your own metal.
Model strategy is a business decision, not a vendor default. AnyForge is radically provider-neutral: bring your own keys, route every call to the model that earns it, and change your mind later without losing memory, audit history, or a single workflow.
Every frontier cloud
Anthropic Claude (including Fable 5, Sonnet, Opus, and Haiku), OpenAI GPT and o-series, and Google Gemini, plus AWS Bedrock and Google Vertex for teams that buy through their cloud. Live model catalogs per provider, automatic failover, one endpoint.
Open-weights, first-class
GLM-4.6 and GLM-4.5 Air, DeepSeek V3 and R1, Qwen Coder, Llama, Mistral, and Grok are selectable anywhere a frontier model is used: for agent roles, Code Studio, and routing targets. Use 200+ models through OpenRouter, or route directly to the vendor to skip the middleman.
Custom & private LLM hosting
Run models on infrastructure you control: your own Ollama daemon, a vLLM or llama.cpp cluster, or a dedicated private gateway we operate for you, on-prem, in your VPC, or air-gapped. You get the same governance, audit chain, and flat platform fee as cloud. Built for data-residency clauses, regulated industries, and defence.
Claude Max subscription routing
Already paying for Claude Max? Paste an OAuth token and route agent traffic through your flat subscription instead of metered API rates, with a self-pacing throttle against Anthropic’s utilization windows, extended prompt-cache TTL, and automatic fallback to an API key.
Smart routing & caching
Intent-based routing uses rules, regex, or embeddings to send each call to the cheapest model that clears the quality bar, with price-intelligence recommendations per agent role. Prompt caching typically shaves 30–40% off model spend. Most teams save more than the platform fee costs.
Benchmark two models on the same work
Stop guessing which model is worth it. Run one work item twice as two isolated arms on their own branches, each with its own per-role run profile. A side-by-side report scores tests, open findings, duration, tokens, and cost, and a deterministic recommendation ranks them by outcome, then tests, findings, and price. The recommendation is advisory; you pick the output to take forward, review its PR the normal way, and retry a failed arm in one click. Neither arm can auto-merge.
Dial reasoning effort, per model
Set a reasoning level, low, medium, high, xhigh, or max, on any agent. AnyForge translates it into whatever the underlying model understands: Anthropic thinking effort, OpenAI reasoning_effort, Gemini thinking budgets, or OpenRouter reasoning. It’s capability-gated, so it only appears on models that actually reason and clamps to a supported level when you switch. It’s available on the Oracle, the Genesis planners, and every Crew role.
Budgets that actually stop spend
Per-codebase and per-initiative token budgets have a soft warning at 80% and a hard stop at 100%. They are enforced in the execution path, not reported after the fact. Failed calls never accrue. One flat platform fee, identical for cloud and self-hosted.
Providers
Open-weights models
Why AnyForge
Compare us to anything. We built it for that.
AnyForge occupies a category of one: the platforms that govern don't deliver, the tools that deliver don't govern, and the frameworks that promise both hand you a build project. Here is the honest comparison.
vs. Coding agents (Claude Code, Cursor, Codex, Copilot)
They are agents; AnyForge is the platform above them. Keep them and govern every call through Control, or use Code Studio and Crew and get planning, delivery, audit and insights the tool layer cannot see. AnyForge is the only place one ledger covers your IDE agent, your browser agent and your background crews.
vs. LLM observability proxies (Langfuse, Portkey, Helicone-class)
Watching spend is not governing it. AnyForge adds enforcement: budgets that halt, policies that block, and approval gates the model cannot skip. It also adds an entire delivery layer (crews, specs, releases, incidents) that observability products do not attempt.
vs. Agent frameworks (LangChain, CrewAI, AutoGen)
Frameworks give you primitives; you still build the harness yourself: checkpointing, HIL, audit, cost control, and tenancy, then maintain it forever. AnyForge is that production harness, already built and governed, with a delivery track record you can audit.
vs. Building in-house
Every team that scales agents ends up building this platform. Routing, portable memory, policy engines, hash-chained audit, escrow, and capacity science add up to years of platform work. Install Control in five minutes instead, and keep your engineers on your product.
Get started
One flat fee. A hundred million free tokens. Five minutes.
Everything on this page is one platform with one price: a flat per-million-token platform fee, bring-your-own-key, no seats, no tiers. New accounts start with a 100M-token trial, and your first Crew run is on us.