HEIMDALL control plane
Pre-launch · for AI-heavy engineering orgs

Run every AI task on the most cost-efficient model that can actually do it.

Heimdall routes each coding task to the most cost-efficient capable model and makes every result prove itself: code changes must pass the compiler; answers must cite file:line sources that Heimdall mechanically verifies against your repo. Escalation happens only when the evidence fails — with DLP, budgets, and audit on every request.

The watchman's promise: as much control as each tool allows — full visibility regardless. Nothing crosses the gate without Heimdall seeing it.

The Gate · cascade + verify $0.000
“Where is the SAML assertion validated?” code · task
gemini‑flashlowest cost ~$0.004
haikulow cost ~$0.011
opuspremium ~$0.130
One schema, every provider
Claude GPT Gemini Your AWS Bedrock Cursor · Figma
The problem

Every company is becoming an AI company. Almost none can run AI like infrastructure.

When large IT-services firms tell employees not to use AI for small tasks, it isn't anti-AI — it's a company with no way to control when and which AI gets used. Today the only choices are both bad:

Choice one

Ban AI for small tasks

You claw back the invoice — and lose most of the productivity the tools were bought for.

versus
Choice two

Let everyone use premium AI

Cursor, Claude, GPT on everything — the bill explodes, and most of it is overkill for the task.

Nobody routes intelligently — small task to a low-cost model, hard task to the best model, automatically, with quality protected and every request governed. That gap is Heimdall.

How it works — cascade + verify

Don't predict which model a task needs. Try low-cost, then prove it.

Prediction is unreliable — we measured it, and so has everyone else. So Heimdall runs the lowest-cost loop-capable model first, demands evidence, and only climbs the price ladder when the evidence fails. The gate is deterministic checks — never an AI judging an AI.

Run the most cost-efficient capable model

An agentic tool-loop lets the model grep, read, and iterate to gather the exact code context it needs — starting on the lowest-cost model that can actually drive that loop.

Demand evidence, verify it mechanically

Code changes must parse, type-check, and build. Answers must cite file:line sources — Heimdall verifies every cited file and symbol actually exists in your repo, and a second independent sample must ground in the same code. No LLM judges.

Ship, or escalate one tier

Pass: done, paid pennies. Fail: Heimdall recovers and escalates one tier up, then verifies again. Most tasks stay low-cost; only the hard tail reaches premium.

You always see which model ran, what evidence was checked, and how deeply the result was verified. Nothing ships on trust — only on proof.
Anatomy of a request

Follow one request from prompt to shipped — through the gate.

This is the exact path every task takes. Governance happens before any provider sees your data; the model is chosen by difficulty; the cascade climbs only when a verify check fails. Watch it run.

Request
“Where is the SAML assertion validated?” code · task · signed-in dev
Gateway
DLP · no secrets policy · allowed budget · 62% used
Router
risk consequence × complexity signals · retrieval · history start → gemini-flash
Model cascade + verify
harder · premium
opus~$0.13
sonnet~$0.078
haiku~$0.035
gemini-flash~$0.013
lower cost · pennies · measured per-task, our benchmark
escalate ↑ evidence failed — climb one tier
Ship & meter
evidence verified · shipped $0.048 · 2 tiers used metered + audited
tieridle running cost$0.000 statuswaiting
The product

One governed control plane between your developers and every AI provider.

Routing is the mechanism. The product is everything an enterprise needs to actually deploy AI at scale — and the memory that makes it better and less costly over time.

Intelligent routing

Each task lands on the most cost-efficient capable model, with the routing decision, difficulty, per-request cost, and projected savings shown for every call.

The moat

Shared org memory

One memory store shared across models — and, opt-in per project, across the org. Tell Claude something; ask GPT and it knows, because the knowledge lives in Heimdall, not the model. Every verified task enriches the next one.

Governance, built in

DLP secret-scanning, per-team budgets and token caps, and allow / deny / limit policies evaluated at the gateway before each request — with an append-only audit trail.

Outcomes you can prove, not just spend

Cross-provider spend in one schema, waste detection (idle seats, right-sizing, redundant tools), and — via the AI Ledger — cost per shipped, production-healthy story, not just token counts. See how →

Lives in VS Code

The full surface inside the editor: sends the exact code spans a task needs — not whole files — with the model, tokens, cost, and % saved shown on every turn, and native diff → Apply.

Runs in your account

Heimdall is not a reseller. Traffic runs through your own AWS Bedrock account; keys are vaulted and never sent to clients. No-log mode, SOC 2 & GDPR on the roadmap.

The honest part

As much control as each tool allows. Full visibility, regardless.

Closed tools can't be governed like an API you proxy — so we don't pretend otherwise. Every AI surface is tagged by exactly how much control is real. That honesty is what a CISO can actually deploy.

Governed

Traffic proxied through the gateway

Hard budget and DLP enforcement before the request reaches a provider. Full control — for the API models you route.

Managed

Access & seats provisioned

Usage pulled from the tool's admin API. You control who has a seat and see what they spend, even when you can't proxy the traffic.

Tracked

Visibility only

For the most closed tools: usage and cost surfaced in one place, so nothing is invisible — even where enforcement isn't possible.

Three levels, one schema. Nobody else has every provider's cost, usage, and policy in a single normalized store.

Proof, not promise

Measured on a 1,100-file production repo. Graded against verified ground truth.

Two benchmarks on a 1,105-file production repo: 31 hard questions for model economics, and a sealed 16-question held-out exam — brand-new questions, no tuning possible, one attempt — for the evidence-gated router itself.

27/31 Haiku tied Sonnet a low-cost model matched premium on the bulk of work
~90% cache-read discount reused context costs ~0.1× — break-even at 2 requests
100% turns metered & audited every call attributed to the real signed-in user
31/32 two sealed exams, two repos held-out questions, JS + TS, frozen rules — incl. a perfect 16/16
Real session — unedited producttyped question → low-cost model answers with verified citations → a break is caught, repaired one tier up, and proven before apply → the org dashboard
Where the 16 held-out tasks landed under the evidence cascademeasured
62% low-cost tiers 38% escalated
Answers that PASSED mechanical evidence checks — pennies each Escalated on failed evidence — premium only where it's earned
Exam 1, same-exam fixed models: router 15/16 · gemini-only 14/16 · opus-only 13/16 · haiku-only 12/16 · Exam 2 (different repo & language): router 16/16rankings flip between workloads — the router was never beaten
The router beat every fixed model on held-out questions — including premium at half its cost. Evidence checks are deterministic — no LLM judges; fabricated citations cannot pass.
Prove the outcome, not the activity

Everyone measures AI activity. Nobody can prove AI outcomes.

Engineering-analytics tools see commits and tickets but not what was AI-made, what it cost, or whether it shipped and stayed healthy. Billing dashboards see spend but no outcomes. Heimdall is in the request path and joins that spend to the work it produced — so it can put cost per shipped, production-healthy story on one screen. Team-level by design, never ranking individuals.

Every request carries its work key

The extension stamps each call with the session and task it belongs to, so spend is joinable to outcomes from the moment it happens — not reconstructed after the fact.

Only we can

Spend joins to commits, stories & releases

The AI Ledger connects dollars → commits → shipped stories → releases via your CI/CD and issue tracker. Because we own both ends of the chain, the join is real — not a guess a dashboard makes from the outside.

Cost per shipped, production-healthy story

Release health and incidents fold back in, so a story only counts once it shipped and stayed up. The CFO number stops being "tokens spent" and becomes "dollars per outcome that held in production."

Both ends of the chain. In the request path for cost, joined to CI/CD and production for outcomes — the one place they meet. And it's team-level by design, which is what makes it sellable past a works council.

Why now & the moat

The router is copyable. Your org's context is not.

Adoption

AI is arriving bottom-up, tool by tool

Teams adopt Cursor, Claude, GPT, and agents faster than any company can govern — with no single place to manage who uses what, at what cost, under what policy.

The flywheel

Every routing outcome is logged and learned

Every one of your org's routing outcomes — accepted evidence, escalations, failures — teaches the router; better context lets an even lower-cost model succeed next time. That compounds with use.

Switching cost

Leaving means losing years of context

When Heimdall becomes where your engineering knowledge accumulates, it's the durable moat a competitor can't copy — the same job Okta did for SaaS identity.

Both ends

We own the request path and the outcome

Analytics vendors see commits but not what was AI-made or what it cost; billing dashboards see spend but no outcomes. We're in the request path and joined to shipped, production-healthy work — so only we can put cost-per-outcome on one screen.

The thesis

“Everyone sells AI usage. We're building AI accountability.”

Same quality, a fraction of the cost — with shared org memory as the moat we build next.

Pricing

Each role gets only the AI it needs — bought at volume, on one governed bill.

The customer pays less than buying tools separately, because API usage routes to the most cost-efficient model that works. We don't make Cursor cheaper than Cursor — we make the whole stack cost less.

Developer

Full-stack AI seat

For engineers shipping code every day.

  • Cursor at volume + routed Claude / GPT pool
  • VS Code extension + shared memory
  • Full governance & audit
Knowledge

Routed API seat

For everyone else who works with AI.

  • Routed GPT / Gemini / Claude pool
  • No premium seat to pay for
  • Full governance & audit
Designer

Creative AI seat

For design and product teams.

  • Figma AI at volume + image / chat pool
  • Shared memory across tools
  • Full governance & audit

Role-based bundles priced per seat on your mix and volume. We keep a healthy margin; you still pay less than separate tools.

Why this team

The hard part isn't another chat box near code. It's deciding who can do what, what data can leave, what a task should cost, and whether the result is trustworthy.

AA Abhishek Anand Founder — identity, product,
engineering systems

Background in identity & access control — direct founder-market fit around policy, provisioning, and developer adoption.

MJ Mudit Jaroli Co-founder — product
& technical execution

Focused on making the control-plane thesis work as a daily tool engineering teams actually want to use.


See Heimdall on your repo

Let developers use AI. Just route it, verify it, and govern it.

A 30-minute walkthrough on your own codebase: watch the cascade route, verify, and escalate — with cost and policy on every request.