Arkheia · API Proxy
Stop flying blind
on AI calls
Point your model calls through Arkheia. Get runtime risk verdicts, cost attribution, and receipts without changing your application logic.
§ 00 · First Useful Call
Three steps to the first verdict
The self-serve path is intentionally small: make one AI call through Arkheia, see whether it was trustworthy, and keep the receipt.
01
Get a key
Create a free key or use an existing Arkheia key. The first useful verdict should not require a sales call.
02
Change one base URL
Point your model calls through Arkheia's OpenAI-compatible endpoint and keep your existing provider key.
03
Read the verdict
Responses keep their provider shape and add Arkheia risk, confidence, flags, recommendation, cost, and receipt metadata.
Endpoint swap
OPENAI_BASE_URL=https://api.arkheia.ai/v1 ARKHEIA_API_KEY=ak_live_... OPENAI_API_KEY=sk-...
What gets added
"arkheia": {
"risk_level": "HIGH",
"confidence": 0.84,
"recommendation": "hold_for_review",
"receipt_id": "rcpt_abc123"
}§ 01 · Who It's For
Designed for the 10× development workflow
Primary audience
Developers building with AI agents
You're using Claude Code, Cursor, or similar AI coding tools — or building agent pipelines yourself. You ship fast. Arkheia sits in your API path and surfaces risk in real time, so you know which outputs to trust and which to verify. No workflow changes. Integration in 15 minutes.
Also works for
Anyone comfortable in a terminal or IDE
If you can change a base URL and add a header, you're set. No infrastructure to run, no SDK to install. Works with any language or framework that makes HTTP calls to OpenAI, Google, or Anthropic.
§ 02 · Integration
Three steps. Fifteen minutes.
01
Change your endpoint
Point your API calls to api.arkheia.ai instead of api.openai.com (or api.google.com, api.anthropic.com). One line of config.
02
Pass your API key
Use your existing provider key. We forward your requests on your behalf — no new model accounts, no new quotas.
03
Get enriched responses
Standard responses return unchanged. Risk metadata is appended: risk level, confidence, which signals fired, and a recommendation.
Enriched response shape
{
"choices": [ /* standard provider response, unchanged */ ],
"arkheia": {
"risk_level": "HIGH",
"confidence": 0.84,
"flags": ["signal_anomaly"],
"recommendation": "verify_manually",
// What we measured for this model, so the verdict can be
// weighed. Recall never appears without its false-alarm rate.
"provenance": {
"measured": true,
"measured_on": "2026-08-27",
"held_out_recall": 0.745,
"held_out_fpr": 0.120,
"auc_ci_95": [0.738, 0.907]
},
// Not a fabrication verdict — a workflow signal, raised even
// when risk is LOW. A step that produced nothing, or a request
// whose premise the model appears to doubt.
"advisory": {
"surface": true,
"code": "premise_may_be_unsound"
}
}
}Claude Code · Cursor · MCP agent frameworks
Using Claude Code or Cursor? The MCP Trust Server integrates at the tool layer — no endpoint changes needed.
§ 03 · Use Cases
Where detection matters
Agent pipelines
Multi-step agentic workflows where one abnormal output cascades. Per-invocation detection lets you gate, retry, or escalate before the pipeline continues.
Code generation
Flag when models reference non-existent libraries, generate uncertain patterns, or produce output that deviates from their established baseline.
Legal and research tools
Detect when AI cites sources or precedents with atypically high confidence across uncertain terrain. Reduce verification overhead on high-stakes output.
Customer-facing AI
Know when a response needs human review before it reaches a user. Reduce incorrect information at the point of generation, not after.
§ 04 · Coverage Posture
Built to follow the models serious teams actually use.
Detection is profile-based, not a generic threshold. Current public model families are covered as named profiles or held as explicit onboarding targets before production use. We do not claim support for models we cannot access or characterize.
83
models with demonstrated detection
141
validated profiles
18
confirmed in the last 45 days
300
held-out splits behind every figure
OpenAI
- Current GPT frontier family
- Mini / nano variants
- Codex / coding profiles
- GPT-OSS where telemetry permits
Anthropic
- Current Claude Opus family
- Current Claude Sonnet family
- Current Claude Haiku family
- Fable 5: not claimed until accessible
- Current Gemini Pro family
- Current Gemini Flash family
- Flash-Lite / low-latency variants
- Prior Gemini family fallbacks
xAI
- Current Grok frontier family
- Fast / lightweight variants
- Coding / build profiles where available
Open Weights
- Llama / Mistral / Qwen
- DeepSeek
- Kimi / Moonshot
- GPT-OSS
- Customer local models
Visible Model Lab registry snapshot supplied 2026-06-27; this is one surface of the broader six-month characterisation corpus. Production gating requires a pinned model profile. Newly released or inaccessible models are observe-only until characterisation passes.
§ 05 · Data Policy
Your keys. Your data. Your control.
What we observe
- Requests pass through in transit only
- Behavioural signals: token probabilities, timing
- Never stored — extracted and discarded
What we store
- Aggregated risk metrics for your dashboard
- Usage statistics (request counts)
- NOT prompts · NOT responses · NOT history
Your API keys
- Bring your own keys — no new accounts
- We forward on your behalf
- Encrypted at rest
- You control access and rate limits
§ 06 · Confidence and Limits
What the numbers mean, and what they don't
Detection you cannot interrogate is not evidence. Every figure is stated with what qualifies it, because a number without its qualifier is the easiest way to mislead without lying.
Why a missed detection costs far more than a false alarm
A false alarm costs one regeneration — bounded, immediate, recoverable. A miss does not stop at reaching a commit. It enters the repository, the documentation, the knowledge base, and becomes an organisational asset. Once inside the trusted corpus it becomes grounding: a model asked about it later is no longer uncertain, it is grounded — just wrongly. The very signal that would have caught it disappears, and retrieval now confirms it. One error is recoverable; the other compounds while erasing the evidence of itself. Thresholds are set accordingly.
Recall never travels without its false-alarm rate
A detector can catch everything by flagging everything. Quoted alone, either number is meaningless — so results always carry both. On one current profile that is 74.5% of affected responses caught at a 12.0% false-alarm rate; on another, 98.4% caught at 25.6%. The trade-off is a dial, and on some models the same detection rate is available at half the false alarms.
Every figure is held out
Thresholds are fitted on half the truthful responses and scored on the half they never saw, averaged over 300 random splits. Results carry a 95% confidence interval, and where that interval is wide we say so rather than quoting the midpoint as settled.
A profile is a dated observation, not a permanent claim
These were the thresholds on the day of measurement, on the telemetry the vendor returned that day. Models change underneath profiles — one moved its reasoning burn by two-thirds on identical prompts inside six weeks. Every verdict carries the date it was measured against, and a model we have never characterised returns provenance.measured: false rather than staying silent.
It detects a questionable request as much as a fabricated answer
The benchmark asks about things that do not exist, and frontier models usually get that right and say so. What is measured is the signature of a model working on a request whose premise is false — often the user's, not the model's. Catching that early stops a workflow building on it. It is not the same as catching a model that is confidently wrong and believes itself, and we have not tested that case.
There is no judge model
No second model grades the first, no reference answer, and the prompt and response are never inspected. Detection reads the observable process of generation — which is why it works from outside with no privileged access, and why what a vendor exposes sets the ceiling on what can be measured.
§ Get Started
Add Arkheia to one AI call
Start in observe-only mode. Upgrade when you need retention, team projects, policies, or audit exports.