# Arkheia — Claims & Evidence

*A take-away diligence document for the technical evaluator. Same facts and maturity labels as the
diligence page at https://synesis.arkheia.ai/site/diligence/ — readable in isolation.*

## How to read the maturity labels

Every capability below is labelled. This is a governance company; a diligence document that
overclaims is self-defeating, so the labels are deliberate.

- **live** — public and testable today.
- **demo-grade** — a working demonstration of the capability.
- **dogfood** — we run our own operations on it.
- **pilot** — offered/run with an evaluation customer, not a public surface.
- **roadmap** — a stated direction, not yet built.

**On metrics:** this document publishes **no** precision, recall, accuracy, or effect-size numbers.
We have not measured them for your prompts, and quoting numbers you can't reproduce would be the
overclaim we are trying to avoid. Where a number would normally go, the honest answer is "measure it
on your own prompts" — see `verify-it-yourself.md`.

---

## What Arkheia is

A **governed AI capability company whose unfair advantage is that it built the trust layer first.**

Read this document from the **capability down**: useful AI work is the point; doing real AI work needs
trust; trust needs governance on the execution path; governance needs a runtime signal that can tell a
good answer from a bad one. Arkheia built that signal — runtime detection — first, and turned it into a
control plane the capability work runs on top of. **Detection is the moat, not the doorway.** And
**autonomy is earned, not granted**: an AI agent starts by watching and earns the right to act.

**Built from detection up. Sold from capability down.**

---

## The claims

### 1. Governed AI workflows — what it actually does

These are the workflows the governed runtime does today. Each carries a maturity label; none asserts a
metric beyond the single observed POC datapoint in claim 2.

- **Support triage — demo-grade / in build, in live POC (see claim 2).** An AI support-triage agent
  over a live ticket queue: it reads ticket history, the knowledge base, code, and docs to classify,
  route, and draft a reply. It runs along the earned-autonomy path (below) — critique-only → private
  note → bounded → supervised → full auto — with the detection signal and the governance gate on the
  execution path throughout.
- **CRM and revenue lifecycle — dogfood.** Governed CRM, quote-to-cash, and revenue workflow surfaces
  built and operated on the same runtime.
- **Data-trust scoring — demo-grade.** Scoring the trustworthiness of an incoming data stream so
  downstream automation acts on signal, not noise (see the manufacturing-monitoring demo, claim 5).
- **Reconciliation — demo-grade.** Handwritten-docket ↔ invoice reconciliation under a governed
  decision gate (see the movement-control demo, claim 5).
- **Other governed surfaces — dogfood.** The same governed runtime builds and operates surfaces across
  **procurement, finance close, legal ops, HR service, PPM, marketing, and assurance.**

This is **proof of governed capability** — not a product you must buy before the governance is real.

### 2. Live customer POC — **POC / pilot, not production**

A current evaluation, run as a **proof-of-concept, not a production deployment**:

- The support-triage agent runs **on a support agent's own laptop**, against that team's **real
  Freshdesk ticket queue and knowledge base**.
- The underlying model is **GPT-5.4**; **Arkheia detection** scores each invocation; the agent's output
  is written as **private internal triage notes** (the human stays in the loop — nothing is auto-sent to
  customers at this stage).
- **Observed:** a triage task that took roughly **3 hours** dropped to roughly **2 minutes** in this
  setup.

**Honest framing:** this is one observation in one POC on one queue, with a human reviewing the output.
It is not a benchmark, not a production SLA, and not a claim about your queue. It is what we saw, stated
plainly, with the customer un-named.

### 3. Earned-autonomy path — how an agent earns the right to act

Autonomy is **staged and earned, never granted up front.** An agent moves up only as it demonstrates
trustworthy behaviour under the detection signal and the governance gate:

1. **Critique-only** — the agent observes and comments; it takes no action.
2. **Private note** — it drafts an answer visible only to the operator, never to the end user/customer.
3. **Bounded** — it may act, but only inside a narrow, explicitly allowed envelope.
4. **Supervised** — it acts on the real workflow with a human approving each action.
5. **Learning loop** — its decisions, the verdicts, and the human corrections feed back to tighten its
   behaviour.
6. **Full auto** — it runs the workflow unattended, still scored and gated on every invocation.
7. **Workflow ownership** — it owns the end-to-end workflow, with governance still on the execution path.

The detection signal and the execution gate ride along at every stage. That is what makes turning
autonomy *up* a controlled decision rather than a leap of faith.

### 4. Governance / control plane — **dogfood + live**

- The detection verdict becomes an **execution gate** on the path: **allow / block / require human
  approval / kill-switch / rollback / decision receipt.**
- Governs **API models** (OpenAI, Anthropic, Gemini, xAI & compatible), **local / self-hosted
  models** where deployment permits, and **MCP / tool calls**.
- Emits **tamper-evident decision receipts**, **cost attribution**, and **audit evidence**.
- A **learning loop** feeds verdicts and human corrections back so the governed agent improves over
  time rather than repeating mistakes.

### 5. Runtime fabrication / deviation detection — the moat — **live (public demo-grade)**

This is the layer built first, and the reason the rest can be turned up safely.

- Runs at the **invocation boundary**.
- Builds **per-model behavioural fingerprints** and scores each invocation against *that model's own
  baseline* — a runtime per-model baseline, not a generic content filter.
- Produces a **runtime risk verdict** (e.g. LOW / MEDIUM / HIGH), a **confidence**, and a
  **detection id**.
- That verdict is what feeds the **execution gate** in claim 4.
- A live public demo exists at **https://synesis.arkheia.ai/demo/** — anyone can try it now. The
  browser calls a public Arkheia backend proxy, **not** the model vendor directly.

**Evidence you can check:** open `/demo/` and run your own prompts.

**The honest caveat — evidence-limited LOW.** Coverage depends on per-model profiles. A model with no
profile returns an **evidence-limited LOW**, which means *"couldn't assess"* — not a clean bill of
health. This is surfaced, not hidden.

**Content-minimising posture — live (design posture).** The detection signal can run
**content-minimising / without retaining prompt or response content** — **zero-retention by default**.
The behavioural method measures behaviour, so it does not require keeping the text it scores.

### 6. Proof demonstrations — **demo-grade**

- **A UK manufacturing-monitoring SaaS** — manufacturing event stream → operational value, data-trust
  scoring, scheduled analyses.
- **An earthworks / movement-control SaaS** — a governed earthworks movement-control +
  handwritten-docket ↔ invoice reconciliation demo (available on request).

### 7. Intention mapping / capability teardown — **demo-grade / in build**

Model what a product or workflow is really trying to do — the underlying intents, not the screens —
and surface where value leaks: steps that are manual, brittle, duplicated, or unmeasured. The output
is a map a buyer can act on. We do not quote an effect size here; the value-leakage findings for a
given workflow are measured in a pilot.

### 8. In-stack evaluation — **self-serve**

Run detection against your own prompts via the **API proxy** or **MCP server**, inside your
environment, from one public repo: **https://github.com/arkheiaai/arkheia-mcp**.

- **MCP Trust Server** — one-line install (`npx @arkheia/mcp-server`); ~60-second installer (needs
  Node 18+, Python 3.10+, and an API key).
- **API / Enterprise Proxy (self-host)** — clone the repo and `docker compose up`; point your SDK base
  URL at the local proxy so inference is scored **in your own environment** (secure transit — data
  doesn't leave). The proxy is a thin shim; detection logic runs server-side.
- **Free tier:** 1,500 detections / month, no credit card. Mint a free key at
  **https://app.arkheia.ai/signup**.

The immediately testable surface remains the public demo. For higher-tier or managed evaluation keys,
contact Arkheia — the public download is the primary, self-serve route.

---

## What we deliberately do **not** claim

- No "replacement-grade SaaS" without a maturity label.
- No "build anything."
- No agent head-counts.
- No detection accuracy / precision / recall / effect-size figures on a marketing surface.

---

## Reachable today

- Home: https://synesis.arkheia.ai/site/
- Diligence pack: https://synesis.arkheia.ai/site/diligence/
- Live detection demo: https://synesis.arkheia.ai/demo/

The earthworks movement-control + docket↔invoice reconciliation walkthrough is available on request.

For in-stack evaluation, the primary route is the public download: the MCP Trust Server and the
self-host Enterprise Proxy both ship from **https://github.com/arkheiaai/arkheia-mcp**
(`npx @arkheia/mcp-server`, or clone + `docker compose up`), with a free key
(1,500 detections/month, no card) at **https://app.arkheia.ai/signup**. For higher-tier or managed
evaluation keys, contact Arkheia. No internal endpoints, tokens, or auth flows are published — none are
public.
