The platform

Evidence-backed proof for supply chain AI agents.

Aptha Assure is our evaluation platform for AI agents in supply chain. Right now, it does one thing and does it rigorously: finds where an agent's code, logic, or governance actually breaks down, and turns that into a Trust Score you can act on.

Mandatory for every evaluation

The Agent Functional Template

An agent can run cleanly and still be judged against the wrong yardstick, if no one has written down what it's actually supposed to do. Before Aptha Validator scores an agent, you complete an Agent Functional Template — a structured description of its scope, inputs and outputs, responsibilities, and decision policies. It's a required input, not an optional detail: without it, there's no reliable way to tell whether a behavior is a genuine defect or simply outside the job the agent was built for.

Scope, defined up front

What the agent is responsible for, and just as important, what it's explicitly not — so a finding never penalizes an agent for something it was never meant to handle.

Inputs, outputs, and decision policies

The data the agent should use, the decisions it's allowed to make, and the constraints those decisions have to respect.

The reference every result is measured against

Scenario generation, defect detection, and the Trust Score all trace back to this template — it's what turns a plausible-looking answer into one that's actually verified.

How it works

Four steps, the same way every time.

Every submission runs through the identical process, so a score means the same thing on your tenth agent as it did on your first.

1

Submit

Point Aptha Validator at a repository, a ZIP file, or a set of agent files, together with the agent's Agent Functional Template.

2

Classify

The agent's task and domain are identified automatically, so there's no manual setup per agent.

3

Generate & run

A tailored set of supply chain scenarios is generated and exercised inside an isolated sandbox.

4

Score & report

Findings are measured against independent reference results and turned into a Trust Score, backed by evidence.

Aptha Validator

Find out where a supply chain agent actually breaks down.

Aptha Validator holds AI agents to the kind of rigor enterprises expect from database systems. Point it at a repository, a ZIP file, or a set of agent files — along with the agent's Agent Functional Template — and it classifies the agent automatically, then runs it through an evidence-based assurance process, surfacing defects and turning them into a Trust Score your team can track and trend release over release.

Findings, backed by evidence

What gets checked

Code logic

Runtime and architectural defects — unhandled errors, fragile entrypoints, missing persistence — the kind of thing that breaks silently the first time a real input doesn't match the demo.

Business logic

Whether the agent's decisions actually hold up against supply chain reality — reorder math, supplier weighting, demand assumptions — not just whether the code executes.

Governance

Whether a decision can be traced back to the evidence that produced it, and whether the agent stayed inside the boundaries defined in its Agent Functional Template.

Risk

How the agent behaves under dirty data, missing fields, and environment failures — the conditions a clean demo never has to face.

Compliance

Documented, repeatable evidence of how the agent was tested — the kind of record regulation is starting to require for high-risk deployments.

Security

Resistance to manipulation — an input or context crafted to push the agent past the limits it was scoped to.

Want to see Aptha Validator run against one of your own supply chain agents? Get in touch and we'll walk you through it.

Proof in a single run

A real score, on a real build.

In one run, we validated two successive builds of the same inventory-reorder agent through an identical 58-scenario suite, holding every condition constant so the result reflects a real change in the agent's own behavior.

Build A1

Trust Score

0 / 100

Defects found

0

Build A2

Trust Score

0 / 100

Defects found

0

Every run also produces documented evidence of the kind regulators are starting to expect for high-risk agent deployments — generated automatically, at the moment of every release.

Ready to see Aptha Assure on your workflows?