Evaluations

We publish the method, and you measure the result on your telemetry.

AI Detection and AI Investigation · Updated Sep 23, 2026

How do we know it works?

You measure it on your own telemetry, against a scorecard agreed with us before the first case. Nothing is scored against criteria set afterwards.

Diagram: three steps. Agree the scorecard first, run on your telemetry, then score against it. The first step is highlighted.

Measures

AI Detection runs beside your tooling on labeled replay scenarios and at least one representative week. AI Investigation runs on a labeled set of hard cases, with success and stop criteria in writing.

Detection

Time to disposition, precision, false positives, duplicates, findings routed to shadow, event-to-finding latency including the hourly wait, missing-feed handling, and the CPU, memory and storage used.

Investigation

Correct and incorrect conclusions, unsupported assertions, cases parked incomplete, reviewer effort, minutes per case and cost per analyst-accepted investigation.

Limits

Scores

This page publishes the method, not a score. A number measured on other data would not predict yours.

Outcomes

An evaluation reports minutes per case and cost per accepted investigation. It promises no percentage saving.

Status

AI Investigation is in early access, so its evaluation is a design-partner evaluation.

Standards

Naming a clause is not a claim of certification.

NIST

AI RMF MEASURE 2.1 asks for documented test sets and metrics, and MEASURE 2.5 for limits on generalizing. Measuring on your telemetry answers the second.

EU

Articles 13 and 15 of the EU AI Act ask for declared accuracy metrics. The scorecard declares ours before a case runs.

Related

Responsible Deployment

Explainability

System Cards

Cyberdroid AI Detection

Cyberdroid AI Investigation

Contact
ARRTECH

© 2026 ARRTECH Corporation. All rights reserved.