Evaluations and Benchmarks

Every evaluation runs on your own telemetry.

AI Detection and AI Investigation · Updated Sep 24, 2026

What would a trustworthy score require?

Your telemetry, your labels and criteria agreed before the first result. Cyberdroid AI Detection and Cyberdroid AI Investigation are evaluated that way, in your environment.

See the method

Method

Every evaluation runs on your telemetry, against a scorecard agreed before it begins.

AI Detection runs side by side on labelled replay scenarios and at least one representative week of your telemetry.

AI Investigation is tested on malicious cases, benign look-alikes, contradictory evidence and missing telemetry.

Event-to-finding latency includes the hourly wait. Missing-feed handling is scored, not excused.

Cost is measured per analyst-accepted investigation, and savings as minutes per case, never headcount.

Measures

For Detection, precision, false positives, duplicates, time to disposition, and CPU, memory and storage. For Investigation, correct and incorrect conclusions, unsupported assertions and reviewer effort.

Limits

We publish no benchmark score today. Ground truth depends on who labels, and benchmarks for investigation agents are young.

Sources

Related

Detection Model Training

Method

Agent Architecture

Method

AI Detection

Product

AI Investigation

Product
Contact
ARRTECH

© 2026 ARRTECH Corporation. All rights reserved.