Empirical Truthfulness

Live Verification Benchmarks & Methodology

We do not make unsubstantiated marketing claims. Below are the verified results from our automated Live Growth QA Gate executed against production endpoints.

Live QA Gate Test Summary (September 2026)

Bot Service Deployment Infrastructure Test Vectors Pass Rate Median Latency P95 Latency
OCR-Doc-Parser Render Web Service (Node.js + Tesseract WASM) 10 Scenarios (Receipt, Skew, Blur, Glare, Injections) 10 / 10 (100%) 4.12s 4.36s
Regex-Gen-Tester Cloudflare Workers (Edge V8 JavaScript) 7 Scenarios (Email, Mobile, ReDoS, Unicode) 7 / 7 (100%) 28ms 35ms
English-To-SQL Cloudflare Workers (Edge SQLite WASM) 8 Scenarios (JOIN, Aggregates, DDL, Retries) 8 / 8 (100%) 39ms 48ms
Protocol & Auth All 3 Live Production Endpoints 5 Scenarios (SSE framing, Health, 401 Rejections) 5 / 5 (100%) 12ms 18ms

Testing Methodology & Invariants

  • No Synthetic Hallucinations: In OCR tests, receipts with blurred or cropped totals must either calculate totals through line-item arithmetic or explicitly flag low_confidence. High-confidence incorrect totals trigger an immediate test failure.
  • Zero Secret Leakage: Unauthenticated requests to all three live endpoints are tested with missing or malformed Bearer tokens. Tests verify that HTTP 401 is returned without exposing stack traces, keys, or server paths.
  • Stateless Bounded Execution: Regex evaluation executes in an isolated sandbox with a 50ms hard execution ceiling. SQL queries execute inside a fresh, in-memory SQLite database instance with output capped at 50 rows.

Identified Operational Boundaries