qbrin field notes
Evidence you can inspect.
Clear guides to trustworthy AI, honest benchmark results, and recorded experiments from systems that can actually move, switch and decide.
Actual recorded run
Featured field test
One destination. One rule: never guess.
Watch the complete PX4 mission, including every ALLOW and HOLD decision.
The qbrin editorial series
How the trust layer works—and where it fits.
Six clear guides to qbrin's architecture and evidence, followed by three grounded opportunity maps for industrial and physical AI.
Product architectureHow qbrin Works: Evidence Before Answers
A visual walk through the path from connected sources to a cited answer—or a deliberate stop when the evidence is not enough.
System designqbrin Trust-Layer Architecture Explained
A three-plane architecture for grounding what an agent knows, checking what it claims, and governing what it may do.
AI reliabilityAI Abstention: Why Knowing When to Stop Matters
The business case, interaction design, and measurement discipline behind a system that sometimes says “not enough evidence.”
Agent governanceEvidence vs Verification vs Authorization in AI
A precise way to separate source retrieval, claim support, and permission to act—three checks that are often collapsed into one.
Technology comparisonqbrin vs RAG, LLM Observability, and Evals
A fair category comparison: retrieval supplies context, observability supplies traces, evals measure behavior, and qbrin gates live claims and actions.
Benchmark evidenceqbrin Benchmark Results: An Honest Guide
The strongest numbers, the sample sizes behind them, and the important places where the published report says qbrin does not lead.
Space and aerospaceVerification for Space and Aerospace AI
An opportunity map for putting evidence, policy, and HOLD decisions between AI recommendations and mission actions.
Manufacturing and IIoTA Verification Layer for Manufacturing AI and IIoT
How to place evidence and policy checks around maintenance, quality, energy, and supervisory-control workflows without bypassing plant safety.
Physical AIA Trust Layer for Physical AI Across Industries
One reusable trust pattern—identity, evidence, verification, policy, and trace—adapted to four sectors with very different failure modes.
From the Qbrin lab
Real systems. Hard questions.
Fifteen measured experiments on what happens when AI agents meet infrastructure, other agents and unreliable evidence.
Telemetry GroundingA safety gate that blocks 100% of real work
An exact-match telemetry check refused 100% of truthful operator actions.
Security ArchitecturePrompt injection stopped by retrieval
Our detector caught 5 of 7 attacks. A retrieval boundary leaked 0 of 8 records.
LLM VerificationA 400-token cap silently dropped answers
JSON truncation crashed the parser, dropping 18.4% of correct answers into refusals.
Critical infrastructureAn AI agent ran a power grid
Without an evidence check, the grid carried out every command and blacked out seven times.
Industrial controlAn AI agent ran a water treatment plant
The plant's own interlocks carried out 23 of 23 attack commands without objecting.
Benchmark integrityOur benchmark said we were perfect. Then we broke it.
An independent audit found fifteen defects—twelve of them in our favour.
MeasurementHow to measure your AI's hallucination rate
Define the failure, build testable traps, score four outcomes, then audit the scorer.
Agent identityWho said that—and can they prove it?
Identity proves who is speaking. It does not prove that what they say is true.
AI groundingWhy a citation isn't proof
A citation is a pointer, not a proof. Four ways a cited answer can still be wrong.
Multi-agent systemsHow one bad claim becomes consensus
In a multi-agent system, a hallucination becomes another agent's input.
Agent verificationWhen one AI agent believes another
A drone called a sector clear, citing a sweep that never happened.
Agent assuranceWe let an AI run a chemical plant
63,589 decisions over eight hours. No invented reason ever moved a valve.
Agent containmentWe let an AI run our hardware
It asked for maximum power 87 times. The evidence allowed it only four.
Groundedness0 invented answers across 120 traps
The benchmark, the audit and the evidence behind the number.