qbrin.blog

qbrin field notes

Evidence you can inspect.

Clear guides to trustworthy AI, honest benchmark results, and recorded experiments from systems that can actually move, switch and decide.

Recorded Gazebo frame of the official PX4 X500 drone reaching its survey destination while qbrin shows ALLOW Actual recorded run Featured field test One destination. One rule: never guess. Watch the complete PX4 mission, including every ALLOW and HOLD decision. 10-minute film
ArchitectureVerificationBenchmarksPhysical AIRecorded experiments

The qbrin editorial series

How the trust layer works—and where it fits.

Six clear guides to qbrin's architecture and evidence, followed by three grounded opportunity maps for industrial and physical AI.

Start with how qbrin works →
Flow diagram showing sources entering qbrin, evidence retrieval and claim checks, then a cited answer or abstention Product architecture

How qbrin Works: Evidence Before Answers

A visual walk through the path from connected sources to a cited answer—or a deliberate stop when the evidence is not enough.

Layered architecture diagram with knowledge, verification and control planes connected by an evidence boundary System design

qbrin Trust-Layer Architecture Explained

A three-plane architecture for grounding what an agent knows, checking what it claims, and governing what it may do.

Decision fork diagram where verified evidence leads to answer and insufficient evidence leads to hold or abstain AI reliability

AI Abstention: Why Knowing When to Stop Matters

The business case, interaction design, and measurement discipline behind a system that sometimes says “not enough evidence.”

Three sequential gates labeled evidence, verification and authorization before an AI action Agent governance

Evidence vs Verification vs Authorization in AI

A precise way to separate source retrieval, claim support, and permission to act—three checks that are often collapsed into one.

Comparison diagram aligning RAG, observability, evaluation and qbrin by when each operates and what it controls Technology comparison

qbrin vs RAG, LLM Observability, and Evals

A fair category comparison: retrieval supplies context, observability supplies traces, evals measure behavior, and qbrin gates live claims and actions.

Benchmark chart comparing qbrin and traditional RAG on unanswerable false answers and citation trust Benchmark evidence

qbrin Benchmark Results: An Honest Guide

The strongest numbers, the sample sizes behind them, and the important places where the published report says qbrin does not lead.

Mission assurance diagram linking telemetry, a spacecraft and ground systems through a verification and authorization layer Space and aerospace

Verification for Space and Aerospace AI

An opportunity map for putting evidence, policy, and HOLD decisions between AI recommendations and mission actions.

Factory and IIoT flow diagram where sensor evidence passes through verification before a bounded action Manufacturing and IIoT

A Verification Layer for Manufacturing AI and IIoT

How to place evidence and policy checks around maintenance, quality, energy, and supervisory-control workflows without bypassing plant safety.

Radial diagram linking a central trust layer to robotics, drones, energy and biotech sectors Physical AI

A Trust Layer for Physical AI Across Industries

One reusable trust pattern—identity, evidence, verification, policy, and trace—adapted to four sectors with very different failure modes.

From the Qbrin lab

Real systems. Hard questions.

Fifteen measured experiments on what happens when AI agents meet infrastructure, other agents and unreliable evidence.

Follow via RSS →
Telemetry grounding precision experiment visual Telemetry Grounding

A safety gate that blocks 100% of real work

An exact-match telemetry check refused 100% of truthful operator actions.

Prompt injection retrieval boundary visual Security Architecture

Prompt injection stopped by retrieval

Our detector caught 5 of 7 attacks. A retrieval boundary leaked 0 of 8 records.

Verifier token cap parse error visual LLM Verification

A 400-token cap silently dropped answers

JSON truncation crashed the parser, dropping 18.4% of correct answers into refusals.

Transmission network experiment visual Critical infrastructure

An AI agent ran a power grid

Without an evidence check, the grid carried out every command and blacked out seven times.

Water treatment control experiment visual Industrial control

An AI agent ran a water treatment plant

The plant's own interlocks carried out 23 of 23 attack commands without objecting.

Benchmark integrity audit visual Benchmark integrity

Our benchmark said we were perfect. Then we broke it.

An independent audit found fifteen defects—twelve of them in our favour.

AI hallucination measurement visual Measurement

How to measure your AI's hallucination rate

Define the failure, build testable traps, score four outcomes, then audit the scorer.

Inter-agent identity and evidence visual Agent identity

Who said that—and can they prove it?

Identity proves who is speaking. It does not prove that what they say is true.

Citation grounding visual AI grounding

Why a citation isn't proof

A citation is a pointer, not a proof. Four ways a cited answer can still be wrong.

Multi-agent hallucination cascade visual Multi-agent systems

How one bad claim becomes consensus

In a multi-agent system, a hallucination becomes another agent's input.

Multi-agent verification visual Agent verification

When one AI agent believes another

A drone called a sector clear, citing a sweep that never happened.

Chemical process agent assurance visual Agent assurance

We let an AI run a chemical plant

63,589 decisions over eight hours. No invented reason ever moved a valve.

Hardware containment experiment visual Agent containment

We let an AI run our hardware

It asked for maximum power 87 times. The evidence allowed it only four.

Grounded answers benchmark visual Groundedness

0 invented answers across 120 traps

The benchmark, the audit and the evidence behind the number.