Your agents made 1.4 million decisions last quarter. You can reconstruct none.

Enterprises put autonomous systems into production faster than any technology in thirty years. The controls did not go in with them. Venture Vertex builds the layer that governs, prices and defends those systems while they are running — four platforms, in production, in regulated environments.

THE PROBLEM
2024 — 2026

Every control an enterprise owns was designed to be applied before the system ran.

Model risk management. Third-party review. Change advisory boards. Pre-production sign-off. Architecture review. All of it assumes the same shape, and it is a shape that has worked for forty years: a system is built, a committee inspects it, the system is released, and it behaves in production the way it behaved during inspection.

Autonomous agents broke that assumption in about eighteen months. An agent chooses its own path through tools. It calls other agents. It reads data nobody approved for that purpose. It produces a different action on Thursday from the one it produced on Tuesday, with no code change, no release, and no ticket to point at. The artifact the committee inspected is not the thing that is running.

The questions changed at the same time. Boards stopped asking whether the model was validated and started asking what the agent did last Tuesday, who authorised it, and whether it can do it again. The Federal Reserve extended model risk governance to AI agents in production. The EU AI Act put Articles 11, 12 and 14 into force with fines at six per cent of global turnover. The FCA published explicit guidance. The CFPB and OCC made clear that an agent making a credit decision is subject to fair lending law regardless of intent.

Very few institutions can answer those questions. The ones that can are answering from application logs that were never designed to be evidence — no scoring model attached, no thresholds captured, no ownership chain, no signature. A log is not evidence. It is a story you are asking someone to take on faith, six months after the fact, from a system that has since been redeployed four times.

This is not a modelling problem and it is not a tooling problem. It is a control gap at the moment of the decision, and that is the only place it can be closed.

Pre-deploymentWhere every existing enterprise AI control operates today
RuntimeWhere autonomous behaviour actually diverges from what was approved
Post-incidentWhere most organisations first discover the difference
THE PORTFOLIO
FOUR PLATFORMS

Four business problems, each one created by a system that acts without waiting to be asked.

They share an engine and a discipline: observe the system at the moment it acts, score what it did against what it was supposed to do, and leave behind a record that holds up when someone senior asks about it six months later.

Two of them govern behaviour — VARC in the decision path, SafetySignal in drug safety. Two of them govern the factory that behaviour runs on — InferPulse on AMD Instinct, AgentPulse on NVIDIA. Same question, asked at four different altitudes.

A regulator asks why an agent released a payment. Nobody in the building can answer.

The logs exist. They show that a call was made and a response came back. They do not show what the agent was permitted to do, which thresholds were in force that day, what the scoring model was at the time, or who owned the consequence. Six months later, none of it reconstructs to a standard anyone would put in front of an examiner.

VARC intercepts the action before it executes. It normalises every agent action to a canonical event, scores it across twelve behavioural dimensions, applies a five-level enforcement ladder from observe to decommission, and issues a SHA-256 signed attestation token. The whole chain completes in under fifty milliseconds. The evidence is a by-product of running, not something assembled after an incident.

Read the VARC brief

Behavioural governance for enterprise AI agents

ACTS AT
Pre-execution, in the decision path — verdict in under 50ms
SCORES
12 behavioural dimensions per interaction
ENFORCES
GRO ladder, L0 observe through L4 decommission with IAM write-back
ATTESTS
A-JWT hash chain, independently verifiable without trusting VARC
BUYER
Chief Compliance Officer, Chief Risk Officer, GSI managed-service practice
IN PRODUCTION · GCP MARKETPLACE

You bought an AMD AI factory. Nobody can tell you whether it pays.

Finance sees the invoice. Engineering sees utilisation. Neither can tell you the return on a specific workload, which means the whole estate gets defended as one number or cut as one number. When the pressure comes, the cuts land on whatever is easiest to switch off rather than whatever is least productive.

InferPulse puts one ratio on every workload — the Workload Economic Index, value produced per hour over cost per hour, recomputed every five seconds. Every inference microservice lands in one of four zones: scale, maintain, optimise, or kill. It reads AMD infrastructure natively rather than through a generic GPU abstraction — ROCm counters, AMD-SMI telemetry, Helios rack power, XGMI fabric, KFD process mapping, vLLM latency.

Read the InferPulse brief

InferPulse

inferpulse.live

AMD AI factory intelligence and workload economics

MEASURE
WEI — value score over cost score, per workload, every 5 seconds
ZONES
Scale ≥1.4 · Maintain 1.0–1.4 · Optimise 0.6–1.0 · Kill <0.6
DRIFT
6-dimension CUSUM statistical control, continuous
OUTPUT
Seven board- and CFO-ready reports, no placeholder content
BUYER
CFO, CIO, neocloud operator, AMD factory owner
IN PRODUCTION · PATENT PENDING

Eighty-two per cent of your GPU estate is idle, and the part that is busy may be producing the wrong answer.

The industry has spent roughly half a trillion dollars on GPU infrastructure and uses a fraction of it. On a hundred-thousand-GPU cluster that is billions in idle capital. Worse: a cluster running at eighty-five per cent on an agent that has silently drifted from quality ninety-two to sixty-seven has zero effective utilisation. The spend is real and the outcomes are wrong. Industry mean time to detect that drift is measured in hours.

AgentPulse is the operating system for an NVIDIA AI Factory. Five capability areas — factory economics, output quality assurance, agent intelligence, edge and KPI scheduling, and compliance evidence. It runs as a side-channel observer against NIM, DCGM, OpenShell and Mission Control, adds no latency to the inference path, and generates EU AI Act Article 11, 12 and 14 documentation as a by-product of normal operation.

Read the AgentPulse brief

AgentPulse

agentpulse.live

AI factory operating system for NVIDIA AI Factory

FINDS
Idle capital quantified in real time — the found-capital calculator
DETECTS
Quality drift in under 5 minutes across 8 behavioural dimensions
REMEDIATES
Under 60 seconds autonomously for standard-tier agents; HITL for critical
PROVES
Hash-chained immutable evidence trail, 2-year retention
BUYER
CTO, Head of AI Platform, neocloud and GPUaaS operator
LIVE · NVIDIA NATIVE

A safety signal sat in the case narratives for nine months before anyone saw it.

Pharmacovigilance teams are measured on volume and judged on what they missed. The signal is rarely in the coded fields. It is in the narrative, in the phrasing a reporter used, in a pattern across cases that no individual reviewer sees because no individual reviewer reads all of them. General-purpose assistants make confident guesses here, which is worse than making none.

SafetySignal is an AI-native pharmacovigilance platform built on SafetySignal-PV-7B, a 7.6-billion-parameter model fine-tuned on real adverse event data. Three-algorithm causality consensus — Naranjo, WHO-UMC and Kramer — in under three seconds. Submission-ready ICH E2B(R3) XML in milliseconds. BRAT benefit-risk assessment aligned to ICH E2C(R2). It connects to Oracle Argus Safety, Empirica Topics and Argus Mart rather than replacing them.

Read the SafetySignal brief

SafetySignal

safetysignal.ai

AI-native pharmacovigilance

MODEL
SafetySignal-PV-7B v3 — 7.6B parameters, fine-tuned on real AE data
CAUSALITY
Naranjo, WHO-UMC and Kramer consensus in under 3 seconds
REGULATORY
ICH E2B(R3), GVP Module VI/IX, CIOMS XIV, 21 CFR Part 11
INTEGRATES
Oracle Argus Safety, Empirica Topics, Argus Mart, OCI Kubernetes
BUYER
Head of Pharmacovigilance, QPPV, Chief Medical Officer
IN PRODUCTION · 350+ ENDPOINTS
RUNTIME
WORKED EXAMPLE

One decision, governed end to end, before the agent acts.

An agent proposes an action. VARC normalises it, scores twelve behavioural dimensions, checks provenance, applies the enforcement ladder, and signs the record. Run it and watch what gets written down — including the runs where the action does not happen.

varc · runtime controlsession idle
$ awaiting decision
Illustrative sequence. Not connected to a live tenant.
SUPERVISION
VARC-VERIFY

Supervisors are being asked to examine AI systems using tools built for institutions to grade themselves.

Every AI governance product on the market points the same direction: inward, at the institution, on the institution's terms. There is nothing pointed the other way — nothing an examiner can sit behind and use to reach a conclusion they would be willing to defend in a report of examination.

VARC-VERIFY is that side of the table. Nine questions, the same nine a federal bank examiner will ask, scored against the same evidence standard, producing a Governance Examination Readiness Rating on the 1–5 scale every banker and examiner already knows from CAMELS.

For an institution, the point is to know your GERR before the examiner arrives. For an examiner, the point is a verified assessment without needing access to institutional systems. There is no exam date to prepare for. The exam is ongoing.

Run the self-assessment at varc.live/verify

MODE A

Read the runtime

Connected to a governed estate, the examination pulls attestation records and scores the institution on what its agents actually did.

MODE B

Structured examination

No instrumentation required. A guided sequence producing a comparable score from evidence the institution can put on the table.

RATING

GERR 1–5

Strong through Critical, on the CAMELS convention. A GERR 4 means the finding is already written. You have not seen it yet.

OUTPUT

A citable finding

Each assessment issues a verifiable examination identifier, so a conclusion reached today can be reopened and re-checked a year from now.

WHO BUILT THIS

This was not built by people learning what enterprise risk looks like.

Venture Vertex was founded by Vyasa Murthy after twenty-eight years running enterprise technology businesses — carrying quota, owning P&L, standing in front of boards, and sitting on the other side of the table from exactly the buyers these platforms are built for.

That history is the reason the products start where they do. A governance platform designed by someone who has never had to defend a number to a board tends to optimise for dashboards. These optimise for the moment a senior person is asked a question they cannot answer, because that moment is a familiar one.

Leadership and background

SCALE
Built a cloud consumption business from zero to $87M through global systems integrators
P&L
Owned a $160M portfolio at 35% growth and 145% of quota
PORTFOLIO
Ran a $245M strategic business unit with a 35-person team
CHANNEL
Standing relationships across the major global systems integrators
DELIVERY
Carve-out and integration technology delivery on live M&A transactions
IP
Every platform designed, built and owned in-house

Bring us the system you would not want to be asked about.

A walkthrough runs about forty minutes against a real estate rather than a slide. Tell us what is running, and who has started asking.

VARC · BEHAVIOURAL GOVERNANCE FOR ENTERPRISE AI AGENTS

Your agents are in production. Your governance is not.

VARC is the runtime behavioural governance layer that intercepts every agent action before it executes, scores it across twelve compliance dimensions, and generates cryptographic evidence a regulator can independently verify. The full chain completes in under fifty milliseconds.

THE FORCING FUNCTION

Regulators have arrived. The governance gap is now an examination finding, not a slide.

For three years the AI governance conversation was a strategy conversation. Committees, principles, frameworks, a policy document nobody read. That period is over, and it ended faster than most institutions have adjusted to.

Every major financial regulator now applies model risk governance explicitly to autonomous agents. The exposure is no longer reputational. It is examination exposure, enforcement exposure, and in Europe, balance-sheet exposure.

AuthorityInstrumentWhat it requiresStatus
US Federal ReserveSR 26-2 IV.CModel risk management extended to AI agents in production: pre-deployment validation, continuous monitoring, behavioural drift detection, and examination-ready evidence packages.Effective 2026
European UnionEU AI Act, Articles 9–15Risk management systems, data governance, transparency, human oversight, accuracy and robustness for high-risk systems. Fines to €30M or 6% of global turnover.In force
UK FCAAI governance guidanceDocumented governance chains and audit-ready behavioural evidence for agents touching payments and consumer credit decisions.Active
CFPB & OCCECOA, Fair Housing, UDAPAlgorithmic discrimination enforcement extends to agent decisions in credit, underwriting and servicing — regardless of intent.Enforcement active
HOW IT WORKS
FIVE STAGES

Intercept. Score. Enforce. Validate. Attest.

A five-stage governance chain that completes before the agent acts, not after. Every stage produces a tamper-proof record. The distinction matters more than it sounds: a system that observes and reports is a monitoring product. A system that returns a verdict inside the decision path is a control.

StageEngineWhat happens
01CEM interceptAgent action normalised to a Canonical Event Model. Platform-neutral ingress across Vertex AI, Bedrock, Copilot, Databricks and custom REST.
02BEV scoreTwelve-dimensional behavioural scoring: authority claims, governance bypass, financial actions, harm potential, PII access, deception patterns, cross-agent delegation, RAG provenance, compliance boundary, intent drift, data classification, temporal context.
03GRO enforceFive-level graduated enforcement from L0 observe to L4 decommission. IAM write-back fires at L4, reducing agent permissions in Okta, Azure AD or GCP IAM automatically. No binary allow-or-block.
04RAG validateSource trust registry check, content integrity scan, and data lineage confirmed before the agent reads any external source.
05A-JWT attestSHA-256 signed attestation token on a tamper-proof hash chain. Independently verifiable without trusting VARC.
WORKED EXAMPLE

Watch a decision get governed.

Four scenarios drawn from regulated production patterns. The outcome is not fixed — run it more than once and you will see the enforcement ladder land at different levels, including the runs where the action does not happen at all.

varc · runtime controlsession idle
$ awaiting decision
Illustrative sequence. Not connected to a live tenant.
INSIDE THE LOOP

Boundary tools see the edge. VARC operates inside the loop.

SIEMs, EDRs and prompt firewalls sit at the boundary. They see instructions arrive and results leave. That was a reasonable place to stand when the thing in the middle was a model answering one question at a time.

It is no longer sufficient. The three attack classes OWASP added for agentic systems in 2025 all happen mid-session, inside the execution loop, invisible to anything watching the perimeter. An agent whose objective is corrupted by a poisoned tool response looks perfectly normal from the outside. Each individual turn is clean. The attack only exists in the arc.

VARC scores every instruction, every tool response re-entry, and every session turn against a behavioural baseline. Goal hijacking is caught by multi-turn consistency scoring. Memory corruption is caught by CUSUM drift detection against the session baseline, before the corruption compounds.

ClassVectorHow VARC sees it
LLM07Agentic goal hijackingAgent objective corrupted through adversarial tool responses mid-session. Detected via multi-turn consistency and authority-dimension scoring.
LLM08Session memory corruptionContext poisoned across turns. Detected by CUSUM drift against the session baseline before corruption compounds.
LLM09Embedding inversionTraining data extraction via crafted inputs. Addressed through the information-seeking dimension and RAG source governance.
COGNITIVE LAYER
VCL

VARC now reasons about behaviour. It does not just score it.

Six cognitive capabilities sit above the enforcement engine, and the layer is deliberately model-agnostic — Claude, Gemini, GPT-4o, Mistral, or a local sovereign model, swapped per tenant with no code change. VARC governs Gemini agents and can reason about Gemini behaviour using Gemini. That is not a party trick; it is what lets a regulated tenant keep reasoning inside its own approved model boundary.

ECONOMIC IMPACT

Consequence, not just behaviour

Scores the business consequence of an action before it fires: reversibility, blast radius, regulatory exposure.

PREDICTIVE GRO

Escalation forecast

Forecasts the enforcement level with a confidence score before the behavioural threshold fires, and detects coordinated cross-agent patterns across a sixty-minute window.

GOVERNANCE DEBT

History changes the verdict

Weighted risk accumulated per agent on a thirty-day rolling window. Same prompt, different history, different response. Behavioural intelligence rather than rules.

CHAIN OF THOUGHT

A verdict an examiner can read

Six-step audit-grade reasoning chain per decision. Five steps deterministic from computed signals, the sixth a written synthesis. Under five seconds.

ARCHITECTURE
FOUR TIERS

One governance brain. Policy flows down, evidence flows up, isolation is absolute.

Venture Vertex sets the behavioural floors at the root. A managed-service partner can tighten them for a client portfolio; nobody can loosen them. Evidence aggregates back up the same hierarchy for reporting and attestation. Partner A cannot see Partner B, and no client tenant can see another.

This matters because the realistic delivery route into regulated institutions runs through the systems integrators who already hold the account. A governance layer that cannot be operated as a multi-tenant managed service is a governance layer that has to be sold one institution at a time.

TierWhoHolds
Tier 0Venture VertexRoot policy authority, master policy, behavioural floors, adversarial scenario library, framework registry
Tier 1MSP / GSI partnerClient portfolio governance, billing cascade, domain isolation
Tier 2Client institutionBoundary contract per agent, trust commerce, mandatory pre-production simulation gate
Tier 3AgentThe signed behavioural contract the agent is deployed under
EXAMINATION
VARC-VERIFY

The examination instrument for both sides of the table.

Nine questions. The same nine a federal bank examiner will ask. Scored against the same evidence standard, producing a Governance Examination Readiness Rating on the 1–5 CAMELS convention that every banker and every examiner already understands.

For an institution: know your GERR before the examiner arrives. For an examiner: a verified assessment without needing access to institutional systems. Mapped to IOSCO OR/07/2026, SR 26-2, OSFI B-13, the EU AI Act and FFIEC.

A GERR 4 means the finding is already written. You have not seen it yet.VARC-VERIFY rating scale

Run the nine-question self-assessment

The governance layer your agents need before your next examination.

Production-ready, enterprise multi-tenant, available on Google Cloud Marketplace against existing committed spend.

INFERPULSE · AMD AI FACTORY INTELLIGENCE

Does your AMD AI factory pay?

InferPulse measures the economic output of every inference microservice on your AMD cluster and tells you which ones to scale, fix, or stop. One ratio, recomputed every five seconds, on infrastructure it reads natively rather than through a generic GPU abstraction.

THE PROBLEM

Nobody buys a factory without a P&L. Somehow we did it with GPUs.

If you commissioned a plant, you would know the unit economics of every line in it before the first shift. You would know which lines were profitable, which were marginal, and which were running because someone started them and nobody stopped them.

AI infrastructure got bought differently. It was bought on strategic urgency, in the middle of a capability race, with a business case written at the level of the whole estate. That was defensible in the first year. It is not defensible in the third, when the CFO asks which half of the spend is working and the honest answer is that nobody has ever measured it at that resolution.

The failure mode is predictable. Under pressure, the cuts land on whatever is easiest to switch off rather than whatever is least productive — which means the workloads that get killed are the ones with the weakest internal sponsor, not the ones with the weakest economics.

InvoiceWhat finance can see today
UtilisationWhat engineering can see today
ReturnWhat neither can see, per workload
THE METRIC
WEI

The Workload Economic Index is a ratio, not a dashboard.

WEI is the value a workload produces per hour divided by what it costs per hour. Every microservice gets a WEI score every five seconds. The number integrates behavioural drift, SLA status, GPU cost profile and an industry value model — so a workload that is technically healthy but producing degraded output does not keep scoring as if it were fine.

Below 1.0, that workload loses money every hour it runs. The point of the metric is that it produces an action, not a conversation.

ZoneWEIWhat it meansAction
Scale≥ 1.4Value production significantly exceeds cost. Your most profitable workloads.Increase GPU allocation
Maintain1.0 – 1.4Stable positive economics. Watch for drift pushing it downward.Monitor weekly
Optimise0.6 – 1.0Declining. Behavioural drift or SLA breach is eroding value.Root cause investigation
Kill< 0.6Negative economics. Every hour costs more than it produces.Suspend or reconfigure now
CAPABILITIES

Behavioural drift is an economic signal, not a quality metric.

InferPulse tracks six behavioural dimensions continuously using CUSUM statistical control — the same technique that has run manufacturing process control for sixty years. When a dimension drifts, WEI adjusts in real time, before a human notices the degradation. That is the whole idea: the economics move the moment the behaviour moves, not at the end of the month.

DRILL-DOWN

Universal workload detail

Click any workload reference anywhere — fabric matrix, drift chart, SLA table, leaderboard — and get full economics, behavioural history and value attribution.

HELIOS

Power overlaid on economics

AMD Helios rack-scale power correlated with WEI. Racks in combined crisis — high power, low WEI — surface immediately as the most expensive problem in the factory.

PORTFOLIO

Classification with a narrative

Every workload classified automatically every tick, with a plain-language explanation and dollar figures. The recommendations table says what to do and what it saves.

SCORE

Six weighted dimensions

Behavioural integrity, instruction governance, hardware reliability, inference performance, compliance coverage, session integrity. The lowest dimension tells you where to focus.

AMD NATIVE

Built for AMD infrastructure, not adapted for it.

Generic GPU abstractions lose exactly the information that makes the economics real. InferPulse reads the AMD stack directly, and every feed plugs into the WEI computation rather than into a separate monitoring view.

FeedWhat it contributes
InstinctMI300X, MI325X and MI355X cost profiles, VRAM capacity and compute characteristics native to the WEI calculation
ROCmPlatform performance counters — compute occupancy, memory bandwidth, kernel efficiency — feeding drift scoring
HeliosRack-scale power management overlaid on WEI, with nine rack economic states from optimal to critical waste
AMD-SMIUtilisation, temperature, power draw and VRAM into the drift engine and economic overlay
XGMIInfinity Fabric bandwidth matrix, with per-GPU workload attribution
KFDKernel Fusion Driver process-level GPU assignment for precise cost attribution
vLLM on ROCmTime-to-first-token and per-token latency; SLA breach feeds the economic penalty computation
AITERKernel mode distribution per node, factored into the hardware reliability dimension
OUTPUT

Seven reports, one for each person who will ask.

Every report prints to PDF in one click and contains no placeholder content — only data from the actual factory. The reason there are seven is that the CIO, the CFO, the CISO, compliance and legal all need the same underlying truth expressed in their own vocabulary, and a single dashboard has never once survived that requirement.

ReportForContents
Factory economics board packCIO and boardFactory grade, executive summary, portfolio matrix, priority actions with dollar savings, factory P&L, 30/60/90-day forecast, compliance status
WEI ROI analysisCFOFull return calculation with payback period, net annual ROI, incidents prevented, compliance risk reduction, quantified
Neocloud ROI reportFactory operatorMulti-tenant factory economics with per-tenant WEI, monthly value and contract coverage analysis
EU AI Act compliance certificateCISOPer-workload certifiable statement covering the AI Act, NIST AI RMF, GDPR and ISO 42001, with hash chain summary and signature block
Behavioural audit reportComplianceEight-week WEI and drift history, value attribution by driver, framework coverage, regulatory cost avoided, full audit trail
Conflict investigation reportRisk and legalStep-by-step causal chain from instruction conflict to business impact, affected workloads, resolution path with owners

Bring your GPU count and your workload mix.

Thirty minutes is enough to show you the economics of your own factory. If the answer is that everything is fine, that is a useful answer too.

AGENTPULSE · AI FACTORY OPERATING SYSTEM FOR NVIDIA AI FACTORY

Your AI factory has three unresolved problems.

The NVIDIA AI Factory gives you the infrastructure. What it does not give you is visibility into what that infrastructure is producing, whether it is producing it correctly, and whether you can prove either to a regulator. Three separate instruments, unified in one operating system.

THREE PROBLEMS

This is a working capital problem wearing a technology costume.

Start with the arithmetic, because the arithmetic is not in dispute. A hundred-thousand-GPU cluster represents three to five billion dollars of investment. At the industry-average eighteen per cent utilisation, eighty-two thousand of those GPUs are producing nothing. That is roughly two and a half billion dollars of capital sitting idle, plus the power bill for keeping it available.

Now the second problem, which is worse because it is invisible. A cluster running at eighty-five per cent utilisation on an agent that has silently drifted from quality ninety-two to sixty-seven has zero effective utilisation. The spend is entirely real. The outputs are wrong. And because nothing crashed, nothing alerted.

The third problem arrives on a date. EU AI Act enforcement for high-risk systems requires immutable audit logs under Article 12, human oversight records under Article 14, and technical documentation under Article 11. Most enterprises have none of the three, and the penalty runs to three per cent of global annual turnover.

The unlock is not subtle. Ten points of utilisation on a hundred-thousand-GPU cluster is ten thousand GPUs. Buying that capacity costs roughly three hundred million dollars and nine months of lead time. Recovering it operationally costs a fraction of that and takes weeks.

18%Utilisation — the capital problem
4.2 hrsDrift detection — the quality problem
Article 12Immutable evidence — the governance problem
THE PLATFORM
FIVE AREAS

Five capability areas covering the full operating lifecycle of a production AI factory.

Not five products bundled. Five views of one estate, sharing one behavioural model, because the economics, the quality and the evidence are the same signal read at different resolutions.

AreaWhat it does
Factory economicsReal-time inference spend attribution down to agent, cluster and business unit. Found-capital quantification. Power-aware routing against live grid pricing across eight regions. ROI waterfall with six quantified levers. AI Estate Efficiency score — utilisation multiplied by quality — as the single board-level number.
Output quality assuranceContinuous monitoring across eight behavioural dimensions: semantic coherence, factual grounding, instruction adherence, tool call accuracy, output completeness, context retention, safety adherence, latency consistency. Drift detected in under five minutes. Causal attribution traces it to the exact instruction, tool version or config change.
Agent intelligenceEight autonomous reasoning agents running a continuous observe-hypothesise-query-reason-plan-act-verify-report loop. Prompt injection detected as a behavioural drift signature rather than by content scanning. Instruction genome snapshots let you diff exactly what changed when drift appears.
Edge and KPI schedulingBehavioural quality promoted to a fourth KPI domain alongside application latency, infrastructure utilisation and economic cost. Machine-readable scheduler feed so any orchestrator can route on quality, not just on capacity.
Compliance and evidenceArticle 11, 12 and 14 documentation generated automatically as a by-product of normal operation. Hash-chained immutable audit log with two-year retention. Human oversight records with approver, rationale and outcome. ISO 42001 readiness and SEC AI disclosure evidence.
INTEGRATION

A side-channel observer. Zero changes to the inference path.

Anything that sits in the inference path adds latency, and anything that adds latency gets removed the first time a workload misses its SLA. AgentPulse reads from the NVIDIA infrastructure that is already running and never touches the path itself.

NIM

Microservice telemetry

Native integration with NVIDIA NIM inference microservices. Quality scoring per session, with token and latency telemetry.

DCGM

GPU correlation

Telemetry from the DCGM exporter, correlating GPU metrics to behavioural quality for predictive early warning.

MISSION CONTROL

Where operators already look

Behavioural quality and cost-per-outcome surfaced alongside GPU metrics in the dashboard the team already has open.

SDK

One line to attach

Behavioural monitoring attached to any agent in a single line. Evaluation mode needs no key and no account.

THE FINANCIAL CASE

Six levers, each quantified. Payback measured in weeks, not budget cycles.

The comparison that matters is not against another monitoring tool. It is against buying more hardware. New GPUs cost hundreds of millions and take the better part of a year to land. Recovering the same capacity from an estate you already own takes weeks and costs a rounding error against the idle capital it releases.

Every figure here is arithmetic applied to public GPU pricing and observed industry utilisation, not a benchmark. Run it against your own numbers and the shape does not change.

8 weeksConservative payback period on a 10,000-GPU estate
$300M+Found capital from ten points of utilisation on a 100,000-GPU cluster
$17M/yrPower savings from off-peak scheduling at that scale

Three problems. One operating system.

Twenty minutes, your cluster profile, a live walkthrough of all five capability areas and a custom ROI waterfall built on your numbers.

SAFETYSIGNAL · AI-NATIVE PHARMACOVIGILANCE

The signal was in the narrative. Nobody had time to read it.

SafetySignal automates adverse event detection, causality assessment, regulatory reporting and benefit-risk evaluation — powered by SafetySignal-PV-7B, a 7.6-billion-parameter model fine-tuned on real-world safety data rather than a general assistant asked politely to be careful.

THE PROBLEM

Drug safety is the only function measured on volume and judged on what it missed.

A pharmacovigilance team processes what arrives. Cases come in from spontaneous reports, literature, clinical studies and regulatory feeds, and every one of them has a clock on it. Fifteen days for an expedited report. The team is staffed to the volume, which means it is staffed to process, not to notice.

The problem is that the signal is usually not in the structured fields. It is in the narrative — the phrasing the reporter used, the sequence of events, the detail that a clinician would flag and a coded field cannot hold. And it is often not in any single case at all. It is in a pattern across cases that no individual reviewer sees, because no individual reviewer reads all of them.

This is exactly the kind of work a domain-trained model should do, and exactly the kind of work a general assistant should not be trusted with. In pharmacovigilance a confident wrong answer is worse than no answer, because it consumes the scarce resource — qualified reviewer attention — and returns nothing.

CAPABILITY

Every PV workflow, from intake to submission.

The platform does not ask a team to change how it works. It removes the manual steps between the case arriving and the reviewer having something worth reviewing.

CapabilityWhat it does
AI causality assessmentThree-algorithm consensus using Naranjo, WHO-UMC and Kramer, run in parallel by SafetySignal-PV-7B in under three seconds per case
Signal prioritisationComposite score combining PRR, ROR, IC and EBGM disproportionality with causality strength, clinical severity, case volume and novelty — producing a ranked queue rather than a list
Benefit-risk assessmentBRAT evaluation aligned to ICH E2C(R2), generated from case data in under fifty milliseconds
Labelling gap detectionAutomated listedness determination, cross-referencing signals against product labelling to flag unlisted events requiring fifteen-day expedited reporting
E2B(R3) XML generationSubmission-ready ICSRs with all required fields — demographics, drug detail, causality, narrative — in FDA and EMA formats
Literature surveillanceAutomated portfolio surveillance extracting signals from abstracts and flagging literature-only signals not yet visible in spontaneous reporting data
Signal trendingWeek-over-week emergence tracking across the drug-event pair space, classifying signals as new, rising, declining or stable
Narrative generationICH E2B-compliant case narratives generated from structured data, with completeness scoring and quality grading
ORACLE ECOSYSTEM

It augments Argus. It does not ask you to replace it.

Nobody rips out a validated safety database. The system of record is validated, audited, and connected to every downstream regulatory process the company owns. Any product that begins by asking a sponsor to replace it has already lost the conversation, and deserves to.

SafetySignal connects directly into the Oracle pharmacovigilance ecosystem and enriches what is already there: bidirectional case exchange with Argus Safety over REST and E2B messaging, signal import from Empirica Topics, direct read access to the Argus Mart warehouse, and validated deployment alongside Oracle Cloud Infrastructure for GxP environments.

The positioning is deliberate. Argus holds the record. SafetySignal supplies the intelligence layer above it, for the installed base that already has the record and still cannot find the signal.

ARGUS SAFETY

Bidirectional sync

Real-time case exchange over REST and E2B messaging, enriching Argus cases with generated causality assessments and narratives.

EMPIRICA

Signal import

Detection results imported from Empirica Topics and augmented with composite priority scoring and benefit-risk assessment.

ARGUS MART

Warehouse read

Direct read access to aggregate safety data and case listings, surfaced in the regulatory intelligence dashboard.

OCI

GxP deployment

Validated Terraform and Kubernetes manifests for deployment alongside Oracle Cloud Infrastructure.

STANDARDS

Built to the standards the inspector will cite.

Pharmacovigilance is one of the few domains where the regulatory format is more prescriptive than the analysis. Getting the science right and the submission format wrong is still a finding. The platform is built to both.

ICH E2B(R3)Individual case safety report submission format, FDA and EMA
GVP VI / IXCollection and management of adverse reactions, and signal management
CIOMS XIVAggregate reporting and benefit-risk methodology
21 CFR Part 11Electronic records and signatures

Show us a month of cases you have already closed.

The useful demonstration is not on our data. It is on yours, retrospectively, where you already know the answer and can judge what the platform would have surfaced and when.

ABOUT VENTURE VERTEX

We build the controls that should have shipped with the agents.

Venture Vertex is an independent enterprise AI company based in Dallas. Four platforms in production, every line of the intellectual property designed, built and owned in-house, and a delivery practice that keeps us in front of the same buyers the platforms are built for.

THE STORY

Every technology cycle produces the same gap. The capability ships first, the controls arrive years later, and in the interval a category of company gets built to close the distance. Cloud adoption produced the cloud access security broker. Mobile produced mobile application management. Neither of those categories existed because a vendor invented a need; they existed because enterprises had already deployed something they could not govern, and a regulator eventually asked about it.

Autonomous AI agents are the current instance, and the interval is shorter and sharper than anything that came before. Agents went from demonstration to production inside eighteen months, in credit decisioning, claims adjudication, financial crime monitoring, clinical safety. They went in through platform teams and innovation budgets, mostly without a model risk file, and in many cases without a named owner who could say what any given agent was permitted to do.

Venture Vertex was founded to build the missing half of that deployment. Not a policy framework, not an advisory practice, not a maturity model — software that sits at the moment of the decision and produces evidence as a by-product of the system running.

MISSION
Every consequential decision made by an autonomous system should be reconstructable by someone who was not there.Venture Vertex mission statement

That sentence is deliberately narrow. It does not say that AI should be safe, or fair, or aligned — those are worthy arguments and other people are having them well. It says something more mundane and more enforceable: if a machine took a consequential action, a human being who was not in the room should be able to establish what it did, what it was permitted to do, who was accountable, and whether it would do the same thing again.

That standard is not new. It is the standard every regulated industry already applies to human decisions. Trading has it. Lending has it. Clinical practice has it. The only reason it does not yet apply to autonomous systems is that nobody built the instrumentation, and the instrumentation has to sit at runtime because that is the only place the behaviour actually exists.

HOW WE WORK

Four operating principles, arrived at the hard way.

01

Start with the business problem

Every deliverable begins with the problem in the buyer's language, not the architecture. If the first slide is a diagram, the conversation has already been lost.

02

Evidence, not dashboards

A dashboard tells you something is wrong now. Evidence tells someone else what happened then. The second is much harder and much more valuable.

03

Augment the system of record

Nobody rips out a validated safety database or a working identity platform. We connect to what is running and add the layer that was missing.

04

Build it before selling it

Every platform in the portfolio is in production before it is in a deck. There are no roadmap demos here.

THE DELIVERY PRACTICE

We still do the hard programme work, and it is not a side business.

The platforms came out of large-scale enterprise delivery, and that practice is still running. Separation and integration technology delivery on live M&A transactions. Cloud migration from assessment through cutover. AI governance advisory for organisations that need the outcome before they are ready for a product.

There is a strategic reason to keep it, beyond revenue. Delivery work is how you stay honest. It puts us inside real estates, on real timelines, with real constraints — which is where you find out whether a governance model survives contact with a carve-out at day one, or whether it only works in the demonstration environment.

M&ACarve-out and integration technology delivery on live transactions
CloudAssessment through cutover across the major hyperscalers
AdvisoryBoard and regulator-facing AI governance work

Come and argue with us about it.

The most useful conversations we have start with someone telling us why this will not work in their environment.

LEADERSHIP

Twenty-eight years on the other side of the table.

Venture Vertex is founder-led and deliberately small. The platforms are designed and built in-house, and the person who writes the code is the person who has spent three decades in front of the buyers who will be asked to trust it.

FOUNDER

Vyasa Murthy

FOUNDER AND MANAGING PARTNER

Vyasa Murthy founded Venture Vertex after twenty-eight years running enterprise technology businesses — carrying quota, owning profit and loss, building global systems integrator channels, and standing in front of the boards and risk committees that these platforms are ultimately built to serve.

The commercial record is specific. He built a cloud consumption business from zero to eighty-seven million dollars through the global systems integrator channel. He carried a hundred and sixty million dollar profit and loss at thirty-five per cent growth and a hundred and forty-five per cent of quota. He ran a two hundred and forty-five million dollar strategic business unit with a thirty-five person team. Across HPE, Cognizant, Wipro, Infosys, Compaq and NEC, the through-line is the same: large regulated enterprises, long sales cycles, and technology that has to survive a procurement process and an audit.

That history is the reason the products start where they do. A governance platform designed by someone who has never had to defend a number to a board optimises for what looks impressive in a demonstration. These optimise for the moment a senior person is asked a question they cannot answer — because that moment is a familiar one, and because it is the moment that actually moves a budget.

He is also the sole architect and developer of every platform in the portfolio. VARC, InferPulse, AgentPulse and SafetySignal are designed, written, deployed and operated in-house. There is no outsourced core, no white-labelled engine, and no third party holding a piece of the intellectual property.

He is the author of No Exam Date, a book on the supervision of autonomous AI systems in regulated financial institutions, and he works out of Dallas, Texas.

ROLE
Founder and Managing Partner, Venture Vertex LLC
EXPERIENCE
28+ years in enterprise technology leadership
SCALE
$0 to $87M cloud consumption business built through the GSI channel
P&L
$160M portfolio at 35% growth, 145% of quota
PORTFOLIO
$245M strategic business unit, 35-person team
BACKGROUND
HPE, Cognizant, Wipro, Infosys, Compaq/HP, NEC
CHANNEL
Standing relationships across the major global systems integrators
AUTHOR
No Exam Date — on the supervision of autonomous AI in regulated institutions
BASED
Dallas, Texas
STRUCTURE

Small on purpose, with a partner network that does the scaling.

Venture Vertex is not trying to build a large direct sales organisation. Regulated institutions do not buy runtime governance from a company they have never heard of; they buy it from the systems integrator who already holds the account and already carries the delivery risk.

So the architecture assumes it. VARC runs as a four-tier managed service, with the partner governing a client portfolio underneath a policy floor they cannot loosen. The product is built to be operated by someone else, at scale, without giving them the keys to the intellectual property or visibility into another partner's tenants.

The same logic runs through the other platforms. SafetySignal augments the Oracle safety ecosystem rather than competing with it. InferPulse and AgentPulse sit natively on AMD and NVIDIA infrastructure respectively rather than abstracting over both badly. In every case, the strategy is to be the missing layer in an ecosystem that already has customers, not a new stack asking for a migration.

If you are a systems integrator, this is a channel conversation.

Every client is asking the same governance question. Building a bespoke answer per client is unsustainable and unbillable.

CAREERS

Small team. Consequential problems. No layers.

We are looking for people who have operated inside regulated enterprises and are tired of watching the same failure repeat. Remote-first, senior-weighted, and honest about what an early-stage company is: high ownership, high ambiguity, and direct access to the customer from week one.

WHAT IT IS LIKE

There is no product management layer between you and the problem. You will talk to chief risk officers, examiners, safety physicians and factory operators directly, and what they tell you on Tuesday will change what you build on Wednesday. If you need a roadmap that holds still for two quarters, this will be uncomfortable.

In exchange: the domains are genuinely hard, the regulatory ground is being written in real time, and the work has a shape that is rare — the thing you build is used by people who have to defend it under examination. That is a good discipline. It removes a lot of arguments about what matters.

We hire for judgement over credentials, and we are explicitly interested in people who have spent years inside banks, insurers, pharmaceutical safety organisations or hyperscaler infrastructure teams and know where the bodies are buried. Domain scar tissue is the scarce input here, not framework familiarity.

OPEN ROLES

Open roles

Founding Engineer — Runtime Governance

Own parts of the VARC decision path: interception, behavioural scoring, the enforcement ladder, attestation. Python, FastAPI, PostgreSQL, GCP. You will be writing code that returns a verdict inside fifty milliseconds in a regulated production path, which means correctness and latency are both non-negotiable.

REMOTE · SENIOR

Solutions Architect — Financial Services

Take VARC into banks and insurers alongside the partner channel. You have sat in a model risk committee, you know what an examiner actually asks for, and you can hold a technical conversation and a control conversation in the same meeting.

REMOTE · US / UK

Senior Pharmacovigilance Scientist

Ground SafetySignal in real safety practice. Case assessment, signal management, benefit-risk methodology, and the judgement to tell us when a model output is plausible and wrong. QPPV or equivalent experience welcome.

REMOTE

AI/ML Engineer — Life Sciences

Work on SafetySignal-PV-7B: fine-tuning, evaluation against clinical ground truth, and the unglamorous work of making a domain model behave reliably on the twenty per cent of cases where a general model gets it wrong.

REMOTE

Infrastructure Engineer — AI Factory Telemetry

Own the telemetry paths for InferPulse and AgentPulse. ROCm, AMD-SMI, DCGM, vLLM, Kubernetes. You care about the difference between a metric that is available and a metric that is trustworthy.

REMOTE

Partnerships Lead — Global Systems Integrators

Build the managed-service channel. You have carried a number through a GSI alliance before and you know why most of them produce meetings rather than revenue.

REMOTE · US

Regulatory Affairs Director

Own the mapping between what the platforms produce and what supervisors actually accept as evidence, across SR 26-2, the EU AI Act, IOSCO, OSFI B-13 and ICH.

REMOTE

Nothing that fits? If you have spent a decade somewhere that would have needed this and you can explain precisely why, write to us anyway. Send a note rather than a resume: what you have operated, what broke, and what you would build instead.

CONTACT

Tell us what is running, and who has started asking.

A walkthrough runs about forty minutes against a real estate rather than a slide deck. The most useful version starts with your problem, not our architecture.

HOW TO REACH US

Write directly. Every enquiry reaches the founder, and you will get a substantive reply rather than a routing message.

It helps if your note contains three things: what is running today, which regulator or committee has started asking about it, and what you have already tried. That is usually enough to tell whether we can help, and we would rather say no early than run a discovery process that wastes a quarter.

vyasa.murthy@venture-vertex.com

COMPANY
Venture Vertex LLC
LOCATION
Dallas — Frisco, Texas
EMAIL
vyasa.murthy@venture-vertex.com
INFERPULSE
inferpulse.live
AGENTPULSE
agentpulse.live
SAFETYSIGNAL
safetysignal.ai
BEFORE YOU WRITE
IF YOU ARE AN INSTITUTION

Bring the awkward system

The most productive session starts with the agent you would least like to be asked about, not the one that demos best.

IF YOU ARE AN INTEGRATOR

Bring the repeated question

If every client is asking you the same governance question, that is a channel conversation and it should start at the portfolio level.

IF YOU ARE A SUPERVISOR

Bring the evidence gap

VARC-VERIFY is built for your side of the table. We are interested in what you actually need to reach a defensible conclusion.

CONFIDENTIALITY STATEMENT

What we do with what you show us.

Venture Vertex operates inside regulated environments, on live estates, under examination pressure. This statement sets out how client information, client data and intellectual property are handled. It is written to be read, not to be survived.

EFFECTIVE 2026
VENTURE VERTEX LLC

INSIGHTS

Written from inside the problem.

Notes on runtime governance, AI factory economics, supervision and the commercial reality of selling controls into regulated institutions. No thought leadership, no predictions for next year — arguments we are actually having with customers.

Disagree with something here?

That is the more useful email. Tell us where the argument breaks in your environment.

There is no exam date. The exam is ongoing.

Institutions are preparing for an AI examination as though it were an event on a calendar. It is not. The evidence either exists at the moment the agent acts, or it does not exist at all.

I have sat through a lot of examination preparation over the years, and it always has the same rhythm. A date appears. A programme forms. Somebody builds a tracker. For eight weeks, a group of capable people assemble artefacts, reconcile them against a request list, and rehearse the answers. The examiner arrives, works through the file, and issues findings. The programme disbands. The tracker goes stale.

That rhythm works when the thing being examined holds still between examinations. A lending policy holds still. A capital model holds still, more or less, and when it changes there is a change record with a date on it and a person who signed it.

An autonomous agent does not hold still, and this is the part that has not landed yet.

The artefact you validated is not the thing that is running

An agent chooses its own route through the tools available to it. It calls other agents. It reads sources that were not in scope when it was reviewed. Its behaviour on Thursday can differ materially from its behaviour on Tuesday with no code change, no release, and no ticket to point at — because a tool returned something different, or the context window filled differently, or an upstream agent phrased a handoff another way.

So when an examiner asks what the agent did in March, the honest answer in most institutions today is a reconstruction: here is the application log, here is the prompt template we believe was in force, here is the model version we think was deployed, here is a person who remembers approving something adjacent. Every step of that chain is an inference, and every inference is a place where the examination goes badly.

A log is not evidence

This distinction is worth being precise about, because a lot of programmes are being run on the assumption that logging is the same as evidence.

A log records that something happened. Evidence establishes what was permitted, what was applied, what was decided, and by whose authority — in a form that a party who was not present can verify without taking your word for it.

Concretely, an evidentiary record of an agent decision needs to carry, at minimum:

  • The action proposed, normalised to a form that does not depend on which platform the agent was running on.
  • The behavioural contract in force for that agent at that moment, identified by revision.
  • The scoring model applied, identified by version — because a score of 0.72 means nothing without knowing what produced it.
  • The thresholds in force at that timestamp, snapshotted rather than looked up later.
  • The enforcement decision and the level at which it landed.
  • The accountable owner, resolved at the time, not reconstructed from an org chart that has since changed.
  • A signature over all of it, on a chain that breaks if anything is modified afterwards.

None of that can be added retrospectively. Every one of those fields is only available at the moment the decision is made. This is why runtime is not a preference or an architectural style — it is the only point in the lifecycle where the evidence physically exists.

What the supervisors have actually said

The regulatory picture stopped being ambiguous some time ago. The Federal Reserve extended model risk governance explicitly to AI agents operating in production, with continuous monitoring, behavioural drift detection and examination-ready evidence packages named as requirements rather than aspirations. The EU AI Act put Articles 9 through 15 in force for high-risk systems, with penalties calibrated to global turnover. The FCA published guidance for agents touching payments and consumer credit. The CFPB and OCC made clear that fair lending law applies to an agent's decision regardless of whether anyone intended the outcome.

And IOSCO published a supervisory toolkit for AI examination — a document, not software. Which is the more interesting fact, because it tells you where the supervisory community actually is: they have agreed what should be examined, and they do not yet have the instrument to examine it.

The uncomfortable implication

If you accept the argument to this point, one conclusion follows that most governance programmes are not structured to handle.

An institution that instruments its agents today has evidence starting today. It does not have evidence for last quarter, and no amount of programme effort will produce it. The window for the period you are currently operating in closes continuously, every hour, whether or not anyone has approved a budget.

That is what the phrase means. There is no date to prepare for, because the record is being written or not written right now. When the examination does arrive, it will look backwards at a period during which you either had the instrumentation or you did not.

A rating of four on a governance readiness scale does not mean you are going to receive a finding. It means the finding is already written. You have not seen it yet.

Vyasa Murthy is the founder of Venture Vertex LLC and the author of No Exam Date.

Bring us the system you would not want to be asked about.

Eighty-two per cent idle, and the busy part might be wrong

The GPU utilisation problem is a working capital problem wearing a technology costume. The quality problem underneath it is worse, because a cluster producing wrong answers at high utilisation looks healthy on every dashboard you own.

Nobody commissions a manufacturing plant without knowing the unit economics of every line in it. You would know which lines were profitable, which were marginal, and which were running because somebody started them and nobody ever stopped them. That is not sophisticated financial management. It is the minimum condition for operating a capital asset.

AI infrastructure got bought differently, and for understandable reasons. It was bought fast, on strategic urgency, in the middle of a capability race, with a business case written at the level of the entire estate. That was defensible in year one. In year three, when the CFO asks which half of the spend is working, the honest answer in most organisations is that nobody has ever measured it at that resolution.

Start with the arithmetic

Take a hundred-thousand-GPU cluster. At current acquisition costs that is three to five billion dollars of investment. At the industry-average utilisation figure that keeps showing up in operator data — around eighteen per cent — eighty-two thousand of those GPUs are producing nothing at any given moment. Call it two and a half billion dollars of idle capital, before you count the power bill for keeping it available.

This is not a benchmark or a projection. It is arithmetic applied to public pricing and observed utilisation. Run it against your own estate and the shape does not change, only the magnitude.

The leverage point is what makes it interesting. Ten points of utilisation on that cluster is ten thousand GPUs. Buying that capacity costs roughly three hundred million dollars and takes the better part of a year to land. Recovering the same capacity from hardware you already own is an operational exercise measured in weeks. The gap between those two options is the entire argument.

The second problem is worse because it is invisible

Now consider a cluster running at eighty-five per cent utilisation — a number any operator would be pleased with — on an agent whose output quality has silently drifted from ninety-two to sixty-seven.

The effective utilisation of that cluster is zero. The spend is entirely real. The outputs are wrong. And because nothing crashed and no threshold was breached, no alert fired. Industry mean time to detect that kind of drift is measured in hours, and the honest version is that a lot of it is detected by a customer complaint rather than by an instrument.

This is why utilisation on its own is a misleading metric, and why the number that actually matters is utilisation multiplied by quality. A busy cluster producing degraded output is more expensive than an idle one, because you are paying for the compute and then paying again to fix what it produced.

Behavioural drift is an economic signal

Here is the reframe that changes how the problem gets managed. Drift is usually treated as a quality issue, which routes it to an engineering team and a backlog. Treated as an economic signal, it routes to the person who owns the budget, and it moves.

The mechanism is straightforward once you accept the framing. Track behavioural dimensions continuously using statistical process control — CUSUM has run manufacturing quality for sixty years and works perfectly well here. When a dimension drifts, adjust the workload's economic score in real time. A workload that is technically healthy but producing degraded output stops scoring as if it were fine, immediately, without waiting for a human to notice.

Express the result as a single ratio: value produced per hour over cost per hour. Below one, that workload loses money every hour it runs. Above 1.4, it is your best use of capital. The point of the ratio is that it produces an action rather than a conversation.

Why the cuts always land in the wrong place

The predictable failure mode, when the budget pressure finally arrives, is that the cuts land on whatever is easiest to switch off rather than whatever is least productive. Which in practice means the workloads killed are the ones with the weakest internal sponsor, not the ones with the weakest economics.

That is not a failure of will. It is a failure of instrumentation. If you cannot rank workloads by return, you will rank them by politics, because politics is the only ranking available.

Give an organisation a per-workload economic score and the conversation changes shape entirely. It stops being a negotiation about whose project survives and becomes an operational review of a portfolio. That is a much better meeting, and it is the meeting the plant manager has been having for a hundred years.

Vyasa Murthy is the founder of Venture Vertex LLC and the author of No Exam Date.

Bring us the system you would not want to be asked about.

What a behavioural contract actually says

Policy documents describe intent. A behavioural contract is the machine-enforceable version: what this agent may do, scored across twelve dimensions, with a graduated response and a signature over the result.

Most enterprise AI governance today lives in a document. The document says the right things: agents shall operate within approved scope, sensitive data shall be protected, human oversight shall be maintained for consequential decisions. It is reviewed by a committee, approved, published, and then plays no part whatsoever in what happens at three in the morning when an agent decides to release a payment.

The gap between an approved policy and an enforced control is the whole problem. Closing it means expressing governance intent in a form a machine can evaluate in single-digit milliseconds, and doing it without collapsing into a rules engine that anyone with a novel phrasing can walk around.

Twelve dimensions, not one score

The instinct is to produce a single safety number. Resist it. A composite score with no structure tells an operator that something is wrong and nothing about what.

The dimensions that matter in regulated environments, in our experience, are these: authority claims, governance bypass, financial action, harm potential, personal data access, deception patterns, cross-agent delegation, retrieval provenance, compliance boundary, intent drift, data classification, and temporal context.

Score each of those independently and two useful things happen. First, the verdict becomes explicable — not "risk 0.78" but "authority claims elevated to 0.91 while provenance fell to 0.34," which a compliance officer can read and act on. Second, the dimensions can carry different weights in different industries, because what constitutes an escalation in banking is genuinely different from healthcare or defence.

The scoring is not the interesting part. The ladder is.

Binary allow-or-block is the single most common design mistake in this space, and it fails in both directions. Set the threshold tight and you block legitimate work until an operator disables the control. Set it loose and you have a monitoring product wearing an enforcement badge.

A graduated ladder solves this. Five levels, roughly: observe, flag, human-in-the-loop, block or freeze, and decommission — with the top level writing back to the identity platform to reduce the agent's permissions automatically. The important property is that most decisions land in the lower levels, which means the control stays switched on, which means it is there on the day it matters.

The other important property is that the response can account for history. The same prompt from an agent with no prior violations and an agent that has accumulated governance debt over a thirty-day window should not receive the same treatment. That is behavioural intelligence rather than rule matching, and it is the difference between a control that adapts and one that gets gamed.

What gets written down

Every verdict produces a record, and the contents of that record are what determine whether any of this survives an examination.

  • The scorer version. A score is meaningless without knowing what produced it. Pin the engine version to the record.
  • The configuration hash. Governance configuration changes. The record must say which configuration was in force.
  • The threshold snapshot. Thresholds move as an institution tunes its posture. Capture them at decision time rather than looking them up later.
  • The effective-from timestamp. So the configuration in force on any historical date can be established without archaeology.
  • A signature over all of it, on a hash chain that breaks if anything is altered.

That last property is the one that matters most and is most often skipped. The record has to be independently verifiable — a third party should be able to confirm it has not been modified without trusting the vendor that produced it. A governance record you have to take a vendor's word for is not evidence. It is marketing with a timestamp.

Why the boundary is the wrong place to stand

Existing security tooling sits at the perimeter and watches instructions arrive and results leave. That was a reasonable place to stand when the thing in the middle answered one question at a time.

It is no longer sufficient, and the reason is structural rather than a matter of vendor capability. The attack classes that matter for agentic systems now happen mid-session, inside the execution loop. An agent whose objective has been corrupted by a poisoned tool response looks entirely normal from outside: every individual turn is clean. The attack exists only in the arc across turns.

Catching it requires scoring every instruction, every tool response as it re-enters the loop, and every turn against a session baseline. That is an inside-the-loop position. You cannot get there from the perimeter, no matter how good the perimeter product is.

The test

Here is the question I would ask of any runtime governance design, including ours. Take a consequential decision an agent made ninety days ago. Can a person who was not present establish what was proposed, what was permitted, what was applied, what was decided, who was accountable, and whether the record has been altered — without asking anyone to trust the vendor?

If the answer is yes, the design is sound. If the answer requires a person to vouch for a step, it is a monitoring product, and it will be found out at the worst possible moment.

Vyasa Murthy is the founder of Venture Vertex LLC and the author of No Exam Date.

Bring us the system you would not want to be asked about.

Why we did not try to replace the system of record

Every enterprise software category eventually produces a challenger that asks the customer to migrate. In validated environments, that pitch is not brave. It is a misunderstanding of what the customer is protecting.

When we started building in drug safety, the obvious commercial move was to position against the incumbent safety database. The incumbent is expensive, the workflows are dated, and there is a genuine capability gap on the intelligence side. Every instinct trained by a decade of enterprise software says: build the modern replacement and go take the seat.

We did not, and the reasoning generalises well beyond pharmacovigilance.

What the customer is actually protecting

A validated safety database is not just software. It is a validated system in the regulatory sense — qualified, documented, audited, connected to every downstream submission process the company owns, and named in inspection responses going back years. Replacing it is not a migration project. It is a revalidation programme with regulatory exposure attached, run by a team whose day job is already oversubscribed.

So when a vendor opens with "replace your system of record," the head of pharmacovigilance is not hearing a product pitch. They are hearing a proposal to take on institutional risk in exchange for features. That conversation ends politely and quickly, and it should.

The same structure appears everywhere in regulated enterprise. The core banking platform. The identity provider. The general ledger. The claims system. These are not chosen on merit any more; they are load-bearing, and the cost of replacement is dominated by revalidation and regulatory exposure rather than by licence fees.

The layer that is actually missing

Once you stop trying to take the seat, the real gap becomes visible. The incumbent holds the record extremely well. What it does not do is read the narrative, run causality assessment in seconds rather than hours, rank a signal queue by anything other than volume, or produce a benefit-risk assessment from case data.

That is an intelligence layer, and it sits above the record rather than in place of it. It reads from the system of record, enriches what it finds, and writes back in the formats the record already understands. The customer keeps the validated system, keeps the inspection history, keeps the downstream integrations, and gets the capability they were missing.

The commercial consequence is significant. The addressable population is not "customers willing to migrate," which is a small and reluctant group. It is "customers who already own the incumbent," which is the entire installed base.

The same logic across the portfolio

This is not a pharmacovigilance-specific insight, and once we saw it clearly, it shaped everything else.

  • Runtime governance for AI agents does not ask an institution to change its agent platform. It normalises actions from whatever platform is running and governs them from the decision path.
  • Factory intelligence does not ask an operator to change their scheduler or their GPU vendor. It reads existing telemetry as a side-channel observer and never touches the inference path, because anything that adds latency gets removed the first time a workload misses an SLA.
  • Examination tooling does not require the institution to be instrumented at all. Where a governed estate exists, it reads the evidence. Where it does not, it runs as a structured examination and still produces a comparable score.

The discipline this imposes

Choosing the augmentation position is easy to say and harder to hold, because it forces you to be genuinely good at integration rather than merely adequate. You inherit someone else's data model, someone else's API, someone else's release schedule, and someone else's edge cases. None of that is glamorous and all of it is real work.

It also forces honesty about where the value actually is. If your product only wins because the customer had to throw away their incumbent to use it, you never find out whether the capability was worth anything on its own merits. Sitting on top of a working system of record is a harsher test, and passing it is a much better signal.

Vyasa Murthy is the founder of Venture Vertex LLC and the author of No Exam Date.

Bring us the system you would not want to be asked about.