I build AI systems
that can be proven,
not just trusted.
I audit RAG and agent pipelines for faithfulness, hallucination, and regulatory compliance. I show my work against public documents so you can check it yourself. Twenty-six years of FICO Blaze and Drools in regulated industries where a wrong rule was a compliance incident, not a bug. That discipline is what makes LLM evaluation possible to do rigorously.
The AI Pipeline Audit
A fixed-scope engagement for teams that have shipped a RAG or agent system and need an independent, audit-grade evaluation before a compliance review, a high-stakes demo, or a regulated production launch.
Structured Evaluation
Three automated test suites: extraction accuracy, RAG faithfulness (DeepEval and Ragas), and adversarial resistance. Every score traced back to source data, not a summary judgment.
Teams That Have Shipped AI
You've deployed a RAG or agent system and have no independent way to answer "how do you know it works?" Especially true before a compliance review, a high-stakes client demo, or a regulated launch.
Audit-Grade Deliverables
A Scorecard Delta showing before and after scores, plus a prioritized fix roadmap. PDF for compliance teams, HTML dashboard for internal review, JSON for your CI/CD pipeline.
SERFF Hallucination Audit: Scorecard Delta Demo
Don't take my word for it. A live demonstration applied to P&C insurance SERFF filings: baseline Faithfulness 0.23 (hallucinated state-specific clause) vs. post-fix 0.97 after Hybrid Search and Cross-Encoder Re-ranking. Every step documented in an audit packet that maps directly to NAIC Exhibit C.
A redacted sample report will be published here shortly. Reach out directly if you'd like to see it before the public release.
Active Work
Novus Forge
AI governance infrastructure for regulated industries. Every agent call is intercepted for PII scanning, policy evaluation, and audit logging before a decision leaves the pipeline. The entry point is Inspector: three automated evaluation suites that prove your AI is faithful to source documents, adversarially resilient, and compliant with regulatory axioms before a single production decision is made.
Novus Inspector
Standalone AI evaluation lab. Three automated suites test extraction accuracy, RAG faithfulness, and adversarial resilience. Audit-ready PDF, HTML, and JSON reports every session.
Centenarian OS
Privacy-first longevity tracking for iOS. Designed for people serious about the science of healthspan. No subscription, no data harvesting.
From the Practitioner's Desk
Technical articles on AI governance, evaluation, and decision intelligence.
25 Years Building Systems That Have to Be Right
My career started in business rule management, at a time when explainability and auditability were the default requirement, not an afterthought. Decades of FICO Blaze Advisor, FICO Decision Modeler, Apache Drools, JESS, and OFBiz taught me that the logic driving enterprise decisions has to be traceable, testable, and defensible.
When LLMs arrived, I saw the same problem repeating: powerful capability, insufficient governance. Novus Forge is the answer I built. Inspector is how I help teams find out what their AI is actually doing before a regulator does.
Years in business rule management and enterprise decision systems
Industries including insurance, healthcare, finance, and logistics
FICO Blaze Advisor, FICO Decision Modeler, Apache Drools, JESS, and OFBiz: production implementations across regulated industries
Outside of work: currently deep in longevity science (hence Centenarian OS), and I spend time writing, photographing, and building things that have no immediate business case.
If Your AI Has Never Been Independently Evaluated, Let's Fix That.
Book a 20-minute call to scope an AI Pipeline Audit. If your system has shipped but never been evaluated by someone outside your own team, that gap is fixable and provably so.