A Python Model Context Protocol (MCP) suite of life-sciences data servers — ClinicalTrials.gov, openFDA (labels/events/recalls/approvals), and PubMed — with pgvector hybrid semantic search, cited LLM tools, cross-server joins, incremental refresh, usage metering, and an API-key/tier authorization gate.
A governed AI tool hub that removes the prompt — 13 cited-or-abstain tools on one no-prompt shell, safe on sensitive data (PII/PHI sanitized before egress), with a measured 20/20 trust eval (0 hallucinations). Live demo included.
Test LLM systems like software — grade against golden + adversarial sets (LLM-as-judge + deterministic rule checks), gate regressions in CI, and filter PII/hallucinations.
AI flags likely payment-integrity issues and explains each in plain English; a human approves or dismisses (never auto-decides), with a dashboard tallying dollars identified.
Structured intake + LLM scoring of AI use cases across five dimensions — flags where AI is a poor fit and exports a prioritized portfolio. An AI-adoption / prioritization tool.
An AI agent that plans a goal into ordered steps and runs them through registered tools — with step/cost/latency guardrails, a human checkpoint before anything finalizes, and a full audit log.