- Define and execute the firm-wide AI and data strategy, embedding GenAI, agentic systems, and cloud-native analytics across the entire venture lifecycle: sourcing, due diligence, and portfolio support.
- Architected and shipped an end-to-end agentic platform: it ingests unstructured pitch decks, uses LLMs (OpenAI and Claude via AWS Bedrock and AgentCore) to extract entities into DynamoDB, agentically enriches profiles against firm criteria in Snowflake, and auto-generates ranked tear sheets and full investment due-diligence memos delivered through Outlook and Slack.
- Architected an MCP-exposed knowledge graph over a 40k+ document corpus, routing local (Qwen/Ollama) and frontier (Claude) models by task for cost and data-privacy control.
- Recruit, budget for, and lead a multi-national (US/UK/India) cross-functional team of engineers, data scientists, and statisticians, shifting investment decisions from fully manual to automated and evidence-led.
My current role sits where AI strategy meets hands-on engineering: I set the
roadmap and standards for responsible AI, and I still architect and ship the
platforms myself. The work spans applied ML and decision science.
Propensity-score acquisition-prediction models. Time-series analysis of market
downturn predictors. Publication-driven signal detection for emerging biotech
growth areas. All of it on AWS and Snowflake, with data governance and model
evaluation built in.
- Engineered a scalable pharmacovigilance pipeline (PhD dissertation) mining massive unstructured social-media corpora with transformer models and weak-supervision frameworks (Prodigy, SpaCy), architecting custom programmatic labeling and fine-tuning workflows before the commercial LLM paradigm.
- Developed and validated statistical and ML methods for clinical and regulatory decision-making, disseminated through peer-reviewed publications and national conference talks.
- Delivered analyses and dashboards supporting federal broadband initiatives, translating high-volume data into actionable policy with BigQuery, Python, and Tableau.
This is where the AI-engineering foundation was built. The work was
pre-generative frontier NLP: synthesizing novel, high-fidelity clinical and
epidemiological datasets out of noisy social-media text. That turned out to be
exactly the groundwork production LLM and RAG systems would later require.
- Led cross-functional teams of researchers, engineers, statisticians, clinicians, and regulatory experts through complex drug-development programs, including animal trials and first-in-human clinical trials.
- Directed the design and execution of causal studies for drug efficacy and safety, owning statistical analysis plans, study protocols, and SOPs aligned to FDA standards.
- Built the firm's first quality-management program and supported QA/QC of cGMP batches used in human studies.
- Managed large multi-national, cross-functional teams running Phase I clinical trials under strict timelines and federal regulations, overseeing protocol development, statistical analysis plans, and SAS code review.
- Harmonized SOPs for Phase I project and program management; mentored clinical coordinators through promotion to project manager.
- Led public-health research and statistical evaluations on immunizations, early-childhood mental health, and rural health.
- Produced Health Impact Assessments used by state legislators in policy decisions.
- Presented findings directly to state legislators, agency staff, and community audiences, translating statistical evidence into terms each room could act on.