Product Engineer · New York City

I build AI products for experts, and prove they work.

Princeton computer scientist and co-founder of Verata (YC W23), a people-and-company data platform for private-equity talent teams and executive-search firms.

I led product, design, sales, and our data and AI research, and personally built our AI sourcing and evaluation systems. Now I'm looking for my next product engineering role in New York.

Portrait of Josh Gardner
JGListen → Build → Measure
01

Owned the outcome

Took Verata from 100+ discovery interviews to 30+ paying customers across PE and search firms, including two of the ten largest PE firms. Core subscription ARR roughly doubled, ~$52K → ~$110K (Jun 2025 → Sep 2026).

02

Shipped at customer speed

Customer requests shipped in 1–2 weeks. I designed a 103-screen applicant-tracking product in 2 days, and it went from decision to launch in 8 weeks.

03

Built our AI sourcing systems myself

Solo-built two AI systems in about 3 months (~100K lines, 726 commits, 2,000+ tests). One turns a role brief into a sourced executive shortlist for a few dollars per search.

04

Measured honestly

Showed that the platform's legacy candidate scorer ranked a client's real hire decisions backwards (AUC 0.178, n=14). My corrected system scored 0.740 on the same labels.

Case study 01

Verata

  • Co-founder & CTO
  • YC W23
  • 2022 — present

A two-founder company that gives PE talent partners and executive recruiters partner-level market knowledge: executive profiles with transaction timelines, revenue estimates with visible comps, backchannel reference finding, network mapping and an applicant-tracking system.

Discovery → product

  • Started in 2023 building AI due diligence for PE, then killed it when the models weren't reliable enough for the job. A customer's request for executive and network search became the product.
  • Ran 100+ discovery interviews with PE firms and recruiters. The core insight: recruiters stitch PitchBook and LinkedIn together by hand.
  • Tested prices from $20 to $1,000 per seat; launched at $150/seat/month.
  • Cut features on evidence: hid an AI market-mapping feature that caused the most confusion and replaced it with deterministic, explainable tags.

Design & delivery

  • Designed the product in Figma and specced customer requests; my co-founder implemented them, usually within 1–2 weeks.
  • A portfolio-management module, designed live on a customer call, shipped in about 12 days and closed an annual contract.
  • Designed a 103-screen applicant-tracking product (40 conventional + 63 AI-native screens) in 2 days; it launched 8 weeks after the decision to build.

Customers

  • Sold as the founder-seller: 30+ paying customers across PE and search firms, including two of the ten largest PE firms. Warburg Pincus bought after a head-to-head evaluation; the firm uses both tools.
  • Ran a $72K-annualized paid pilot at a top-5 global search firm: 40 licenses, 37 users trained in 6 sessions, weekly usage reports.
  • Took a two-person company through enterprise procurement at a top-10 PE firm.
  • About $1M raised; cash-flow neutral while paying both founders.
  • “This is the best product I've ever used.”

    Talent partner, mid-market PE firm

  • “I can find the same people, if not better ones, in Verata within a couple of hours, compared with recent graduates working LinkedIn Recruiter for days.”

    Head of talent, growth-equity firm

  • “Every result is relevant.”

    Talent team, global PE firm

Case study 02

AI sourcing & evaluation systems

  • Solo build
  • Python + TypeScript
  • Jun — Sep 2026

In 2026 I wrote two AI systems end to end. One turns a company and a role ("CEO of a paper company") into a ranked, evidenced shortlist of real executives. The other rebuilds a dated timeline of what a company did from its employees' role descriptions, with every fact linked to a verbatim quote.

lines of code
~100K
commits
726
tests
2,000+
on a live CEO search
$7.21
The sourcing pipeline
  1. 01ResolveTarget company, cited facts only
  2. 02Find peersExact vector search over 1.53M companies
  3. 03ProfileMerge only corroborated facts
  4. 04RankAnonymized LLM tournament, groups of 8
  5. 05Verify12-class problem check: flags, never silently drops
  6. 06ApproveHuman gate before anything ships
  7. 07DeliverClient-ready evidence report

Engineering

  • A shared model gateway: schema-validated outputs with validate-and-repair, a content-addressed cache, and per-call cost accounting.
  • Ported the extraction engine from Python to TypeScript (~55K lines, 1,080 test cases) with a parity check.
  • Found that 92% of verifier input was one repeated ~1,860-token prompt, so prompt caching was the main cost lever.
  • Production model calls ran on Azure OpenAI.

Product judgment

  • Built after roughly 90% of customers and prospects asked for agentic search (my estimate). An earlier natural-language search had been removed because users asked for data no system holds.
  • Human-approval gates between every stage; database-written and model-written content flagged separately in the report.
  • Wrote into the handoff that the system is not a recall lift: its wins are reproducibility, ranking stability and groundedness.

On live searches

  • Ran the AI sourcing pipeline and wrote the 2,590-word search spec for a live PE-backed CEO search: 255 evidence cards and a 23-page client report for $7.21.
  • Directed three coding agents (Claude Code and Codex) under a written coordination protocol; 579 of my 726 commits were co-authored with Claude Code.

Case study 03

Evaluation as a product discipline

  • Pre-registered
  • Blinded
  • Budget-capped

I write the test plan first, cap the budget, and report the unflattering answer. Most quality wins are unglamorous, and the worst outcome is shipping a confident number that's wrong.

AUC against client hire decisionsLegacy scorer AUC 0.178; corrected system AUC 0.74 (95% confidence interval 0.622 to 0.846); n = 14. 0.5 is a coin flip.00.250.50.751.0Legacy scorer: AUC 0.178Legacy scorer0.178 · ranked backwardsMy corrected system: AUC 0.74My corrected system0.740 · 95% CI 0.622–0.8460.5 = coin flip
Ranking quality against a client's real hire / no-hire decisions on the same 14 candidates. 0.5 is a coin flip; below 0.5 means the ranking is backwards. n = 14
  • 0.855 → 0.656Recall after I rebuilt the benchmark's ground truth. The old one hid misses.
  • 16 for $34Pre-registered experiments, run against a $150 budget cap.
  • 12–0A model tournament beat the weighted composite in blinded panels, at ~$3 vs ~$166 per slate (research branch, not merged).
  • 31 → 0Attribution violations after one prompt phrasing change. A 7.8× cheaper model was rejected on quality.
  • 5 / 5Planted error types caught by the hallucination gate in a smoke test, with 0% false flags.

I'd rather report the honest number.

Verata × HSiQ · Feb 2026

From Pedigree to Performance

Does CEO background predict PE exit success?

Lead author

A study of 12,174 PE-backed CEO appointments (2000–2018), co-published with Hunt Scanlon's research arm. About 23 primary tests plus robustness checks, with false-discovery-rate control: 4 of 23 traits survive, and career traits explain under 1% of exit variance. Red-teamed and corrected before release.

Read the study
SHAP feature-importance chart: appointment timing (0.222) dominates every CEO-level trait; the model relies more on when a CEO was hired than on who they are.
What actually drives the model: when a CEO was hired matters more than who they are. Source: From Pedigree to Performance, Verata × HSiQ, Feb 2026.

WINE 2023 · arXiv

Optimal Stopping with Multi-Dimensional Comparative Loss Aversion

Co-author with Linda Cai and S. Matthew Weinberg

Grew out of my Princeton M.S.E. thesis in algorithmic game theory.

arXiv 2309.14555

Education

  • 2022

    M.S.E., Computer Science

    Princeton University

    GPA 3.883. Graduate coursework in NLP, systems & ML, computer vision and algorithms. Thesis became a WINE 2023 paper.

  • 2020

    B.S.E., Computer Science

    Princeton University

    Four years of Mandarin, including a summer in Beijing.

Experience

  • 2022 — present

    Co-founder & CTO

    Verata (AiFlow Inc.), YC W23

  • 2020 — 2022

    Assistant in Instruction

    Princeton Computer Science

    Taught full-stack development and algorithmic game theory.

  • 2020 — 2021

    Presales engineer intern

    Laiye (AI & automation)

Technical toolkit

Languages

  • TypeScript
  • Python
  • SQL

Product & UI

  • React
  • Next.js
  • Vite
  • Figma
  • Customer discovery
  • Pricing

AI & evaluation

  • LLM pipelines
  • Embeddings & vector search
  • Structured outputs
  • LLM-judge panels
  • Pre-registered evals
  • Cost modeling

Data & ML

  • Pandas
  • LightGBM / SHAP
  • Survival & causal stats
  • Entity resolution
  • OpenSearch
  • Pinecone

Off the clock, I go looking for the hidden city.

I'm a New York urban explorer. My favorite weekends go deep into outer-borough enclaves, to the corners of the city most people never see and the neighborhoods that don't make the guidebooks.

At home it's a full house: a giant, very fluffy Great Pyrenees mix, a cat who runs the place, and a large aquarium I'm always tinkering with. I also spent years walking dogs at NYC animal shelters.

Let's build something experts trust.

I'm looking for founding, forward-deployed and product engineering roles in New York at teams building AI for demanding users. In person.