Measure what good search actually means.
Mohamed A M Elansary, PhD — evaluation frameworks, measurement under uncertainty, multi-model forecast evaluation, and production agent evaluation sets for search quality in an LLM world.
Scientific measurement
- Designed multi-model forecast comparisons across basins and hydroclimates.
- Quantified uncertainty and validated imperfect USGS, NOAA, and NASA observations.
- Ran reproducible Python, R, Bash, Linux, and HPC workflows.
Production systems
- Builds GPT, Claude, and Gemini agent workflows at Vertexium.
- Maintains regression evaluation sets with retrieval and query routing.
- Ships tenant isolation, provenance, and validation systems.
Proposed evaluation approach
Define intended search behavior for developers and agents and a small failure taxonomy; build a golden/synthetic/agentic evaluation set with provenance and ambiguity labels; implement statistical analysis, stratification, and uncertainty; compare simple baselines; and report what the signal does and does not support before embedding it in Exa’s research feedback loop. This is a proposed measurement approach, not a claim of prior Exa-internal search evals.
Honest fit boundary
Exa search-product eval stack ownership and retrieval-benchmark authorship are a stretch. I have not authored Exa-internal search evals or published Exa-scale retrieval leaderboards. I do not invent search metrics, safety research, or RLHF. My contribution is evaluation frameworks, measurement under uncertainty, multi-model forecast evaluation, and production agent regression evaluation.
Role and location
Research, Evals · San Francisco, California · OnSite. The posting states this is an in-person opportunity in San Francisco. Relocation with a support package is an honest discussion point; remote eligibility is not asserted.
“$180K – $350K • Offers Equity” · Official role posting