Kayvan Zahiri

Kayvan Zahiri

AI engineer in San Francisco. Founder of ResumeAI.

ParkCast SFParking predictions for 12.7K SF blocks ATS share, from ResumeAI data Merged openai/whisper #2836 Fix English number normalizer

Work

ResumeAI editor with a live resume preview
Example ResumeAI score card reading 92
Bar chart of applicant tracking system share among large employers ATS share, from ResumeAI data

Founder. Built solo.

ResumeAI

AI-powered ATS resume optimizer: ATS compatibility scoring, Claude-powered bullet rewriting, AI cover letter generation, PDF/DOCX export and a companion Chrome extension. Reached active users across 7+ countries with 80%+ week-over-week new-user growth.

Next.js, TypeScript, Supabase, Claude API, Stripe, Chrome extension

ParkCast SFPredicted occupancy around Chase Center

Team project

ParkCast SF

End-to-end MLOps system predicting parking occupancy for 12.7K SF blocks at 8.98 MAE / 0.73 R² from a 33-feature LightGBM residual model. Fully automated GitHub Actions retraining pipeline with MLflow tracking and a champion-challenger gate that blocks silent regressions before Cloud Run rollout.

FastAPI, LightGBM, Docker, GCP Cloud Run, GitHub Actions, MLflow

PantrIQ Discover screen with featured recipes
PantrIQ meal planner week view
PantrIQ app icon

iOS app

PantrIQ

Full-stack AI mobile app for pantry management, with recipe suggestions, automated meal planning and shopping lists from scan, voice or text input.

React Native, TypeScript, Supabase, OpenAI, TailwindCSS

Word error rate, Whisper

Twenties6.53% Seventies4.67%

Utterances with a 700 ms pause

Twenties8.0% Sixties19.7%
2,760 Common Voice clips matched on accent, gender and speaker

Benchmark

asr-age-gap

A published benchmark that set out to find an age penalty in speech recognition and found the opposite. On 2,760 Common Voice clips matched on accent, gender and speaker, Whisper scores 4.67% WER on people in their seventies against 6.53% on twenties, replicated on a wav2vec2 CTC control. The penalty is in turn-taking instead: at a fixed 700ms endpoint, 19.7% of sixties utterances carry a pause that long against 8.0% of twenties, which word error rate cannot see.

PyTorch, Whisper, wav2vec2, Common Voice

Meridian Core, a legacy banking test app, on its Confirm Funds Transfer screen

Agents and automation

agent-hands

Computer-use automation for legacy back-office apps with no API. A model drives the app once and emits a reviewable, versioned capability artifact; a deterministic engine replays it with no model in the loop. Targets are recorded as a ranked list of accessibility-tree strategies, so a renamed field degrades to a fallback instead of breaking, and the degradation is logged. Built against a deliberately hostile fixture: frameset, table layout, no labels, server-generated IDs.

Python, Playwright, Claude API, Accessibility tree

71/75

answered or refused correctly

False answers
0
Against a no-data LLM
23–3

Evaluation

EARTHBENCH

A grounding harness that tests whether an agent-facing geospatial API can be trusted, scored against oracles outside the system under test. It answers or refuses correctly on 71 of 75 cases with zero false answers, and beats a no-data LLM 23–3. The finding that mattered was reproducibility: the same question at the same coordinate returned different field selections across calls, which for a product sold as audit-ready is the result worth reporting.

Python, Eval harness, Bootstrapping

More projects

  • Search Engine

    Full-stack search engine with inverted index, TF-IDF ranking and a multithreaded web crawler. 1,746 SLOC with 1,910 lines of Javadoc documentation.

    Java 21, Apache OpenNLP, Jetty, Bootstrap, Maven

    GitHub
  • Semiparametric Regression Visualizer

    Interactive web app for exploring semiparametric regression models. Compare OLS against Generalized Additive Models with real-time visualizations on datasets like Boston Housing and World Happiness.

    Python, Streamlit, pygam, statsmodels, Google Cloud Run

    GitHub
  • AI Interviewer Tool

    GenAI research app that simulates user interviews and focus groups with participant personas you define, for product and UX research.

    Python, Streamlit, OpenAI, Prompt engineering

    Live app
  • AI/ML Industry Data Dashboard

    Data dashboard aggregating AI/ML content from arXiv papers and Medium articles, with automated scraping via GCS, article trend visualizations and Anthropic-powered topic clustering.

    Python, Streamlit, GCS, NLP, Anthropic, Google Cloud Run

  • PeopleCodeOpenAI

    Python library that gives easy access to OpenAI’s tools and APIs, including text-to-speech, speech recognition and RAG. Used in introductory courses at USF.

    Python, OpenAI API, RAG, Speech recognition

Open source

33 pull requests merged into third-party production repos, with 67 more in review.

Merged openai/whisper #2836 The English number normalizer rewrote the digit 1 as “one” inside larger numbers, so timestamps and quantities came out wrong. Twelve lines. Merged by one of the authors of the Whisper paper.
Landed facebookresearch/faiss b4c66ba The index factory’s round-trip silently dropped the HNSW storage index. Imported and landed by a Meta engineer.
Merged astral-sh/uv #21144 and #21146 Two fixes into uv. Merged by Astral’s co-founder.
Merged librosa #2087 Input validation in rms. Merged by the library’s creator.
Merged elevenlabs-python #858 Closed as untouchable generated code until I showed the file was listed in their own .fernignore.
In review NVIDIA-NeMo/Speech #16200 A voice-activity detector clipping the last frame of every segment. Approved by an NVIDIA maintainer.
In review pytorch/audio #4227 A strided-window bug crashing feature extraction on non-contiguous audio.

Experience

  1. Jan 2026 – Jun 2026Remote

    Asurion

    AI Engineer

    • Built a voice turn-detection system, improving auto-labeling accuracy from 50% to 88%+ (90%+ on a 400-sample validation set)
    • Designed and scaled an automated labeling pipeline to auto-label 400K+ audio samples, enabling large-scale fine-tuning of a Whisper-based PyTorch model
    • Deployed the model in a real-time voice AI pipeline (LiveKit) for low-latency, end-to-end conversational agent testing
  2. Oct 2025 – Jan 2026Remote

    Spotly Jobs

    AI Engineer

    • Built optimized ATS pipelines, increasing job coverage 5%+ and processing hundreds of postings per run via dlt/dbt
    • Created 7 Dagster assets/sensors to automate ETL workflows, reducing manual intervention 50%+
    • Improved processing speed by 30–40% through optimized scraping logic and API handling
  3. May 2024 – Sep 2025Remote

    Outlier AI

    AI Engineer

    • Trained LLMs using RLHF, improving instruction-following accuracy by 20%
    • Evaluated 1M+ tokens across multiple datasets, surfacing 300+ instruction-following errors
    • Refactored 3+ AI models to generate more precise and succinct output
  4. Sep 2024 – May 2025San Francisco, CA

    University of San Francisco, IDEA Team

    Software Developer

    • Lead developer delivering 3 full-cycle applications for the IDEA team
    • Created 3 GenAI Python apps adopted by 50+ students and professors
    • Researched and evaluated 10+ existing AI tools and technologies
  5. Aug 2024 – Dec 2024San Francisco, CA

    Founders Network

    Software Developer

    • Built Slack integrations into a codebase of over 10,000 lines of code
    • Delivered 8 sprint cycles in Agile with a 95% on-time delivery rate
    • Collaborated directly with the CEO and a team of professional developers
  6. May 2024 – Aug 2024San Francisco, CA

    University of San Francisco

    Software Developer

    • Created and deployed 2 generative AI Thunkable apps, one featured on the site’s front page
    • Lead developer on a full-stack PeopleMuseum project, 5,000+ lines of code
    • Built front-end UI with React JS, achieving load times under 2 seconds
Portrait of Kayvan Zahiri

About

I’m an AI engineer, software developer and the founder of ResumeAI, an AI-powered resume platform I built solo that’s now used across 7+ countries. Most recently I was on the Voice AI team at Asurion, where I built a voice turn-detection system and scaled an automated labeling pipeline to 400K+ audio samples, used to fine-tune a Whisper-based model for real-time conversational AI.

With a B.S. in Computer Science and an M.S. in Data Science & AI from the University of San Francisco, I work across the full AI/ML lifecycle, from training and fine-tuning LLMs with RLHF to shipping production MLOps pipelines and full-stack apps end to end.

AI / ML
LLM fine-tuning, RLHF, RAG, voice AI, prompt engineering and model evaluation
Full-stack
React, Next.js, TypeScript, Python. Apps from database to deployment
Cloud & MLOps
AWS and GCP certified. MLOps pipelines, ETL, Docker, distributed computing

Skills

Languages
Python, TypeScript, Java, C, SQL, NoSQL, JavaScript, HTML/CSS
AI / ML
Generative AI, LLM fine-tuning, RLHF, voice AI, prompt engineering, model evaluation, RAG, deep learning, PyTorch, Whisper, Claude API, OpenAI
Data science
PySpark, SparkSQL, Pandas, NumPy, Scikit-Learn, Matplotlib, Plotly, ETL pipelines, regression, classification, unsupervised learning
Tools & platforms
AWS, GCP, Docker, Git / GitHub, Bash, LiveKit, MLflow, Dagster, Apache Airflow, dlt / dbt, MongoDB, distributed computing
Frameworks
React / React Native, Next.js, FastAPI, Supabase, Stripe, TailwindCSS, Streamlit, Bootstrap, Eclipse Jetty, Maven

Education

  • M.S. in Data Science & Artificial IntelligenceUniversity of San FranciscoJul 2025 – Jun 2026
  • B.S. in Computer ScienceUniversity of San FranciscoAug 2021 – May 2025
  • AWS Certified Cloud PractitionerAmazon Web Services
  • Google Cloud Digital LeaderGoogle Cloud

Contact

kayvanandre@gmail.com

The fastest way to reach me is email. I’m looking for AI engineering roles and interesting problems to work on.