Bilal Shihab / Projects

MedAdvisor

An iPhone app that scores a recorded medical consultation against a 16-criterion clinical rubric, without the audio or transcript ever leaving the phone.

Built with a Stanford surgeonJun 2026 – presentTestFlight pilot
Specmeasured on iPhone 17
Model
Qwen 3.5-4B, Q4_K_M, 3.0 GB
Runtime
llama.cpp on Metal GPU
Active RAM
< 500 MB (memory-mapped weights)
Full analysis
~144 s mean
Transcription
Apple SpeechAnalyzer, 3.1% WER
Rubric
16 criteria, quoted evidence
Model choice
8+ models, 240 labeled decisions
Leaves the phone
Redacted scores, after review

Write-up coming soon.

Why I built it

Who it was for and what problem they had.

The hard part

One decision I had to make and what I gave up.

What I'd do next

Honest limits and the next measurement.

Qwen 3.5-4B 85% Qwen 2.5-7B 79% 0 50 100%
Fig. 2Agreement with 240 labeled rubric decisions. The smaller model scored higher and is a 30% smaller download, so it shipped.