Senior Data Scientist with 8+ years of experience applying machine learning, NLP/LLM, and statistical modeling to healthcare, education, testing, and psychometrics.
I'm a results-oriented data science and analytics professional who loves using the transformative power of data, technology, and analytics to solve complex, real-world problems.
Currently, I lead the development of predictive models and NLP/LLM solutions on 10M+ clinical records at Kentro, driving policy and care decisions across 15+ VA hospitals. Before that, I built LLM chatbots and psychometric analytics at Inteleos and researched patient health outcomes at UC San Diego Health.
I also teach data analytics and visualization to undergraduates at the University of Maryland Global Campus — because the best way to master something is to explain it. Beyond the classroom, I'm active in technical training, career coaching, and mentoring, from leading career panels to judging AI competitions.
My background bridges clinical psychology and computer science, which gives me an unusual lens: I care as much about the humans behind the data as the models built on top of it.
Beyond final-answer accuracy: auditing reasoning faithfulness in test-time-scaled medical LLMs. A pilot audit of HuatuoGPT-o1 on MedXpertQA Reasoning.
View on GitHub →Hybrid NLP pipeline (keyword tagging + DistilBERT classification + BERTopic clustering) analyzing 50,000+ referral records — reclassified 70%+ of "Unknown" entries and improved label accuracy by 35%.
Professional work — VA healthcareMulti-modal models for early detection of Mild Cognitive Impairment using acoustic, linguistic, and social-determinant features. Top 30 of 400 submissions in the NIH/NIA PREPARE Challenge.
NIH/NIA PREPARE ChallengeComparison of NLTK vs. SacreBLEU for machine-translation quality evaluation, highlighting differences in standardization, preprocessing, and reproducibility when assessing LLMs.
View on GitHub →RAG-style chatbot built with Python and LlamaIndex that reads imported documents and answers natural-language questions, improving document accessibility by 30%.
View on GitHub →Integrated machine learning and NLP (topic models, LDA) into test development — item bank analysis and automatic test assembly — improving test validity and balance by 30%.
View on GitHub →R Shiny application presenting a Behavioral Health Equity Index (BHEI) to support mental health service planning in San Diego County, enhancing resource allocation by 20%.
View on GitHub →Deep learning models (PyTorch) deployed with AWS SageMaker, Lambda, and API Gateway to classify chest X-rays for COVID-19 detection at 90% accuracy, with a real-time web app.
View on GitHub →Machine learning models leveraging social determinants of health data for early prediction of Alzheimer's disease and related dementias. arXiv:2503.16560 [q-bio.QM].
Read publication →Feasibility evaluation of a cognitive training program designed to strengthen executive function skills in autistic teens.
Read publication →Baker-Ericzén, M. J., Smith, L., Tran, A., & Scarvie, K. (2020). Pilot study of a cognitive behavioral intervention supporting driving skills in autistic teens and adults.
Read publication →Led a virtual career panel connecting Vietnamese students and early-career professionals with industry experts in data science and technology. Mentored case competition teams on analytical problem-solving, storytelling with data, and presentation strategy.
Mentored emerging AI talent and judged submissions in a national competition spotlighting Vietnam's next generation of AI builders — evaluating projects on technical rigor, innovation, and real-world impact, and coaching participants on machine learning and career growth in AI.