Primary Focus Area
Responsible AI
As AI becomes more pervasive in our lives, extra care must be taken to ensure that their training and deployment are handled responsibly, with an emphasis on mitigating harmful biases.
Responsible AI · Bias & interpretability
I build tools that make machine learning systems trustworthy: auditing datasets before they become models, and models before they become decisions.
Primary Focus Area
As AI becomes more pervasive in our lives, extra care must be taken to ensure that their training and deployment are handled responsibly, with an emphasis on mitigating harmful biases.
Powerful models deployed in the wrong way can cause unexpected harm. Knowing when to trust a model involves understanding how it makes decisions.
Finding just the right visualization to generate and communicate insights with others is often the most valuable tool in any discussion.
Please check out the Projects tab to see some of my work!
An open source python library for evaluating training datasets
I'm part of the team developing DataEval with the goal of providing data scientists with an easy to use toolkit for quickly surfacing issues with training data including biases, outliers, labeling errors, and incomplete concept coverage. My role centers on adapting our existing image-centric tools to video data.
View DataEval's Documentation
How do models define hate speech?
This project was motived by my belief that a deeper understanding of how models learn to define and detect hate speech may provide guidance on what to focus on in building training datasets, and where the models can be deployed most effectively. I found that a multilingual model trained on both English and Korean hate speech datasets learned a 'clearer definition' of hate speech than either monolingual model, suggesting that each language brings a different view of hate speech to the table and combine to form a more 'complete' dataset.
View Paper on Google Drive
Presented at SPIE 2026
AI systems need continuous performance assurance as models face evolving environments that can degrade performance undetected, but continually labeling new data to check for drift is costly and time consuming. I presented an uncertainty-based drift detection workflow for monitoring motion imagery datasets that requires no labeled operation data, no significant architectural modifications to the deployed model, and is computationally suitable for edge deployment.
View Paper and Presentation
Helping prospective homeowners understand housing affordability across the U.S.
I worked on a team of 4, working to show how housing costs have changed over time, and across different regions in the United States. My focus was on correlations between various factors related to the housing market and housing affordability metrics. We researched, designed visualizations, and conducted a usability study to ensure that our visualization tools were informative and intuitive.
View Dashboard
Published in Computational Materials Science
A client had developed a few shot ML model whose reference examples could be crops extracted from a single image. In order to use this model, one had to first divide an image into a grid of patches, label individual patches, then pass those to the model for fine tuning. My team designed a user interface to make that process much easier for the user.
View Paper
This site is still a work in progress
I can't wait to share more of my work soon!
This page is a work in progress, please come back soon for more detail :)