Hi there! This is Yixue (yi·sh·weh). I'm a Computer Scientist by training, a Generalist at heart (knowledge has no boundaries, right?) My focus evolves over time, largely driven by how the world changes and who I meet, such as from Program Analysis to AI for Mental Health. What doesn't change is that I always work on what I enjoy with people I love. What a blessing! 🥰
I founded Yixue Research Institute and became my own boss guided by what makes my heart sing. ✨ Before the leap, I was a Research Computer Scientist at USC Information Sciences Institute (ISI) and AI4Health Center. I was a Computing Innovation Fellow at University of Massachusetts Amherst (UMass) awarded by National Science Foundation (NSF) and Computing Research Association (CRA), and Microsoft Research (MSR) Fellow. I received my Ph.D. in Computer Science from University of Southern California (USC). Before my PhD, I grew up in China and received my Bachelor’s degree in Software Engineering from Harbin Institute of Technology (HIT).
💫 Leadership. Across all my work, people matter most. I'm very “picky” on who I work with (in a good way! 🤗) and I only collaborate in win-win situations. I'm deeply committed to leadership and mentoring: I founded HARMONY, a pioneering interdisciplinary workshop on AI and mental health; Steering Committee of ICSE's Student Mentoring Workshop; launched the Africa Initiative to bring African students to research conferences. My mentoring experience (20+ students, including high schoolers and underrepresented groups) can be found here.
🪷 Hobby. My work is my biggest hobby and I'm always learning something new (it feels like getting new toys!). Outside of research, I enjoy rock climbing, watching anime with my husband, taking fun classes (see class list), writing blogs (academic blog, happiness blog), hosting the Happiness Buffet podcast, reading books, making videos, Jazz, K-Pop, & Hip-hop dancing, and meditation (my favorite! 🧘♀️). More personal notes live on my misc page.
💌 Contact. For speaking engagements, podcasts, consulting, investments, or research collaborations, please use the contact form. For private 1:1 coaching and advisory work, please review the details on my Ko-fi page to ensure we’re a good fit. I have limited availability and new openings are quietly announced here. Hope to work with you someday! 💖
PhD in Computer Science
University of Southern California (USC), USA
BEng in Software Engineering
Harbin Institute of Technology (HIT), China
[Sep 2026] The HARMONY Substack is now LIVE! 🎉 Come join our growing community of interdisciplinary research on AI and mental health, and watch the HARMONY 2026 opening here! 🥰 Subscribe to stay in the loop! ✨
[Aug 2026] My personal favorite project launched! 🎉 Join this AMAZING Zambia safari meditation retreat of Lovingkindness led by Delson Armstrong and hosted by Dazzle Africa. TL;DR on my LinkedIn post! Hope to see you there!! 🥰
[Aug 2026] The FIRST pioneering interdisciplinary workshop on AI and mental health HARMONY 2026 made its debut in the beautiful Pittsburgh! What a historical moment we created together. THANK YOU ALL!! See my LinkedIn post for the beautiful memories! ✨
[May 2026] Our paper Toward Contemplative LLM: A Modular Framework for Evaluating and Enhancing LLM Alignment in Mental Health is accepted to HARMONY 2026! 🥳
[May 2026] My VERY FIRST YouTube milestone unlocked: 100 subscribers!!! ✅ It means so much to me. 🥰 Thank you for believing in me!! ✨ On to the next milestone! 💪
see CV for the full list :)
We present Trivana, an on-device mental health mobile application that enables holistic, proactive, and privacy-preserving support. The system integrates multimodal user inputs, a forecasting model for anticipating mental health risks, and a local LLM-based chatbot for personalized interventions. By unifying sensing, forecasting, and interaction within a single pipeline, Trivana overcomes fragmentation in existing apps while ensuring privacy guarantees through fully on-device execution.
Mental health challenges among college students continue to rise, motivating scalable and proactive support. Passive smartphone sensing provides an unobtrusive way to capture behavioral signals related to mental well-being, enabling machine learning models that predict future symptom changes. However, most existing work targets mental health detection, while leaving mental health forecasting underexplored. Forecasting remains technically challenging, and we still lack a clear understanding of which modeling choices most effectively improve long-term forecasting accuracy in realistic deployments. In this paper, we present the first large-scale empirical study of college student mental health forecasting using the College Experience Study (CES), a multi-year longitudinal dataset with passive sensing and weekly surveys. We systematically evaluate three practical design dimensions: (1) single-user forecasting under privacy-restricted, data-scarce settings; (2) model granularity, comparing population-level generic models, similarity-based models, and personalized fine-tuned models; and (3) architecture choice, contrasting one-stage end-to-end forecasting with a two-stage decoupled pipeline. Across controlled comparisons, population-level generalization consistently delivers the highest forecasting performance, achieving up to 0.777 accuracy, while similarity-based transfer and fine-tuning provide limited gains. Two-stage pipelines often reduce accuracy due to objective mismatch between stages. These findings provide actionable baselines and guidance for deployable mental health forecasting.
Contemplative traditions have long guided ethical behavior and prosocial interaction, and recent work suggests that contemplative principles (e.g., mindfulness, compassion, non-dual reasoning) may offer a promising paradigm for aligning large language models (LLMs), improving cooperation and reducing ethical violations in LLM outputs. However, as new models, evaluation metrics, and benchmarks emerge rapidly, it remains challenging to systematically assess whether and how contemplative principles enhance LLM alignment across diverse and evolving scenarios, and existing approaches are often ad hoc and fail to generalize. We present a modular, extensible evaluation framework, initially targeted at the mental health domain, that enables seamless integration of new models, metrics, and benchmarks through a reusable pipeline. The framework currently reproduces existing state-of-the-art results and supports systematic cross-evaluation by flexibly mixing and matching models, metrics, and benchmarks, enabling fair comparison and deeper insight. Its plug-and-play prompting module offers a principled pathway for incorporating ethical perspectives such as contemplative principles, allowing domain experts to define alignment criteria without requiring technical expertise. Although initially focused on mental health, the framework is domain-agnostic and extends naturally to areas such as decision-making, moral reasoning, and human-AI collaboration. By bridging computational evaluation with human-centered ethical reasoning, this work lays the groundwork for interdisciplinary research spanning cognitive science, behavioral economics, philosophy, and system design, toward robust, trustworthy, and socially beneficial human-AI ecosystems.
Mental health issues among college students have reached critical levels, significantly impacting academic performance and overall wellbeing. Predicting and understanding mental health status among college students is challenging due to three main factors: the necessity for large-scale longitudinal datasets, the prevalence of black-box machine learning models lacking transparency, and the tendency of existing approaches to provide aggregated insights at the population level rather than individualized understanding. To tackle these challenges, this paper presents I-HOPE, the first Interpretable Hierarchical mOdel for Personalized mEntal health prediction. I-HOPE is a two-stage hierarchical model that connects raw behavioral features to mental health status through five defined behavioral categories as interaction labels. We evaluate I-HOPE on the College Experience Study, the longest longitudinal mobile sensing dataset. This dataset spans five years and captures data from both pre-pandemic periods and the COVID-19 pandemic. I-HOPE achieves a prediction accuracy of 91%, significantly surpassing the 60-70% accuracy of baseline methods. In addition, I-HOPE distills complex patterns into interpretable and individualized insights, enabling the future development of tailored interventions and improving mental health support.
The prevalence of social media and its escalating impact on mental health has highlighted the need for effective digital wellbeing strategies. Current digital wellbeing interventions have primarily focused on reducing screen time and social media use, often neglecting the potential benefits of these platforms. This paper introduces a new perspective centered around empowering positive social media experiences, instead of limiting users with restrictive rules. In line with this perspective, we lay out the key requirements that should be considered in future work, aiming to spark a dialogue in this emerging area. We further present our initial effort to address these requirements with PauseNow, an innovative digital wellbeing intervention designed to align users’ digital behaviors with their intentions. PauseNow leverages digital nudging and intention-aware recommendations to gently guide users back to their original intentions when they “get lost” during their digital usage, promoting a more mindful use of social media.
The COVID-19 pandemic has intensified the urgency for effective and accessible mental health interventions in people's daily lives. Mobile Health (mHealth) solutions, such as AI Chatbots and Mindfulness Apps, have gained traction as they expand beyond traditional clinical settings to support daily life. However, the effectiveness of current mHealth solutions is impeded by the lack of context-awareness, personalization, and modularity to foster their reusability. This paper introduces CAREForMe, a contextual multi-armed bandit (CMAB) recommendation framework for mental health. Designed with context-awareness, personalization, and modularity at its core, CAREForMe harnesses mobile sensing and integrates online learning algorithms with user clustering capability to deliver timely, personalized recommendations. With its modular design, CAREForMe serves as both a customizable recommendation framework to guide future research, and a collaborative platform to facilitate interdisciplinary contributions in mHealth research. We showcase CAREForMe's versatility through its implementation across various platforms (e.g., Discord, Telegram) and its customization to diverse recommendation features.
Writing and maintaining UI tests for mobile apps is a time-consuming and tedious task. While decades of research have produced automated approaches for UI test generation, these approaches typically focus on testing for crashes or maximizing code coverage. By contrast, recent research has shown that developers prefer usage-based tests, which center around specific uses of app features, to help support activities such as regression testing. Very few existing techniques support the generation of such tests, as doing so requires automating the difficult task of understanding the semantics of UI screens and user inputs. In this paper, we introduce Avgust, which automates key steps of generating usage-based tests. Avgust uses neural models for image understanding to process video recordings of app uses to synthesize an app-agnostic state-machine encoding of those uses. Then, Avgust uses this encoding to synthesize test cases for a new target app. We evaluate Avgust on 374 videos of common uses of 18 popular apps and show that 69% of the tests Avgust generates successfully execute the desired usage, and that Avgust's classifiers outperform the state of the art.
Prefetching and caching is a fundamental approach to reduce user-perceived latency, and has been shown effective in various domains for decades. However, its application on today’s mobile apps remains largely under-explored. This is an important but overlooked research area since mobile devices have become the dominant platform, and this trend is reflected in the billions of mobile devices and millions of mobile apps in use today. At the same time, user-perceived latency has been shown to have a large impact on mobile-user experience and can cause significant economic consequences. ❧ In this dissertation, I aim to fill this gap by providing a multifaceted solution to establish the foundation for exploring prefetching and caching in the mobile-app domain. To that end, my dissertation consists of four major elements. As a first step, I conducted an extensive study to investigate the opportunities for applying prefetching and caching techniques in mobile apps, providing empirical evidence on their applicability and demonstrating insights to guide future techniques. Second, I developed PALOMA, the first content-based prefetching technique for mobile apps using program analysis, which has achieved significant latency reduction with high accuracy and negligible overhead. Third, I constructed HiPHarness, a tailorable framework for investigating history-based prefetching in a wide range of scenarios. Guided by today’s stringent privacy regulations that have limited the access to mobile-user data, I further leveraged HiPHarness to conduct the first study on history-based prefetching with “small” prediction models, demonstrating its feasibility on mobile platforms and in turn, opening up a new research direction. Finally, to reduce the manual effort required in evaluating prefetching and caching techniques, I have devised FrUITeR, a customizable framework for assessing test-reuse techniques, in order to automatically select suitable test cases for evaluating prefetching and caching techniques without real users’ engagement as required previously.
UI testing is tedious and time-consuming due to the manual effort required. Recent research has explored opportunities for reusing existing UI tests from an app to automatically generate new tests for other apps. However, the evaluation of such techniques currently remains manual, unscalable, and unreproducible, which can waste effort and impede progress in this emerging area. We introduce FrUITeR, a framework that automatically evaluates UI test reuse in a reproducible way. We apply FrUITeR to existing test-reuse techniques on a uniform benchmark we established, resulting in 11,917 test reuse cases from 20 apps. We report several key findings aimed at improving UI test reuse that are missed by existing work.
A large number of mobile-app analysis and instrumentation techniques have emerged in the past decade. However, those techniques’ components are difficult to extract and reuse outside their original tools, their evaluation results are hard to reproduce, and the tools themselves are hard to compare. This paper introduces DECREE, an infrastructure intended to guide such techniques to be reproducible, practical, reusable, and easy to adopt in practice. DECREE allows researchers and developers to easily discover existing solutions to their needs, enables unbiased and reproducible evaluation, and supports easy construction and execution of replication studies. The paper describes DECREE's three modules and its potential to fundamentally alter how research is conducted in this area.
Reducing network latency in mobile applications is an effective way of improving the mobile user experience and has tangible economic benefits. This paper presents PALOMA, a novel client-centric technique for reducing the network latency by prefetching HTTP requests in Android apps. Our work leverages string analysis and callback control-flow analysis to automatically instrument apps using PALOMA's rigorous formulation of scenarios that address “what” and “when” to prefetch. PALOMA has been shown to incur significant runtime savings (several hundred milliseconds per prefetchable HTTP request), both when applied on a reusable evaluation benchmark we have developed and on real applications.
see CV for the full list :)