top of page
Search

Research Abstract: ML for Travel-Associated Disease Risk


I’ve finally submitted my paper to the National High School Journal of Science. This has been such an interesting process. I came in with just a question and no idea how to even begin to answer it and now I have a fleshed out research paper full of new information I’ve learned. To be clear, it definitely wasn’t all smooth sailing. I had to completely restart with a new dataset, reframe my paper to address the limitations of using synthetic data, and constantly work to improve my model’s performance. It obviously taught me a lot about Random Forest models, global health, and the lack of accessible travel-linked health data; however, it also taught me how to be patient with the formulation of a research paper and how to make sure it’s readable. One of the last things I did was exchange a few messy tables for clear color-coded graphs and charts, and I’m so happy I decided to do so. 

I’m pasting my final research abstract below. Hopefully it’ll interest you in reading the whole paper, which I’ve pasted the link for at the bottom. If you’re struggling with understanding any particular part while reading, reach out and I’ll make a blog post about it :) Thanks for coming along with me!


Title: A Proof-of-Concept Machine Learning Framework for Disease Risk Assessment Using Synthetic Travel Data: A Case Study of India-US Traveler Records.


Abstract:

Human mobility, especially in terms of international travel, plays a significant role in disease transmission, as was evidenced during the COVID-19 pandemic. However, healthcare systems are not fully prepared to tackle global pathogens and treatment is less expedient than desired. There is potential for machine learning methods to improve outcomes but international health data is lacking. This exploratory study examines whether a Random Forest model, trained on a synthetic dataset of traveler records, can distinguish between and draw patterns from disease categories among individuals traveling between India and the United States. A synthetic cohort of 300 traveler records was generated using a literature-informed GLMM-hurdle movement model and infection simulation incorporating demographic, geographic, temporal, and clinical features across ten prevalent diseases: COVID-19, Influenza, Dengue, Malaria, Tuberculosis, Hepatitis, Chikungunya, Leptospirosis, Typhoid, and no disease. The tuned Random Forest achieved an accuracy of 78.3% and a macro-AUC of 0.968, outperforming Logistic Regression (75.0%, AUC 0.940) and XGBoost (73.3%, AUC 0.953). Network analysis of city-level travel connectivity within the synthetic dataset identified high-connectivity nodes. These findings are constrained by the synthetic nature of the data and the sample size, and cannot be generalized to real-world epidemiological patterns without further validation. This proof-of-concept study demonstrates the feasibility of applying machine learning classification to multi-feature traveler datasets and identifies key methodological challenges that must be addressed in future work. 

Keywords: Infectious Diseases, Machine Learning, Random Forest Modeling, Health Policy, Epidemic, Global Health, Synthetic Data 


 
 
 

Recent Posts

See All
Past the Lab

IN PROGRESS At VHS, I had the opportunity to speak to the Chief Medical Officer of Infectious Diseases, Dr. Kumaraswamy, on his experience developing and promoting HIV Antiretroviral therapy in India

 
 
 

Comments


bottom of page