Completed UT Austin postgraduate coursework using a supplied scenario, dataset, and starter notebook. The charts below come from my completed notebook.
EasyVisa
Bagging, boosting, model comparison, and tuning
pandas, NumPy, scikit-learn, XGBoost, Seaborn, Matplotlib
The supplied data
The course provided historical labor-certification cases with a certified or denied label. Features described the employer and position, including company size, establishment year, prevailing wage and wage unit, plus applicant education, experience, and geographic fields. These are historical labels in a classroom dataset, not a basis for deciding an individual case.
What I did
- Checked the outcome distribution and explored relationships between case attributes and the recorded status.
- Compared decision trees, random forests, bagging, AdaBoost, gradient boosting, XGBoost, and stacking using classification measures.
- Tuned promising ensembles and compared training with held-out performance to look for overfitting.


Finding in the course comparison
Tuned gradient boosting and XGBoost were among the stronger held-out models in the notebook, with F1 scores around 0.82. Some tree models fit the training data much better than the test data, underscoring the need to compare both.
Learning reinforced
The project made ensemble methods and overfitting tangible. Because the features include education and geography and the outcome affects people, a classroom score and feature-importance chart cannot justify real eligibility decisions without legal, fairness, and data-quality review.
Original notebook export
View the complete HTML report, including code, outputs, and the original written analysis.