SOFTWARE ENGINEER & SDET

Jennifer Montgomery

Backend · full-stack · quality engineering

COURSEWORK / INN HOTELS

Classification and decision trees

Could booking details help identify reservations at higher risk of cancellation?

← Coursework on resume

Completed UT Austin postgraduate coursework using a supplied scenario, dataset, and starter notebook. The charts below come from my completed notebook.

SUPERVISED LEARNING

INN Hotels

COURSE FOCUS

Classification, model evaluation, and decision trees

LIBRARIES USED

pandas, NumPy, statsmodels, scikit-learn, Seaborn, Matplotlib

The supplied data

The course supplied reservation records labeled canceled or not canceled. Fields included lead time, room price, special requests, repeat-guest history, market segment, arrival timing, adults and children, and prior cancellations. The scenario asked which booking patterns related to cancellations and how a classifier might flag higher-risk reservations.

What I did

  • Explored distributions and relationships between booking attributes and cancellation status.
  • Prepared categorical fields and compared a logistic-regression baseline with decision-tree models.
  • Examined confusion matrices, precision, recall, and F1, including how tree pruning changed the balance between fit and interpretability.
Bar chart showing the counts of canceled and not-canceled reservations in the course dataset.
From the notebook: the cancellation label distribution sets the baseline for evaluating a classifier. Open chart ↗
Grouped bars show the percentage of canceled and not-canceled bookings by number of special requests in the course dataset.
From the notebook: cancellation share by special requests. This association does not show that requests prevent cancellations; high-request groups were small. Open chart ↗

Finding in the course report

I favored a pre-pruned tree for the exercise's balance of performance and interpretability. Lead time, special requests, and repeat-guest patterns informed possible service and cancellation-policy questions, not a deployed model or proven causal effects.

Learning reinforced

The project brought together exploratory analysis, a baseline model, nonlinear classification, and the tradeoff between overfitting and useful generalization. It also made precision and recall concrete business choices.