dc.contributor.authorSekabanja, Tinashe
dc.date.accessioned2026-04-15T20:52:28Z
dc.date.available2026-04-15T20:52:28Z
dc.date.graduationmonthMay
dc.date.issued2026
dc.description.abstractRetail analytics enables businesses to analyze purchasing patterns, forecast demand, and improve inventory management. This study explores and compares statistical and machine learning methods for modeling retail demand using a synthesized dataset from Kaggle, the Product Sales Dataset (2023–2024) by Yennewar (2023). The dataset included transaction records across multiple states, product categories, and time periods in the United States. Exploratory analysis was conducted to examine geographic and temporal patterns in product demand, revealing significant regional variations that were confirmed through analysis of variance (ANOVA). Transaction-level data were aggregated to monthly and quarterly product-level demand by state and city, and missing combinations were reconstructed to account for no purchases. Thus, producing a complete dataset suitable for analysis. To model demand, a combination of time series, statistical, and machine learning approaches were implemented. ARIMA models were used to capture temporal trends, while count based regression models (Poisson and Negative Binomial) were applied to model demand. Additionally, classification methods including ordinal logistic regression, decision trees, and XGBoost were applied to predict categorized demand levels. Model performance was evaluated using metrics such as mean absolute error and accuracy. Results indicated that demand varies significantly across both regions and time. Product-level time series models captured underlying trends less effectively at aggregated levels. In the regression approach, the Negative Binomial model outperformed the Poisson model due to overdispersion in the data. Classification models showed moderate performance with stronger predictive results for extreme demand categories (very low demand and high demand). Notably, XGBoost underperformed compared to the other models but had a uniform moderately weak performance across all classes. Overall, the findings highlight the importance of aligning model selection with data structure and distribution as well as problem structure, demonstrating that increased model complexity does not necessarily lead to improved predictive performance in forecasting retail demand.
dc.description.advisorPerla E. Reyes Cuellar
dc.description.degreeMaster of Science
dc.description.departmentDepartment of Statistics
dc.description.levelMasters
dc.identifier.urihttps://hdl.handle.net/2097/47235
dc.language.isoen_US
dc.subjectRetail demand forecasting
dc.subjectOrdinal regression
dc.subjectStatistics retail demand
dc.titleA statistical approach to analyzing and forecasting retail demand
dc.typeReport

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
TinasheSekabanja2026.pdf
Size:
813.14 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.65 KB
Format:
Item-specific license agreed upon to submission
Description: