Home About Skills Projects Experience Education Contact
Data Scientist

Hi, I'm HASNAIN SIZAR

Forecasting demand in Python and SQL

Three years of analytics work across hospitality revenue, pharmacy benefits, and financial services. I build demand forecasts and BI dashboards in Python and SQL that replace manual reporting and feed weekly pricing decisions. Underneath them sit the ETL jobs, reconciliation, and data quality checks that make the numbers safe to publish.

Connect
Illustrated portrait of Hasnain Sizar
Illustrated portrait of Hasnain Sizar

Data Scientist

Demand forecasting, BI dashboards, ETL and data quality, and experimentation in Python and SQL.

Chino, CA · Open to data science roles

I'm Hasnain Sizar

I graduated from UC Irvine with a B.S. in Data Science and have three years of analytics work behind me: hospitality revenue at Choice Hotels, pharmacy benefits at Prime Therapeutics, and financial services at Bank of America.

The work itself is forecasting, ETL, and reporting. I build demand and volume forecasts in Python and SQL, write the ETL jobs and data quality checks that feed them, model the results as star schemas, and publish them as Power BI and Tableau dashboards that replace manual Excel rollups.

I also design experiments, from A/B tests on promotional campaigns to causal studies with matching and difference-in-differences. Whatever I ship comes with tests, documentation, and a stated tradeoff behind every threshold, and that standard carries into the projects below.

6 End-to-End Projects
266+ Tests on Green CI
68,404 Matches in Causal Panel

Skills & Tools

The languages, methods, and tooling behind the projects below.

Languages

Daily drivers for analysis and pipelines

Python SQL R

Python Ecosystem

Modeling and analysis libraries

Pandas NumPy Scikit-learn Statsmodels Matplotlib

Data Engineering

ETL, warehouses, and data quality

ETL Data Validation Snowflake BigQuery Databricks Azure Data Factory AWS S3 dbt PostgreSQL SQLite Schema Design CTEs Window Functions Stored Procedures Query Optimization

Machine Learning

Supervised models, evaluated honestly

Supervised Classification Regression XGBoost Feature Engineering Model Selection Model Evaluation Precision Recall AUC Cross-Validation

Statistics & Experimentation

Testing, forecasting, and causal methods

A/B Testing Hypothesis Testing Forecasting Time Series Causal Inference Propensity Score Matching Difference-in-Differences Event Study Placebo Tests

BI & Visualization

Dashboards and automated reporting

Power BI DAX Power Query Tableau Star-Schema Modeling Excel PivotTables Automated HTML Reporting Jinja2

Business Analysis & Delivery

Requirements, stakeholders, and delivery

Requirements Gathering (BRD/PRD) Stakeholder Presentations Agile/Scrum JIRA Confluence

Tools & Practices

How the code gets shipped

Git GitHub GitHub Actions CI pytest ruff mypy Jupyter Typer Claude Code Unix Command Line

Selected Projects

Six end-to-end systems. Each card shows real output, and each repo ships with tests and documented tradeoffs.

Groundswell expansion brief for a sample account showing a composite expansion score of 37 and a consumption ramp signal Analytics

Groundswell

Account usage analytics and alerting over daily telemetry (compute, seats, workspaces, feature breadth) for a 60-account portfolio. Five rule-based detectors using window medians and weekday-matched baselines separate durable change from noise, feeding composite scoring and automated HTML briefs. Validated at 0.89 precision and 0.94 recall over 10 generated datasets and 600 accounts. 116 tests, green CI.

Python SQLite Typer Jinja2 GitHub Actions
Snowwatch displacement digest showing 43 signals collected over 14 days and a scored displacement signal ETL + Scoring

Snowwatch

Competitive-intelligence signal pipeline that ingests posts and job listings from three public APIs, deduplicates into SQLite, and applies direction-aware rules to surface platform-switching signals mapped to follow-up actions. False positives cut with a skills-list dampener, staffing-firm flags, score floors, and suppression notes. 150 tests, green CI.

Python SQLite Typer Hacker News API Stack Exchange API Adzuna API
Rxdelta report header showing coverage changes between two monthly CMS Part D snapshots across 5,517 plans CLI + Reporting

Rxdelta

Medicare Part D formulary change monitor. Loads two monthly CMS releases, roughly 1.1M formulary rows per month across 5,518 plans, into partitioned SQLite with a full audit trail, then classifies tier moves, prior authorization, step therapy, quantity limits, additions, and drops and ranks them by estimated member cost impact. Ships documented cost-impact ranges, mypy strict, ruff, a pytest coverage gate, and a self-contained HTML report.

Python SQLite Typer pytest GitHub Actions CMS public-use files
Scatter plot of expected goal difference against final league position for Serie A 2019/20 with a Spearman correlation of 0.987 Causal Inference

Causal Impact of Mid-Season Managerial Changes

Propensity Score Matching combined with Difference-in-Differences to estimate the causal effect of mid-season manager firings across 68,404 matches, the top 20 European leagues, and six seasons (2019/20 to 2024/25). Estimated +0.292 xGD per match over 12 matchweeks, 95% CI [0.192, 0.392], p < 0.001. Pipeline built on API-Football and Transfermarkt via Selenium into a 6-table SQLite database with 2,053 firings, validated with covariate balance, event-study pre-trends, and placebo tests.

Python Scikit-learn Statsmodels Pandas SQLite Selenium
The Oracle prediction card for Norway against England picking England at 64 percent with narration Prediction

The Oracle

World Cup 2026 prediction bot. A transparent Elo-style rating model computes win probabilities and passes only the computed numbers to an LLM for commentary, so the narration cannot invent scores or statistics. Renders shareable 1200x720 PNG cards with team flags, pick, probability split, and narration, plus offline fallbacks, cached assets, and a terminal card for local runs.

Python Anthropic API Pillow
Staycast bar chart of pooled backtest WAPE by model, with Prophet at 0.268 and the gradient boosted tree at 0.270 ahead of SARIMA and both seasonal naive baselines Forecasting

Staycast

Daily demand forecasting and price-band flagging for San Diego short-term rentals on public Inside Airbnb data. Turns 949,213 reviews across 13,213 listings into a demand index for 21 neighbourhood and room-type series, then backtests Prophet, SARIMA, and a global gradient boosted tree against weekly and yearly seasonal naive baselines over five rolling 90-day windows. Prophet and the tree beat both baselines in every window, with WAPE 19.7% and 19.1% below the weekly naive. A price model and quartile rule flag 1,280 listings as underpriced or overpriced, and everything exports to SQLite, charts, and a Power BI star schema.

Python Pandas Scikit-learn Statsmodels Prophet SQLite Typer Power BI GitHub Actions

Work Experience

Three years of analytics work across hospitality revenue, pharmacy benefits, and financial services.

Business / Data Analyst

Jun 2025 to Present

Choice Hotels, Los Angeles, CA

  • Built four Power BI dashboards on a star-schema model covering occupancy, ADR, RevPAR, and channel mix, replacing a manual Excel rollup with weekly reporting for hotel, revenue, and marketing leaders.
  • Developed Python and SQL demand forecasts from two years of booking and rate data across 80 properties, cutting MAPE from 18% to 12% to support weekly pricing decisions.
  • Clean, transform, and reconcile booking, rate, and channel data across systems, resolving quality issues and documenting metric definitions before reports publish.
  • Delivered more than 20 analyses for revenue and marketing leaders and brought average turnaround down from about three days to four hours by templating recurring data pulls.
  • Designed A/B tests for three promotional campaigns. The winning treatment lifted loyalty engagement 7% against control across a 25-property pilot.

Junior Data Scientist

Sep 2024 to Jun 2025

Prime Therapeutics, Remote

  • Authored six Python and SQL ETL jobs integrating claims, formulary, and member data from five source systems, with scheduled daily refreshes, run-time alerts, and documented lineage.
  • Built an XGBoost claims-volume forecast on 24 months of data that cut forecast error 10% against a linear baseline for Operations and Finance capacity planning.
  • Developed three Power BI and Tableau dashboards on claims and prescription trends. Targeting changes informed by those dashboards raised campaign ROI 15% year over year.

Data Science Intern

Jun 2023 to Aug 2023

Bank of America, New York, NY

  • Co-developed a Python and SQL fraud-detection pipeline processing more than 2M daily card transactions, tuning thresholds to cut false positives about 10% and lift recall 9% on the weekly evaluation set.
  • Delivered two Tableau dashboards for fraud trends and model performance, replacing two manual Excel reports used by fraud operations.

My Education

The degree behind the projects.

Bachelor of Science in Data Science

University of California, Irvine

Relevant coursework
Machine Learning Statistical Analysis Data Structures and Algorithms Database Systems Data Mining Linear Algebra Probability and Statistics

Let's Connect

Open to data scientist, analyst, and data engineering roles. Email is the fastest way to reach me.

Email

hasnainsizar@outlook.com

Phone

562-386-4852

Location

Chino, CA

Find me online
Copied to clipboard