April 2024 – December 2025
Built and optimized cloud data workflows, backend systems, an internal Azure analytics dashboard, and automated test coverage for cloud consumption, telemetry, and billing analytics.
- Modernized legacy Scala ETL logic by rewriting and simplifying 20+ Azure Synapse data pipelines in PySpark, processing billions of Azure usage and billing records for Fortune 500 customers and internal Microsoft teams.
- Reduced runtime of 3 production Apache Spark pipelines by 75%, from 24 hours to 5–6 hours, through stage-retry fixes, worker scaling, partition pruning, and join optimization, accelerating customer reporting.
- Built end-to-end ETL data pipelines with Cosmos DB, Cosmos SCOPE, and Kusto/KQL to ingest, validate, and surface 80+ Azure telemetry metrics in a stakeholder-facing dashboard for usage and cloud performance reporting.
- Maintained and expanded 150+ Playwright E2E tests in TypeScript, increasing dashboard test coverage and stabilizing flaky tests across 3 dashboard domains.
PythonPySparkAzure SynapseAzure Data FactoryCosmos DBKusto/KQLPlaywrightTypeScriptReactC#Azure DevOps