Hi, I'm Sourabh Savre
I build robust
Hands-on Data Engineer with experience building scalable ETL pipelines, orchestrating workflows with Airflow/Delta Live Tables, and implementing high-efficiency Medallion architectures in cloud environments.
About Me
I am a results-oriented Data Engineer who loves turning massive volumes of raw, unstructured data into clean, highly-optimized pipelines. My engineering journey is centered around designing automated architectures that handle high-velocity datasets without breaking.
Having built end-to-end data systems at scale, I specialize in leveraging PySpark, Databricks, and modern cloud infrastructures to deliver analytics-ready datasets. I approach data engineering not just as a support system, but as a critical driver for business analytics and machine learning applications.
Data Engineering Philosophy
"A great data pipeline is invisible—it must be modular, idempotent, and self-healing. I build with the conviction that data quality is absolute, latency should be minimized proactively, and infrastructure must scale dynamically with business needs."
Robust Pipelines
Fault-tolerant ETL workflows built using Airflow and Spark Streaming.
Optimized Compute
Performance tuning via partition pruning, broadcasting, and Z-Ordering.
Modern Storage
Delta Lake architectures enforcing ACID transactions and schema compliance.
Data Quality
SCD Type 2 schemas and comprehensive validation tests.
Technical Skills
Core technologies and architectures I use to design, build, and optimize scalable data systems.
Languages
Big Data
Cloud Platforms
Data Engineering
DevOps & Tools
Architecture
Work Experience
Data Engineer
Active RoleSmallest.ai
- Designed and developed scalable ETL pipelines using PySpark and Apache Spark to process large-scale datasets efficiently.
- Built and maintained data pipelines on Databricks leveraging Delta Lake for reliable, ACID-compliant data storage and processing.
- Implemented Medallion Architecture (Bronze-Silver-Gold layers) to ensure clean, structured, analytics-ready data.
- Orchestrated data workflows using Delta Live Tables (DLT) and Apache Airflow for automated, fault-tolerant pipeline execution.
- Collaborated with data scientists and analysts to deliver high-quality datasets for ML model training and business reporting.
Featured Projects
Case studies of pipelines and systems built with an emphasis on scale, performance, and clean design.
Real-Time Data Pipeline
High ingestion latency and lack of transactional guarantees for event-driven streaming datasets, causing delays in analytical reports.
Built a multi-tier streaming pipeline ingesting raw events into a Bronze layer, applying cleanups/validations in Silver, and aggregating business metrics into a Gold layer.
Significantly reduced latency by adopting incremental loads and applying Z-ORDER cluster-key optimization on Delta tables for faster analytical queries.
Cloud Data Warehouse Migration
Slow query processing times and costly maintenance of legacy, on-premise database warehouse systems with complex schema transformations.
Migrated the database to AWS Redshift, redesigned schemas as Star Schemas, and implemented SCD Type 2 patterns to track historical changes and automate deduplication.
Reduced query execution times by 40% through partitioning strategies, distribution keys, and broadcast join optimizations in Spark workloads.
Crop Price Predictor
Inability for market players to predict crop price fluctuations due to scattered historic files and high seasonal price variance.
Conducted extensive EDA and feature engineering using Pandas and NumPy, then trained, cross-validated, and optimized multiple machine learning regression models.
Delivered a high-precision forecasting tool, choosing the best-performing regression algorithm to accurately predict future price trends.
Architecture Showcase
Interactive representation of the Medallion Architecture (Bronze-Silver-Gold) pipelines I build for data governance and query performance.
Licenses & Certifications

AWS Certified Data Engineer - Associate
Amazon Web Services (AWS)

Databricks Certified Associate Developer for Apache Spark
Databricks

Python for Data Engineering
Coursera
Education
B.Tech, Computer Science and Engineering
Indore Institute of Science and Technology
Key Focus & Coursework
Get in Touch
Have a question or want to discuss scaling a data pipeline? Drop me a message below or contact me directly.
Contact Information
Notice
If the automated form API fails due to system blocks, the submission defaults back to opening your local mail client.