

Lead Data Engineer with 13+ years of experience designing, building, and scaling batch and real-time data platforms. Expert in Apache Spark, Kafka, Databricks, Azure, and distributed systems, with a proven track record of delivering production-grade pipelines for banking and e-commerce platforms. Experienced in leading technical architecture, mentoring engineers, and optimizing large-scale data workflows processing more than 1 million rows per day across 20+ Kafka topics and 10+ pipelines. Strong background in streaming architectures, cloud-native deployments, and modern lakehouse technologies.
Data processing
- Apache Spark (PySpark and Scala)
- Databricks
- Delta Lake
- Unity Catalog
- Medallion Architecture
- Batch and stream processing
- Structured Streaming
- Auto Loader
- Lakeflow / Delta Live Tables
- Databricks Workflows
- Databricks Asset Bundles
- Data transformation and optimisation
Data pipelines
- ETL/ELT pipeline design
- Data ingestion (Batch and Real-time)
- Azure Data Factory
- Workflow orchestration
- Change Data Capture (CDC)
Performance tuning and optimisation
- Spark performance tuning
- Query optimisation
- Data partitioning strategies
Programming
- Python
- Scala
- SQL
- Go (Golang)
Data quality & governance
- Data Quality and Data Governance
DevOps
- Docker
- Kubernetes
- CI/CD and GitHub Actions
Cloud
- Azure
- GCP