Profil professionnel
Vue d'ensemble
Expérience
Formation
Compétences
Langues
Certificats
Chronologie
Generic
Djim DIOP

Djim DIOP

Profil professionnel

Lead Data Engineer with 13+ years of experience designing, building, and scaling batch and real-time data platforms. Expert in Apache Spark, Kafka, Databricks, Azure, and distributed systems, with a proven track record of delivering production-grade pipelines for banking and e-commerce platforms. Experienced in leading technical architecture, mentoring engineers, and optimizing large-scale data workflows processing more than 1 million rows per day across 20+ Kafka topics and 10+ pipelines. Strong background in streaming architectures, cloud-native deployments, and modern lakehouse technologies.

Vue d'ensemble

2
2
Languages
1
1
Certification
14
14
years of professional experience

Expérience

Lead Data Engineer

Dscale
Dubai
2021.03 - 2026.07
  • Architected and led development of real-time data platforms from scratch, powering event-driven pipelines for e-commerce systems (Kafka Streams, Golang APIs).
  • Managed and optimized 20+ Kafka topics and 10+ data pipelines, ensuring reliable real-time data flow across e-commerce systems.
  • Improved data-processing performance and reduced latency across large-scale real-time ingestion pipelines, enhancing overall system responsiveness.
  • Designed and built a PySpark data pipeline on Databricks processing 1M+ events, implementing a bronze/silver/gold architecture for scalable batch ingestion.
  • Built and maintained source/sink connectors for real-time ingestion, enabling faster product search and filtering via Solr/Algolia.
  • Improved data-processing performance and reduced latency across large-scale real-time ingestion pipelines, enhancing overall system responsiveness.systems.

Data Engineer

ZAND Bank
Dubai
2019.12 - 2021.03
  • Architected and developed real-time Kafka Streams applications from scratch, handling 100K events/sec for [business use case].
  • Developed ETL data-processing pipelines with Spark.
  • Built source connectors (Debezium) to ingest data from SQL Server into Kafka topics, and sink connectors to capture and route error data for analysis.
  • Deployed applications to QA and production environments using Docker and Kubernetes, with unit test coverage to ensure reliability.
  • Translated business requirements into technical specifications in collaboration with cross-functional teams, delivering 3 projects

Data Engineer

Societe Generale
Paris
2016.08 - 2019.11
  • Defined Big Data technical solutions across multiple use cases, translating
    functional requirements into technical specifications.
  • Designed Big Data models and developed ETL pipelines for large-scale data
    extraction, processing, and transformation.
  • Built Spark Streaming and batch data-processing modules, handling
    differences sources and formats of data for bank use cases.
  • Migrated Spark programs from version 1.6 to 2.2, ensuring compatibility and
    leveraging DataFrame API, Catalyst optimizer.
  • Created and managed Elasticsearch indexes, building Kibana dashboards.
  • Automated pipeline scheduling using Oozie workflows.
  • Defined Big Data technical solutions across multiple use cases, translating functional requirements into technical specifications.

Data Engineer/Business Intelligence

Capgemini
Rennes
2012.04 - 2016.04
  • Designed and developed Spark data-processing modules and user-defined functions (UDFs) for large-scale Big Data transformations.
  • Built data pipelines loading flat-file data into Cassandra, integrating diverse sources for streamlined analysis.
  • Designed data warehouses and data marts, and developed SSIS packages to centralize data from diverse sources into SQL Server.
  • Led feasibility analysis for migrating legacy BI applications to a Big Data architecture, defining the target technical solution.
  • Loaded and managed data from flat files and Excel into Oracle using SQL, including creation/upgrade of stored procedures, tables, and queries.
  • Conducted data quality checks and testing (test plans/scripts based on requirements), rectifying inconsistencies to ensure data reliability.
  • Produced and delivered installation documentation for pre-production and production environments.

Formation

Double Degree, MSc - Business Intelligence

Grenoble Management School-ESC Grenoble
Grenoble

Masterʼs Degree - information systems for bank, finance and industry

ESIEA (École Supérieure d'Informatique, d'Électronique et Automatique)
Laval/Paris

Compétences

Data processing

- Apache Spark (PySpark and Scala)

- Databricks

- Delta Lake

- Unity Catalog

- Medallion Architecture

- Batch and stream processing

- Structured Streaming

- Auto Loader

- Lakeflow / Delta Live Tables

- Databricks Workflows

- Databricks Asset Bundles

- Data transformation and optimisation

Data pipelines

- ETL/ELT pipeline design
- Data ingestion (Batch and Real-time)
- Azure Data Factory
- Workflow orchestration
- Change Data Capture (CDC)

Performance tuning and optimisation

- Spark performance tuning

- Query optimisation

- Data partitioning strategies

Programming

- Python

- Scala

- SQL

- Go (Golang)

Data quality & governance

- Data Quality and Data Governance

DevOps

- Docker

- Kubernetes

- CI/CD and GitHub Actions

Cloud

- Azure

- GCP

Langues

English
Bilingue
C2
French
Bilingue
C2

Certificats

  • Spark certified
  • Azure Data Engineering certified

Chronologie

Lead Data Engineer

Dscale
2021.03 - 2026.07

Data Engineer

ZAND Bank
2019.12 - 2021.03

Data Engineer

Societe Generale
2016.08 - 2019.11

Data Engineer/Business Intelligence

Capgemini
2012.04 - 2016.04

Double Degree, MSc - Business Intelligence

Grenoble Management School-ESC Grenoble

Masterʼs Degree - information systems for bank, finance and industry

ESIEA (École Supérieure d'Informatique, d'Électronique et Automatique)
Djim DIOP