I'm a data engineer working on two production platforms at once. At Onyxes I run the telecom Big Data Platform for Zain Iraq on Cloudera (Hadoop/Hive/Tez), which ingests about 185 million CDRs a day across 9 ETL sources and is reconciled against Oracle ExaData. At Polaris I build lakehouse pipelines: NiFi ingestion into MinIO and Apache Iceberg, Bronze → Silver → Gold layers orchestrated by Airflow, with Spark and Trino for compute. I care about data you can trust: automated reconciliation, zero-downtime migrations, and documentation the next engineer can actually use.
185MCDR records / day
5.7Brecords in one month
9production ETL pipelines
450+columns per schema
Experience
Data engineer on Zain Iraq's telecom Big Data Platform: large-scale CDR processing on Cloudera CDH 7.1.9 (Hadoop / Hive / Tez), with Oracle ExaData running as the parallel system of record.
- Own and operate 9 continuous ETL pipelines (Voice, SMS, DAT, MSC, SMSC, Smart_SMS, PCRF) on Hive/Beeline with Kerberos and SSL
- Ingest ~185M SMSC/BulkSMS records a day, peaking at 252M, for 5.7 billion records in March 2026 alone
- Built automated reconciliation between Big Data and ExaData that compares record counts, revenue and files across every source
- Led the CBS FreeUnitMSISDN production migration: live schema changes across Voice, SMS and DAT pipelines with zero downtime (go-live May 2026)
- Built the DAT CDR HIST migration framework: a 4-stage Oracle → Hive pipeline with 450-column schemas, dynamic partitioning and Sqoop extraction
- Backfill missing CDR files from ExaData into Hive and keep shared dimensions (e.g. cell sites) in sync between Oracle and Hive
- Handle ETL incident response for stage-file bugs, partition mismatches, OOM failures and SSL enforcement changes
- Documented all 15 operational areas of the platform for a clean knowledge handover
Stack: Hadoop, Hive, Tez, Cloudera CDH, Beeline, Sqoop, Oracle ExaData, Kerberos, Bash, SQL
Build lakehouse data pipelines for national-scale statistical data, from raw ingestion to curated, analysis-ready tables.
- Design and build scalable pipelines for ingestion, transformation and analytics
- Build real-time and batch ingestion flows with Apache NiFi
- Model Bronze → Silver → Gold medallion layers on MinIO and Apache Iceberg, with fact and dimension tables
- Orchestrate everything with Apache Airflow for reliable, monitored and repeatable runs
- Publish Gold tables with Write-Audit-Publish on Iceberg branches, so data reaches dashboards only after it passes checks
- Transform and query with Spark and Trino, and manage the performance and layout of the MinIO data lake
Stack: Airflow, NiFi, MinIO, Iceberg, Spark, Trino, Python, SQL
Where it started: analytics workflows, spatial data and process automation in Alteryx.
- Built analytic apps that automated manual processes and improved workflow efficiency
- Integrated spatial analytics into projects to deepen the analysis
- Earned Alteryx Designer Core and Designer Advanced certifications
Stack: Alteryx, Spatial analytics, SQL
Selected projects
Telecom CDR ingestion platform 185M rows/day
Onyxes · Zain Iraq
Problem: Nine high-volume telecom sources (Voice, SMS, DAT, MSC, SMSC, Smart_SMS, PCRF) have to land in Hive continuously, securely and on time.
Built: I operate 9 continuous ETL pipelines on Cloudera CDH 7.1.9 that load through Hive/Beeline with Kerberos/SSL into multi-year partitioned tables with up to 450+ columns.
Impact: ~185M records a day, peaking at 252M. 5.7B records in March 2026.
Big Data ↔ ExaData reconciliation Every source checked
Onyxes · Zain Iraq
Problem: Two parallel platforms (Hadoop and Oracle ExaData) must agree, and silent gaps in records or revenue are costly.
Built: Automated multi-source reconciliation scripts that compare record counts, revenue and file-level manifests per source and per day, then backfill missing files from ExaData via Sqoop.
Impact: Discrepancies show up as concrete, fixable lists instead of surprises in reports.
CBS FreeUnitMSISDN migration 0 downtime
Onyxes · Zain Iraq
Problem: A billing-system change required live schema changes across Voice, SMS and DAT, without stopping ingestion.
Built: Planned and coordinated the schema changes and pipeline cutover across all three sources inside one maintenance window.
Impact: Went live in May 2026 with zero downtime and no data loss.
DAT CDR HIST migration framework 450-col schemas
Onyxes · Zain Iraq
Problem: Years of historical DAT CDRs sat in Oracle and needed to move to Hive with exact schema fidelity.
Built: A repeatable 4-stage Oracle → Hive framework: Sqoop extraction → staging → cleanup (null sentinels, types) → dynamic-partition insert into history tables.
Impact: A reusable process for 450-column tables and any future backfill.
Lakehouse medallion pipelines Bronze → Gold
Polaris Technology
Problem: Raw survey and statistical data needs to become clean, modeled, trustworthy tables for analysts and dashboards.
Built: NiFi ingestion into MinIO, Iceberg Bronze/Silver/Gold layers with fact and dimension tables, Spark/Trino transforms, Airflow DAGs, and Write-Audit-Publish for Gold.
Impact: Bad data is caught on a branch before it's published, and every layer can be rebuilt from the one below it.
Skills
Big Data Hive / Hadoop / Tez
Quality Data Reconciliation
Storage Data Lakes (MinIO, Iceberg)
Query SQL / Oracle / Trino
Orchestration Apache Airflow
Ingestion Apache NiFi
Migration Sqoop / ETL migration
Code Python
Compute Spark
Security Kerberos / SSL ops
Certifications
- Cloudera Technical Professional (CTP) Accreditation 2.0
- MinIO Accredited AIStor Iceberg Specialist
- Tools for Data Science V2
- The Complete Programming Course — Python
- Alteryx Designer Core & Designer Advanced
Education
B.Sc. Computer Science, Applied Science University (2021 — 2025)
1st Place — ICE (Innovation & Entrepreneurship) competition on campus
Languages
Arabic — Native · English — Professional working
Leadership & community
- Licensee & Lead Organizer, TEDxASPU (Dec 2024 — Aug 2025): Led 40+ volunteers, 400+ attendees, +40% registrations
- Ambassador, DLytica — Data Analytics & AI (Sep 2024 — Jan 2025): Promoted Data & AI training programs across Jordan
- Ambassador, Tech3arabi.com (Aug 2024 — Jan 2025): Organized tech community events