Matiwos Desalegn
Senior Data Engineer
Professional Summary
Senior Data Engineer with production experience at Kifiya Financial Technology designing and operating cloud-native data infrastructure for African fintech markets. Builds end-to-end batch and streaming pipelines with Apache Spark (PySpark), Apache Airflow, Apache Kafka, and dbt on AWS (EMR, Glue, Redshift, Athena, S3) - delivering 99.9% pipeline reliability, 40% cost reduction, and a multi-source ingestion platform serving institutional partners. Deep expertise in dimensional data modeling, medallion architecture, and lakehouse governance on Redshift and Iceberg.
Experience
Senior Data Engineer
Kifiya Financial Technology PLC
- Owns and governs the multi-source S3 medallion lakehouse - Bronze (raw snapshots and CDC), Silver (typed Apache Iceberg tables), Gold (Redshift-ready aggregates); 30+ Airflow DAGs on EKS
- Built the lakehouse governance plane: OpenLineage to Marquez lineage, OpenMetadata PII cataloging, Apache Ranger + Trino for role-based access and query-time masking
- Leads cross-team knowledge transfer - data-science to data-engineering feature-engineering handoff templates and internal ML enablement for analytics self-service
Data Engineer
Kifiya Financial Technology PLC
- Built the organization's first AWS-managed lakehouse (Athena + Glue Data Catalog) orchestrated with Mage AI - 6-hour micro-batch ingestion for institutional partners, 300+ tables / 7,000+ columns, high-watermark extraction, and date partitioning
- Designed a lightweight real-time CDC pipeline under compute constraints - OLake for MySQL binlog to Apache Iceberg (Polaris catalog), queryable via Trino
- Built and deployed the Batch API microservice (FastAPI + EKS) ingesting domain payloads from partner institutions; persists to MongoDB and pushes features to Feast (Redis) for credit-scoring consumption
- Built production Firebase Analytics export pipelines - BigQuery to S3 Bronze Parquet to Iceberg Silver via PyIceberg, with partner-isolated Airflow DAGs and Secrets Manager + IRSA
Software Developer
HST Consulting PLC
- Designed and implemented the frontend of an enterprise HR/payroll (ERP) system, improving usability and operational efficiency
- Integrated RESTful APIs with backend teams for reliable, end-to-end data flow
- Achieved a 40% reduction in payroll processing time through optimized data handling and a responsive interface
Data Science Fellow
10 Academy
- Intensive 12-week data science and machine learning fellowship program
- Worked on real-world data science projects with industry mentors
- Specialized in MLOps, feature engineering, and model deployment
- Collaborated in agile teams on end-to-end ML projects
Google Developer Student Clubs (GDSC) Lead
Wolaita Sodo University
- Led a community of 200+ developer students in learning cutting-edge technologies
- Organized workshops and events on cloud computing, AI/ML, and web development
- Mentored students in building real-world projects and technical skills
- Facilitated hackathons and coding competitions
Technical Skills
Key Projects
World Bank CDC Analytics Platform
Data EngineeringEnd-to-end analytics engineering pipeline with PostgreSQL OLTP, Debezium CDC, Redpanda, ClickHouse OLAP, dbt marts, Airflow orchestration, and Grafana observability plus a live demo environment.
~3K rows ingested from paginated REST API · CDC replication lag monitored with cdc-monitor alerts
Multi-Bank Digital Lending ETL Orchestration
Data EngineeringMage AI-orchestrated ETL system ingesting loan data from 5 Ethiopian partner bank Metabase instances every 6 hours into AWS S3, with PAR (Portfolio at Risk) transformation and Athena-queryable cleansed layers for digital lending analytics
6-hour ingestion cadence across 5 partner banks · Unified RAW to Application to Cleansed S3 layer structure
Firebase Analytics Data Pipeline
Data EngineeringAutomated daily pipeline exporting 600K–800K Firebase Analytics events from Google BigQuery to AWS S3 in Parquet format with Snappy compression
600K–800K events exported daily · ~70% storage reduction via Parquet + Snappy
Credit Scoring Batch API
Data EngineeringProduction FastAPI microservice for batch ingestion of 18+ financial and behavioral data domains, powering credit scoring for 5 Ethiopian partner banks via Feast Feature Store and ML pipelines
5–10 pod autoscaling under load · 99.9% pipeline reliability
Multi-Source Digital Lending Data Platform
Data EngineeringProduction multi-source data platform ingesting core banking, sharia-compliant, operational, and portfolio risk data from 5+ partner banks into Aurora PostgreSQL and Amazon Redshift, powering credit risk, portfolio analytics, and Power BI dashboards for a digital lending business
5 source systems, daily full-refresh ingestion · Zero-downtime atomic table swap for all partner banks
Digital Lending Dimensional Data Modeling
Data EngineeringEnd-to-end dimensional modeling project for a digital lending platform - from conceptual and logical design through physical Star Schema implementation (fact and dimension tables) in Amazon Redshift - covering Transient Staging, Core Canonical Data Model, and a Semantic Layer with Data Contracts for analytics, credit risk, and BI consumption
99.9% uptime - Kubernetes (EKS) containerization · CI/CD via GitHub Actions + GitLab CI/CD; GDPR-compliant anonymization
Certifications & Education
10 Academy · Remote
Intensive 12-week program — graduated top 10% of cohort. Specialized in MLOps, feature engineering, and end-to-end ML project deployment.
Publications
Full list: medium.com/@mathiwossilvio