← Back

Matiwos Desalegn

Senior Data Engineer

Professional Summary

Senior Data Engineer with production experience at Kifiya Financial Technology designing and operating cloud-native data infrastructure for African fintech markets. Builds end-to-end batch and streaming pipelines with Apache Spark (PySpark), Apache Airflow, Apache Kafka, and dbt on AWS (EMR, Glue, Redshift, Athena, S3) - delivering 99.9% pipeline reliability, 40% cost reduction, and a multi-source ingestion platform serving institutional partners. Deep expertise in dimensional data modeling, medallion architecture, and lakehouse governance on Redshift and Iceberg.

Experience

Senior Data Engineer

Kifiya Financial Technology PLC

Jul 2026 - Present
Addis Ababa, Ethiopia
  • Owns and governs the multi-source S3 medallion lakehouse - Bronze (raw snapshots and CDC), Silver (typed Apache Iceberg tables), Gold (Redshift-ready aggregates); 30+ Airflow DAGs on EKS
  • Built the lakehouse governance plane: OpenLineage to Marquez lineage, OpenMetadata PII cataloging, Apache Ranger + Trino for role-based access and query-time masking
  • Leads cross-team knowledge transfer - data-science to data-engineering feature-engineering handoff templates and internal ML enablement for analytics self-service
Credit-scoring inference APIs with <200ms p95 latency on EKSMulti-agent RAG assistant with hybrid retrieval and SSOEstablished lakehouse governance with lineage, cataloging, and policy-based access control

Data Engineer

Kifiya Financial Technology PLC

Oct 2024 - Jun 2026
Addis Ababa, Ethiopia
  • Built the organization's first AWS-managed lakehouse (Athena + Glue Data Catalog) orchestrated with Mage AI - 6-hour micro-batch ingestion for institutional partners, 300+ tables / 7,000+ columns, high-watermark extraction, and date partitioning
  • Designed a lightweight real-time CDC pipeline under compute constraints - OLake for MySQL binlog to Apache Iceberg (Polaris catalog), queryable via Trino
  • Built and deployed the Batch API microservice (FastAPI + EKS) ingesting domain payloads from partner institutions; persists to MongoDB and pushes features to Feast (Redis) for credit-scoring consumption
  • Built production Firebase Analytics export pipelines - BigQuery to S3 Bronze Parquet to Iceberg Silver via PyIceberg, with partner-isolated Airflow DAGs and Secrets Manager + IRSA
Reduced pipeline failures ~60% by migrating Mage AI to Airflow on EKSAchieved 99.9% pipeline reliability across production data workflowsReduced data processing costs by 40% through optimization

Software Developer

HST Consulting PLC

Nov 2023 - Oct 2024
Addis Ababa, Ethiopia
  • Designed and implemented the frontend of an enterprise HR/payroll (ERP) system, improving usability and operational efficiency
  • Integrated RESTful APIs with backend teams for reliable, end-to-end data flow
  • Achieved a 40% reduction in payroll processing time through optimized data handling and a responsive interface
40% reduction in payroll processing time through optimized data handling and responsive UI

Data Science Fellow

10 Academy

Apr 2024 - Jul 2024
Remote
  • Intensive 12-week data science and machine learning fellowship program
  • Worked on real-world data science projects with industry mentors
  • Specialized in MLOps, feature engineering, and model deployment
  • Collaborated in agile teams on end-to-end ML projects
Completed 5 industry-grade ML projects with 90%+ accuracyImplemented MLOps best practices for model lifecycle managementGraduated in top 10% of fellowship cohort

Google Developer Student Clubs (GDSC) Lead

Wolaita Sodo University

Sep 2023 - Jun 2024
Wolaita Sodo, Ethiopia
  • Led a community of 200+ developer students in learning cutting-edge technologies
  • Organized workshops and events on cloud computing, AI/ML, and web development
  • Mentored students in building real-world projects and technical skills
  • Facilitated hackathons and coding competitions
Grew GDSC membership from 50 to 200+ active membersOrganized 12 technical workshops with 95% satisfaction rateLed 3 successful hackathons with 100+ participants each

Technical Skills

Data EngineeringAWS (EMR, Glue, DMS, S3, Athena, Redshift) · Apache Spark / PySpark · Apache Airflow · Mage AI · dbt Core · Apache Kafka · ClickHouse · CDC (Change Data Capture) · Apache Iceberg · Trino · Airbyte · OpenLineage / OpenMetadata
AI / MLLangChain · LangGraph · CrewAI · RAG (Retrieval-Augmented Generation) · Feast (Feature Store) · TensorFlow · PyTorch · Ollama
DevOpsDocker · Kubernetes (EKS) · GitLab CI/CD · GitHub Actions
CloudAWS Cloud Services
DatabasesPostgreSQL · Redis · MongoDB · ChromaDB
LanguagesPython · SQL · JavaScript/TypeScript · Java · Scala

Key Projects

World Bank CDC Analytics Platform

Data Engineering

End-to-end analytics engineering pipeline with PostgreSQL OLTP, Debezium CDC, Redpanda, ClickHouse OLAP, dbt marts, Airflow orchestration, and Grafana observability plus a live demo environment.

~3K rows ingested from paginated REST API · CDC replication lag monitored with cdc-monitor alerts

PostgreSQLDebeziumRedpandaClickHousedbt

Multi-Bank Digital Lending ETL Orchestration

Data Engineering

Mage AI-orchestrated ETL system ingesting loan data from 5 Ethiopian partner bank Metabase instances every 6 hours into AWS S3, with PAR (Portfolio at Risk) transformation and Athena-queryable cleansed layers for digital lending analytics

6-hour ingestion cadence across 5 partner banks · Unified RAW to Application to Cleansed S3 layer structure

Mage AIAWS S3AWS AthenaPythonDocker

Firebase Analytics Data Pipeline

Data Engineering

Automated daily pipeline exporting 600K–800K Firebase Analytics events from Google BigQuery to AWS S3 in Parquet format with Snappy compression

600K–800K events exported daily · ~70% storage reduction via Parquet + Snappy

PythonBigQueryAWS S3ParquetApache Airflow

Credit Scoring Batch API

Data Engineering

Production FastAPI microservice for batch ingestion of 18+ financial and behavioral data domains, powering credit scoring for 5 Ethiopian partner banks via Feast Feature Store and ML pipelines

5–10 pod autoscaling under load · 99.9% pipeline reliability

FastAPIMongoDBFeast Feature StoreKubernetes (EKS)Docker

Multi-Source Digital Lending Data Platform

Data Engineering

Production multi-source data platform ingesting core banking, sharia-compliant, operational, and portfolio risk data from 5+ partner banks into Aurora PostgreSQL and Amazon Redshift, powering credit risk, portfolio analytics, and Power BI dashboards for a digital lending business

5 source systems, daily full-refresh ingestion · Zero-downtime atomic table swap for all partner banks

Apache AirflowAmazon Aurora PostgreSQLAmazon RedshiftAWS S3Airbyte

Digital Lending Dimensional Data Modeling

Data Engineering

End-to-end dimensional modeling project for a digital lending platform - from conceptual and logical design through physical Star Schema implementation (fact and dimension tables) in Amazon Redshift - covering Transient Staging, Core Canonical Data Model, and a Semantic Layer with Data Contracts for analytics, credit risk, and BI consumption

99.9% uptime - Kubernetes (EKS) containerization · CI/CD via GitHub Actions + GitLab CI/CD; GDPR-compliant anonymization

dbt CoreAmazon RedshiftPostgreSQLSQLAmazon Aurora PostgreSQL

Certifications & Education

Data Science FellowApr 2024 - Jul 2024

10 Academy · Remote

Intensive 12-week program — graduated top 10% of cohort. Specialized in MLOps, feature engineering, and end-to-end ML project deployment.

DataCampDataCamp Premium - Data Community Africa Scholarship; Associate DE certifications: SQL (Jul 2026), Snowflake (Jul 2026); Databricks certification 74% complete
LinkedIn Learning16+ advanced courses completed
CodeSignal LearnMulti-format ingestion path verified; Verified assessments: Multi-format data ingestion (Python) (Sep 2025)
Credly Verified2 Credly badges plus AWS Educate coursework; Credly badges: AWS Educate Machine Learning Foundations (Jun 2025), LFC102: Inclusive Open Source Community Orientation (Aug 2026)

Publications

Full list: medium.com/@mathiwossilvio