Author of Production Data Science with R

Data Engineer.
Author.
Builder.

6+ Years Experience · API & ETL Systems · Published Author · Technical Trainer

I design production-grade APIs and data systems, teach others to do the same, and published a comprehensive book on building production data science pipelines with R. From RESTful API architecture to ETL orchestration and multi-step approval workflows, I build structured, observable systems that handle complex real-world interactions.

6+
Years Experience
8
Book Chapters
3
Published Packages
4
Years Training

From Theory to Production

I do not just analyze data — I build systems that make data reliable, scalable, and actionable.

I am Musa Rioba, a Data Engineer with a BSc in Actuarial Science from Karatina University and 6+ years of hands-on experience designing production APIs, ETL pipelines, and multi-step workflow systems across freelancing, technical training, and open source.

I wrote "Production Data Science with R" — a comprehensive guide covering end-to-end ML pipelines, automated ETL with targets, real-time fraud detection with Redis, interactive dashboards with Shiny, and robust A/B testing frameworks. Every chapter follows production best practices: reproducibility, scalability, observability, and safety. My background in actuarial science and technical training has honed my ability to break down complex tasks into precise, logical steps — a skill I apply to API design and structured system interactions.

Beyond the book, I have published three open-source R packages (ugcensus, sacensus, kenyacensus), built production APIs and dashboards with role-based access control, and spent 4 years as a Technical Trainer teaching data analysis, API design, and programming to over 100 students and professionals.

Published Author

"Production Data Science with R" — 8 chapters on production-grade data systems

Pipeline & API Architecture

RESTful API design, ETL orchestration, multi-step workflows, ML pipelines, and automated systems

Technical Training

4 years teaching data analysis, R, Python, and production best practices

Open Source

3 published R packages, live dashboards, production APIs, and production-tested tools

Production Data Science with R

A hands-on guide for data scientists who want to move beyond notebooks and build systems that run in production.

Production Data Science with R
End-to-end systems for real-world data science · Built with bookdown
1
End-to-end ML Pipelines (tidymodels, xgboost)
2
Interactive Dashboards (shiny, plotly)
3
Marketing Attribution (Markov chains)
4
Demand Forecasting (Prophet + XGBoost ensembles)
5
Customer Segmentation (RFM, K-Means)
6
Real-time Fraud Detection (Redis, isolation forests)
7
Automated ETL Pipelines (targets orchestration)
8
A/B Testing Frameworks (sequential, Bayesian)
Reproducibility Scalability Observability Safety
Read the Book

What I Have Built

Production-ready tools and packages that handle the full data pipeline from raw sources to interactive outputs.

Telematics Kenya — UBI Risk Platform
Flutter · Dart · R · Plumber · PostgreSQL · Redis · Docker · Fly.io

A full-stack usage-based insurance (UBI) platform for the Kenyan motor insurance market. Features a Flutter mobile app and a production REST API backend with 10+ structured endpoints handling authentication, scoring, driver profiles, and audit trails — designed for reliable request/response patterns and stateful multi-step interactions.

  • Flutter Mobile App — “Telematics Kenya” with real-time trip tracking, risk score gauge, fleet overview, and driver profile management
  • Web Backend Dashboard — Live admin portal at fly.dev with driver profiles, risk scoring, trip history, and interactive trip location maps
  • Risk Scoring Engine — Weighted ensemble model (Gini 0.41, AUC-ROC 0.73) with SHAP-like explanations per prediction
  • RESTful API — R Plumber backend with 10+ endpoints for scoring, driver profiles, model performance, compliance, and audit trails
  • GDPR & Compliance — Pseudonymization, consent logging, Article 17 right-to-erasure, fairness audits (gender, age, region)
  • Full-Stack Deployment — Flutter app on mobile + Docker Compose stack (Nginx, R API, PostgreSQL, Redis) deployed on Fly.io
2Platforms
10+API Endpoints
LiveOn Fly.io
View on GitHub Live Backend
End-to-End Data Engineering Pipeline
Python · Airflow · Spark · dbt · Snowflake · Docker

A production-grade ETL/ELT data platform that ingests raw sales, customer, and product data from multiple sources, cleans and transforms it, and delivers analytics-ready datasets for business reporting. Designed to run on a standard 16GB RAM laptop.

  • Apache Airflow (CeleryExecutor) — 3 DAGs orchestrating hourly ETL with task dependency graphs, XCom, and automatic retry policies
  • Dual-Mode Processing — Pandas for <1GB datasets, PySpark for 1-50GB datasets with adaptive query execution
  • dbt Medallion Architecture — Staging to Intermediate to Marts with incremental fact tables, SCD snapshots, and custom macros
  • Multi-Source Ingestion — CSV extractor, Snowflake extractor (key-pair auth), REST API extractor with retry logic
  • Data Quality Gates — Runtime validators, dbt tests (uniqueness, not-null, positive revenue), structured JSON logging
  • Docker Compose Stack — Full environment with Airflow + Spark + PostgreSQL + Redis, tuned for 16GB RAM
  • CI/CD — GitHub Actions, pre-commit hooks, pytest with unit/integration/Spark test matrix
3Airflow DAGs
6dbt Marts
FullStack
View Code
TrialView Pro
R Shiny · Docker · CI/CD · SQLite

A production-ready Phase II diabetes clinical trial monitoring system using CDISC SDTM data. Features a multi-step role-based access workflow (Investigator → Monitor → Admin) with state transitions, audit logging, and safety tracking — deployed via Docker and Shinyapps.io with GitHub Actions CI/CD.

  • Role-Based Access Control — Investigator, Monitor, Admin with scoped permissions
  • Enrollment & Safety Monitoring — AE/SAE tables, SOC drill-down, CSV export
  • Efficacy Tracking — Glucose trajectories, responder analysis, lab results
  • Patient Management — Longitudinal profiles, new subject entry, baseline data
  • Audit Logging — Full SQLite-backed audit trail for compliance
  • CI/CD Pipeline — GitHub Actions with linting, styling, data integrity tests
  • Docker Deployment — Multi-stage Dockerfile + docker-compose
CDISCSDTM Standard
SQLiteDatabase
RBAC3 User Roles
Live Demo View Code
Sales Analytics Dashboard
R Shiny · bslib · plotly · Live App

A production-ready business intelligence dashboard for sales analytics and executive reporting. Features reactive filters, KPI cards with trend indicators, interactive Plotly charts, sortable data tables, and one-click CSV export — all in a responsive dark theme.

  • Reactive Filters — Region, Product, and Time Period update all charts instantly
  • KPI Cards — Revenue, Users, Conversion Rate, Avg. Order Value with trend indicators
  • Interactive Charts — Revenue trends, regional donut chart, product bar chart (Plotly)
  • Sortable Data Table — Recent transactions with styled status badges (DT)
  • One-Click Export — Download filtered data as CSV directly from the dashboard
  • Dark Theme UI — Modern professional design using bslib with custom CSS
  • Fully Responsive — Works on desktop, tablet, and mobile
LiveShinyapps.io
PlotlyInteractive
BISales Analytics
Live Demo View Code
Kenya Insurance Hub
R Shiny · bslib · plotly · Live App

A comprehensive insurance marketplace platform for Kenya's insurance market. Features 4 premium calculators, company directory, market analysis, regulatory compliance tracking, and NHIF benefits — all in one interactive dashboard.

  • 4 Premium Calculators — Motor, Health, Life, and Agriculture with real-time quote comparison across 28 insurers
  • Company Directory — Filterable directory of 28 licensed insurers with financial ratings, contact details, and product listings
  • Market Analysis — Interactive bar, pie, and treemap charts for market share, ratings, and product distribution
  • NHIF Benefits — Full coverage details for 19 NHIF benefits with co-payment and waiting period info
  • Regulatory Compliance — IRA contacts, complaint procedures, agent verification, and unlicensed scheme warnings
  • 8 Structured Datasets — 28 companies, 69 products, 41 motor premiums, 21 health premiums, 27 life premiums, 12 agriculture products, 19 NHIF benefits
  • Consumer & Researcher Focused — Built for consumers shopping for insurance, researchers analyzing trends, and policymakers tracking industry performance
28Insurers
69Products
11Tabs
Live Demo View Code
ugcensus
R Package

A comprehensive toolkit for accessing, analyzing, and visualizing Uganda census data from the Uganda Bureau of Statistics (UBOS). Supports 1991, 2002, 2014, and 2024 censuses.

  • Multi-level geography: national, regional, district, county
  • Built-in data cleaning & standardization utilities
  • Interactive map plotting & population pyramids
  • Population projection tools based on growth rates
  • Time-series comparison across census years
4Census Years
17Sub-Regions
45.9MPop. (2024)
View on GitHub
sacensus
R Package

A comprehensive toolkit for accessing, analyzing, and visualizing South African census data from Statistics South Africa (Stats SA). Supports 1996, 2001, 2011, and 2022 censuses.

  • Multi-level geography: national, provincial, municipal, ward
  • Demographic analysis: population, housing, education, employment
  • Trend visualization & population pyramid generation
  • Data cleaning utilities for standardization
  • Population projection based on growth rates
4Census Years
9Provinces
62MPop. (2022)
View on GitHub
kenyacensus
R Package

A comprehensive toolkit for accessing Kenya's Population and Housing Census data from the Kenya National Bureau of Statistics (KNBS). Covers 1948 to 2019 with 13 themed datasets.

  • 8 historical census years: 1948, 1962, 1969, 1979, 1989, 1999, 2009, 2019
  • 13 datasets: demographics, education, ethnicity, disability, water, ICT, employment
  • 47 counties with full 2019 socio-economic data
  • Built-in visualization: trends, pyramids, county comparisons, maps
  • Growth rate calculations and county-level filtering
  • CC0 Public Domain license
8Census Years
47Counties
13Datasets
View on GitHub

My Toolkit

Deep expertise in R-based production systems, with expanding skills in SQL, cloud, and orchestration.

Languages

R (Expert) Python SQL Bash HTML/CSS Flutter / Dart JSON / Structured Data

Data Engineering & Pipelines

targets (ETL Orchestration) ETL/ELT Pipelines REST API Design Multi-step Workflow Design Data Cleaning Data Modeling Web Scraping API Integration

Machine Learning & Analytics

tidymodels XGBoost Prophet K-Means / RFM Markov Chains Isolation Forests Risk Scoring

Dashboards & Visualization

R Shiny ggplot2 Plotly shinydashboard DT (DataTables)

Production & Deployment

Docker Compose GitHub Actions Shinyapps.io plumber (APIs) Nginx API Authentication Git/GitHub pre-commit

Databases & Cloud

Snowflake PostgreSQL Docker Redis BigQuery (bigrquery) SQLite AWS (Learning)

Experimentation & Testing

A/B Testing Sequential Testing Bayesian Methods Statistical Modeling Actuarial Methods Technical Writing

Growing Into

Kafka Terraform Kubernetes AWS/GCP Production

Resume

Download my updated resume in PDF or Word format. Print-ready and professionally formatted.

Musa Rioba — Resume
Data Engineer & Author · 6+ Years Experience

Times New Roman, 12pt. Clean single-column layout. Includes all projects, skills, and experience.

1 Page PDF + Word

Let Us Build Production Systems Together

Interested in working together?

I am open to opportunities where I can apply my expertise in API design, structured data systems, multi-step workflow logic, and precise technical communication to build high-quality, production-grade solutions.