PROJECTS

Things I’ve built and learned.

Personal projects and proofs of concept, each with the reasoning behind it — plus the learning collections I keep along the way.

01 / FEATUREDHOW I CHOOSE AND BUILD.

01Proof of concept + article

Config-driven ML pipeline

ML pipeline overview

A reusable machine-learning pipeline on Airflow, MLflow, FastAPI, and Docker. Adding a dataset takes a new config folder, not new code — three demo pipelines (hospital readmissions, gene expression, energy load forecasting) run on the same codebase.

Judgment call: in a hospital-readmission forecasting experiment, I investigated a feature that reflected past readmission performance, removed it to test the remaining predictors, and documented the lower score and the limits of the model.

Read the forecasting context

This experiment predicts hospitals’ 2025 pneumonia excess readmission ratios from 2024 quality measures. Removing a summary of historical readmission performance reduced test R² from approximately 0.22 to 0.05–0.07. This shows the limited predictive strength of the remaining features; it is not evidence of deployment readiness. A historical feature may be valid if it is available at prediction time. The feature’s provenance, timing, and comparison with a historical-performance baseline need to be considered when interpreting the result.

  • Airflow
  • MLflow
  • FastAPI
  • Docker
  • CI

02Proof of concept + article

Workflow automation & testing with Prefect

Workflow testing overview

Uses Prefect orchestration to automate API-level and end-to-end integration testing of data workflows, with scheduled and automation-triggered worker pools running in containers.

Why it matters: a lightweight, flexible pipeline keeps data reliable for analysis and ML/AI work — because better AI starts with better data.

  • Prefect
  • Python
  • Containers

03Reference implementation

RAG technique reference

RAG comparison overview

Twelve working retrieval-augmented generation techniques — from naive RAG to Self-RAG, Graph RAG, and RAPTOR — each built in both LangChain and LlamaIndex, run on a local LLM, and scored with RAGAS.

Why it matters: choosing a RAG design means comparing trade-offs side by side, not defaulting to one pattern.

  • LangChain
  • LlamaIndex
  • ChromaDB
  • RAGAS
  • Local LLM

04Comparison study + article

LLM & RAG evaluation, compared

Evaluation study overview

Evaluates the same LLM and RAG systems with DeepEval, MLflow, and Ragas on local models, then maps which framework fits which job.

Why it matters: AI output should be measured, not assumed — and the right evaluation tool depends on the question.

  • DeepEval
  • MLflow
  • Ragas
  • Ollama
  • LM Studio
02 / LEARNINGCURIOSITY BECOMES PRACTICE.

Learning collections.