Skip to content
Data science · ML engineering · AI engineering

Models thatdecide in real time.

Statistics with assumptions stated up front, pipelines that keep data trustworthy over time, and systems that answer in milliseconds and stay calibrated as behaviour shifts. Shipping to production is not the finish line: it is where the work starts.

Project footprint8 departments

AntioquiaAtlánticoBogotáBoyacáCundinamarcaLa GuajiraMetaValle del Cauca

+6
Years of experience

Applied statistics, data engineering and AI

+30
Projects delivered

Modelling, analytics and information systems

12
Systems built

Live and in development, with real users

Explore the data lab

What I've built

Each one end to end: data model, backend, interface and rollout.

Organizations

  • Ministerio del Interior
  • Policía Nacional
  • Ministerio de Igualdad
  • Comisión Legal para la Equidad de la Mujer
  • Gobernación de Boyacá
  • Alcaldía de Cali
  • Gobernación del Magdalena

Services

System active

9 services · 3 disciplines

/Data science

Ambiguous questions turned into defensible estimates, with the assumptions and the uncertainty in plain sight.

03 services

DSC·01

Experiment design and analysis

For knowing whether a change actually worked, rather than merely coincided with something. The metric and the sample size are fixed before launch, and the result is read by segment as well as on average.

Causal inferenceStatistical powerHeterogeneous effects

Delivers: Protocol · Analysis · Documented decision

Ask about this service
DSC·02

Inference on complex survey designs

So a survey with a complex sampling design is not read as if it were a simple random sample. Expansion factors, non-response handling, and intervals computed on the actual design.

Survey designExpansion factorsConfidence intervals

Delivers: Estimates · Methodology write-up · Reproducible code

Ask about this service
DSC·03

Predictive models with calibrated probability

For decisions that depend on how much probability there is, not only on which option is likelier. The model ships calibrated, with its calibration curve in view.

CalibrationClass imbalanceTemporal validation

Delivers: Model · Calibration metrics · Limits of use

Ask about this service

/Machine learning engineering

Models that leave the notebook and run: bounded latency, fresh data, behaviour under watch.

03 services

MLE·01

Models in production on a latency budget

For taking a model from the notebook to an API that answers within an agreed time. Latency is measured at the 50th and 95th percentiles, which is where the tail lives.

Servingp95 latencyContainers

Delivers: API · Container · Service metrics

Ask about this service
MLE·02

Data pipelines and reproducible features

So the model sees the same data when it trains and when it runs. Ingestion, transformation and versioned features under a single definition.

IngestionFeature versioningData tests

Delivers: Pipeline · Data contracts · Automated tests

Ask about this service
MLE·03

Drift monitoring and retraining

For catching that the model has started to fail before the business does. Drift and performance monitoring, a defined retraining cadence, and control of the bias the model itself introduces.

Data driftRetrainingFeedback bias

Delivers: Monitoring dashboard · Alerts · Retraining policy

Ask about this service

/AI engineering

LLM systems measured before they are promised: evaluated retrieval, cost and failure modes.

03 services

AIE·01

RAG systems with evaluated retrieval

For answering over your own documentation while knowing how well it retrieves, not just how well it writes. Retrieval and generation are evaluated separately.

RAGRetrieval evaluationContext management

Delivers: System · Evaluation set · Baseline

Ask about this service
AIE·02

Agents with tools and guardrails

So an agent can query and act on real systems without being able to break anything. Read-only permissions, execution caps, and validation before every action.

Tool useGuardrailsTraces

Delivers: Agent · Documented guardrails · Trace log

Ask about this service
AIE·03

Evaluation suites for LLM systems

For knowing whether the system improved or merely changed. Labelled cases, a baseline to compare against, a failure taxonomy with percentages, and cost and latency per query.

EvalsFailure modesCost per query

Delivers: Eval suite · Failure report · Cost and latency

Ask about this service

Method

  1. 01

    Problem and metric

    Before touching any data: which decision depends on this, which metric would prove it, and which result would refute it. A question that cannot be refuted is not a question.

  2. 02

    Data and baseline

    Sources, data contracts and features defined identically in training and inference. Plus an explicit baseline: with nothing to compare against, any result looks good.

  3. 03

    Modelling and evaluation

    The evaluation set comes before the model. Failure modes are documented with percentages rather than anecdotes, and the write-up states what these data cannot support.

  4. 04

    Putting it into operation

    Served behind an API with a stated latency budget, calibrated probabilities, and explicit guardrails on what the system may and may not do.

  5. 05

    Monitoring and retraining

    Data drift, performance decay, and the bias the system itself introduces into what it will see tomorrow. From here the process returns to the first phase: the metric that worked stops working.

About

Hermes Bonilla

Hermes Bonilla

Data scientist · Bogotá, Colombia

Six years working with data across two worlds that rarely meet: companies such as Aeroméxico, Farmatodo and HDI Seguros, and public institutions in Colombia. In the first, the hard part is scale. In the second, that the data is sensitive and mistakes are costly.

I design the models, build the pipelines that feed them, and keep them scoring in real time. Once an automated decision depends on the value of a probability, calibration stops being a methodological detail.

Data science

Inference, experimentation and predictive modelling

Machine learning engineering

Served models, pipelines and monitoring in production

AI engineering

LLM systems, retrieval and evaluation suites

Systems development

End-to-end platforms, from schema to production

Let's talk about your project

Tech stack

The tools I use to build models, pipelines and systems that run.

Next.js

Web apps and platforms

Python

Statistical analysis and modelling

R

Statistical analysis and visualization

PostgreSQL

Relational databases

Neon

Serverless Postgres

AWS

Cloud infrastructure

Google Cloud

Cloud infrastructure

Cloudflare

Network, DNS and security

Docker

Containers and deployment

Claude / OpenAI

Generative AI and LLMs

LangChain

Agent orchestration

Databricks

Data and ML platform

Power BI

Visualization and dashboards

GIS / QGIS

Geospatial analysis

Git / GitHub

Version control

Cursor

AI-assisted development

Auth.js

Authentication and security

Vercel

Deployment and CI/CD

Supabase

Cloud backend

Let's build the solution

What needs building, what evidence backs it, and when it goes live.

Start a technical conversation

A first conversation to understand the problem, at no cost and with no commitment.

Message on WhatsApp

Confidential

Project information and context are handled with absolute discretion.

No commitment

I review the case and say honestly whether I can add value or not.

Location
Bogotá, ColombiaAvailable across Latin America