Skip to content
Uzair Azhar
Menu

Uzair Azhar, Python and AI engineer

Data, AI and computer-vision systems that are deployed, measured and maintained.

I take a messy, real problem, such as unreliable data feeds or documents nobody can search, and ship the system that solves it, with the tests, monitoring and documentation that keep it running after handover.

Latest pipeline run

December 2010 drop, finished 2 minutes ago

Succeeded
Resolve3 msExtract472 msValidate696 msTransform194 msQuality gate7 msLoad2.27 s
rows read
65,004
rows published
41,713
rows quarantined
447
duplicates removed
22,844
quality score / 100
90.0
run time
3.68 s

Real invoices from a UK online retailer (UCI Online Retail II), processed on this server. 13 of 25 monthly drops loaded.

Open the live dashboard

Selected engineering projects

Each project has a working demo on real or clearly labelled data, a write-up of the design decisions, and its source code.

LiveData engineering

Data Pipeline Observatory

Problem
Monthly exports, a product master and a CRM API never agree. Duplicates, cancellations, stock write-offs and missing customer IDs silently corrupt every report built on them.
Solution
A Python pipeline that validates each row against a data contract, quarantines bad rows with reason codes, loads PostgreSQL idempotently and refuses to publish when a quality gate fails.
Technologies
  • Python
  • pandas
  • PostgreSQL
  • FastAPI
  • Procrastinate
  • Docker

Real data: UCI Online Retail II (CC BY 4.0)

Next on the roadmap

Listed honestly: these are not live yet and have no demo until they are.

  • AI CSV Analyst

    Upload a CSV, get a quality report, ask questions in plain English, see the SQL.

    In build
  • Document Intelligence and RAG Lab

    Question answering over documents with cited passages, plus invoice data extraction.

    Planned
  • Customer Support Agent

    A tool-calling agent that resolves order questions, with approval before refunds.

    Planned
  • Edge AI Benchmark Lab

    PyTorch vs ONNX vs INT8 on CPU: latency, memory and accuracy, measured.

    Planned
  • Invoice Matching Agent

    Reconciles public-sector payments against contract awards and explains exceptions.

    Planned
  • Research Search Intelligence

    From a clinical research question to concepts, MeSH terms and ranked papers.

    Planned
  • Street Scene Intelligence

    Detection, tracking, pose and pedestrian crossing-intent on open street footage.

    Planned

What I build

Python data pipelines
Ingestion from files, databases and APIs, with validation, retries and audit trails.
AI and LLM applications
Question answering over documents, natural-language interfaces to data, tool-calling agents.
Computer vision
Detection, tracking and image analysis, optimised to run on ordinary CPUs.
Automation
Replacing manual spreadsheet and inbox work with scheduled, observable jobs.
APIs and backend systems
FastAPI services with typed contracts, migrations and sensible limits.
Dockerised deployments
Reproducible builds, health checks and rollbacks on a single VPS or the cloud.
Data processing
Cleaning, reconciling and profiling messy tabular data at millions of rows.
Research and scientific AI
Medical-imaging models and literature tooling, written up properly.

Engineering capabilities

Languages and data

  • Python
  • SQL
  • pandas
  • DuckDB
  • PostgreSQL
  • pgvector
  • JSON / CSV

Backend

  • FastAPI
  • Pydantic
  • SQLAlchemy
  • Alembic
  • REST APIs
  • job queues

Machine learning

  • PyTorch
  • ONNX Runtime
  • quantisation
  • computer vision
  • medical imaging

LLM systems

  • RAG
  • embeddings
  • function calling
  • Groq
  • llama.cpp

Delivery

  • Docker
  • Linux
  • nginx
  • Git
  • GitHub Actions
  • Ansible
  • TypeScript / Next.js

Research and specialised work

Medical imaging, skin-lesion classification

SHEL: A Knowledge-Guided Hybrid Representation Learning Framework for Efficient Multiclass Skin Lesion Classification

Manuscript under review at a Springer Nature journal. Details and results will be linked here once it is published.

Related work on this site: research-literature search with medical terminology (Research Search Intelligence) and computer vision on video (Street Scene Intelligence), both on the roadmap.

How I work

  1. Problem

    A written statement of what goes wrong today, who notices and what fixing it is worth.

  2. Architecture

    The smallest design that solves it, with the trade-offs written down before any code.

  3. Implementation

    Typed, reviewed code in small increments you can run from the first week.

  4. Testing

    Automated tests on every change, including checks against real data where it exists.

  5. Deployment

    Container images built and scanned in CI, released with a health check and automatic rollback.

  6. Monitoring

    Structured logs, health checks and alerts, so problems are noticed before users report them.

Have a data or AI problem that needs to work in production?

Freelance projects, contracts and full-time roles. Send a short description of the problem and I will reply with questions or a first plan.