ML systems · CI reliability · open source

I build the feedback loops that keep PyTorch compatible.

I'm Subin George, a software engineer and PyTorch contributor. I build dependable ML systems, CI platforms, and developer tooling—from model validation and computer vision to the infrastructure that keeps upstream changes and hardware backends moving together.

Explore selected work
CRCR / production path
WEBHOOK → HUD
PyTorchsigned PR / push
CRCR relayallowlist + dispatch
Backend CIbuild + test
CRCR callbackOIDC verified
HUD ingestionDynamoDB / ClickHouse
65+test-infra PRs and counting
45+PyTorch PRs and counting
L1–L4partner trust model

About me

01 / 04

Engineering dependable paths through complex systems.

I'm Subin George, a software engineer and open-source contributor based in India. At Red Hat, I work on CI/CD and ML infrastructure, with a focus on the systems that help PyTorch and its out-of-tree hardware ecosystem evolve without losing compatibility.

I enjoy the practical work between an ML signal and a confident decision: verifying who produced a result, preserving the evidence behind it, and giving engineers a clear way to act. My work spans applied machine learning, computer vision, CI/CD, Python, C++, TypeScript, cloud services, and the developer experience around them.

I have worked across model development, edge and GPU validation, cloud ML platforms, and open-source infrastructure—bringing the same emphasis on observability, reproducibility, and useful automation to each layer.

Visit my GitHub profile Connect on LinkedIn
Focus
ML systems, PyTorch CI, developer tooling
Systems
CRCR, HUD observability, model validation, computer vision
Stack
Python, C++, TypeScript, TensorFlow, PyTorch, AWS, ClickHouse

The problem I care about

Compatibility should be discovered while a change is still actionable—not during a release or a week of bisection.

My work sits at the intersection of CI reliability, backend onboarding, and developer experience: make the signal trustworthy, make failures explainable, and keep the integration path proportional to a partner's maturity.

Selected work

Systems with a clear path from signal to owner.

Showing all projects

01PyTorch ecosystem

Cross-Repository CI Relay

A tiered relay that forwards PyTorch change events to out-of-tree backend CI and returns authenticated compatibility results to the PyTorch CI HUD.

  • OIDC
  • AWS Lambda
  • DynamoDB
  • ClickHouse
  • Next.js
02Downstream CI

RHEL CI for PyTorch

Nightly RHEL 9 build and test automation with targeted test selection, container publishing, and relay-backed HUD reporting.

  • Podman
  • GitHub Actions
  • Buildkite
Explore the repository
03Practical learning

PyTorch Reference Guide

A growing collection of practical modules, reference cards, and runnable examples—from tensor foundations to CI debugging.

  • Python
  • PyTorch
  • Documentation
Read the guide
04Core contribution

PyTorch & test-infra

Core fixes and CI/HUD improvements spanning numerical correctness, distributed systems, and out-of-tree backend workflows.

  • Python
  • C++
  • TypeScript
See test-infra

Career

Building from ML workloads to the infrastructure behind them.

Selected experience

  1. Red Hat · PyTorch Engineering

    Senior Software Engineer

    Enable PyTorch on RHEL through reproducible builds, CI/CD automation, and upstream compatibility validation.

    • Hermetic builds and downstream wheel delivery
    • Upstream PyTorch and test-infra contributions
    • CI reliability, CRCR, and HUD observability
  2. AMD India · ROCm testing

    Software Engineer System Designer II

    Led validation of the ROCm stack across MI and Radeon hardware for AI/ML and HPC workloads.

    • Led and mentored a twelve-person team while driving validation quality and feature delivery
    • PyTorch, TensorFlow, JAX, and inference-library validation with performance triage across kernel, user-space, and application layers
    • Jenkins and container-based workflow automation that reduced repetitive engineering effort by 60%
    • Applied LLM and retrieval-augmented generation patterns to document summarization and question-answering workflows
  3. Capgemini

    Process Lead — AI

    Led delivery of cloud-native ML systems for financial and multilingual data products.

    • Automated data labeling, multilingual translation, and SSML audio workflows with Google Cloud Platform APIs
    • Built FINRA-rule classification and image-captioning solutions using classical ML, vectorization, and UiPath automation
    • Improved imbalanced-data modeling with normalization and class weighting; built financial insight generation that reduced manual analysis by up to 75%
    • Delivered model workflows on AWS, Azure Databricks, Azure Cognitive Services, and Microsoft SQL Server
  4. Ignitarium Technology Solutions

    ML/AI Engineer

    Developed and tested computer-vision systems, model-serving workflows, and ML compiler tooling.

    • Designed lightweight neural networks for microcontroller deployment and image-based railway-infrastructure detection
    • Built scalable object-detection and classification services with TensorFlow Serving, including hand detection and segmentation with MediaPipe
    • Performed quality and performance work for custom ML compilers on AI-accelerated chips; optimized CNN, transformer, and ONNX models
  5. Ignitarium Technology Solutions

    ML/AI Intern

    Built proof-of-concept computer-vision pipelines and prepared data for model training and evaluation.

    • Object detection with YOLOv3 and image-classification experiments
    • Dataset annotation and reproducible proof-of-concept pipelines
    • Training, fine-tuning, and deployment preparation with Python, TensorFlow, PyTorch, and Jupyter notebooks

How I work

Reliable systems are designed for their difficult days.

  1. 01

    Make identity explicit

    Derive trust from verified claims and configuration—not from untrusted result text.

  2. 02

    Preserve the diagnostic trail

    A dashboard result should lead back to a run, an attempt, a callback, and an owner.

  3. 03

    Earn stronger integration

    Start observably, measure reliability, then promote partners only when the signal is ready to influence merges.

PyTorch Conference North America · Lightning talk

Scaling PyTorch's Compatibility Promise

A look at the tiered Cross-Repository CI Relay: real-time downstream validation, authenticated result ingestion, and a lower-risk path for hardware backends to join the PyTorch ecosystem.

Read the official PyTorch blog

Now

Making relay health and nightly compatibility easier to trust.

Current CRCR work focuses on surfacing stuck health probes, separating nightly-only partners from pull-request integrations, and making HUD metrics explainable from the underlying job matrix.

Explore CRCR design work

Contact

Have an AI, PyTorch, or ML infrastructure problem worth solving?

Share a little context and I'll receive your note by email. I welcome conversations about open-source infrastructure, downstream compatibility, and practical developer tooling.