Arenesha AI Labs

GPU training and inference infrastructure, and the console used to keep an eye on it.

arenesha.ai
Client
2
Deploy pipelines
GPU
Workload class
24/7
Job oversight
FastAPIPyTorchCUDAAWS EC2S3RedisNext.js

The problem

Model training jobs were being launched by hand and nobody could see the fleet. A job that wedged held a GPU until someone noticed, which meant paying for idle hardware and having no answer for the customer waiting on the run. The failure mode that cost most was not a crash but a slow one: a run misconfigured at the start would consume an hour of GPU time before failing on something that could have been caught in the first second.

What we built

  • FastAPI service orchestrating PyTorch and CUDA training and inference workloads on EC2, with S3 for artefacts and Redis for queueing
  • Operations console: live run status and ownership, active-user visibility, and termination of blocked jobs
  • Front end on branch-based deployment pipelines, development to staging and main to production
  • Search, analytics and permanent redirect handling across the public properties
  • Background workers so training and inference run off the request path, survive a deploy, and can be inspected while still running
  • A model validation service that checks a submitted model before it is allowed to consume GPU time
  • A training assistant that guides configuration of a run, reducing the number of jobs that fail on a bad parameter after an hour of compute
  • Versioned schema migrations, so database changes ship with the service rather than being applied by hand

Decisions worth explaining

A wedged job should not need a human to notice it

Runs were being launched by hand with no fleet view, so a stuck job held a GPU until somebody happened to look. Job state, ownership and progress are tracked centrally and surfaced in an operations console with the ability to terminate a blocked run, which turns idle-hardware cost into something visible and actionable.

Cheap checks before expensive compute

The costly failure was not a crash but a slow one: a run misconfigured at the start burning an hour of GPU before failing on a parameter that could have been rejected immediately. Validation runs before a job is admitted to the queue, so the class of error that used to cost an hour now costs a second.

An operations console is for intervening, not just watching

A dashboard that shows a stuck job without offering a way to kill it just moves the problem to whoever is reading it. Run state carries ownership and the console can terminate a blocked run directly, which is what turns visibility into a shorter mean time to recovery rather than a better-informed wait.

Training work belongs off the request path

GPU training cannot run inside a web request. Work is queued and executed by background workers, so the API stays responsive, jobs survive a deploy, and long-running work can be inspected while it runs rather than only after it fails.