Arenesha AI Labs
GPU training and inference infrastructure, and the console used to keep an eye on it.
arenesha.aiThe problem
Model training jobs were being launched by hand and nobody could see the fleet. A job that wedged held a GPU until someone noticed, which meant paying for idle hardware and having no answer for the customer waiting on the run. The failure mode that cost most was not a crash but a slow one: a run misconfigured at the start would consume an hour of GPU time before failing on something that could have been caught in the first second.
What we built
- FastAPI service orchestrating PyTorch and CUDA training and inference workloads on EC2, with S3 for artefacts and Redis for queueing
- Operations console: live run status and ownership, active-user visibility, and termination of blocked jobs
- Front end on branch-based deployment pipelines, development to staging and main to production
- Search, analytics and permanent redirect handling across the public properties
- Background workers so training and inference run off the request path, survive a deploy, and can be inspected while still running
- A model validation service that checks a submitted model before it is allowed to consume GPU time
- A training assistant that guides configuration of a run, reducing the number of jobs that fail on a bad parameter after an hour of compute
- Versioned schema migrations, so database changes ship with the service rather than being applied by hand
Decisions worth explaining
A wedged job should not need a human to notice it
Runs were being launched by hand with no fleet view, so a stuck job held a GPU until somebody happened to look. Job state, ownership and progress are tracked centrally and surfaced in an operations console with the ability to terminate a blocked run, which turns idle-hardware cost into something visible and actionable.
Cheap checks before expensive compute
The costly failure was not a crash but a slow one: a run misconfigured at the start burning an hour of GPU before failing on a parameter that could have been rejected immediately. Validation runs before a job is admitted to the queue, so the class of error that used to cost an hour now costs a second.
An operations console is for intervening, not just watching
A dashboard that shows a stuck job without offering a way to kill it just moves the problem to whoever is reading it. Run state carries ownership and the console can terminate a blocked run directly, which is what turns visibility into a shorter mean time to recovery rather than a better-informed wait.
Training work belongs off the request path
GPU training cannot run inside a web request. Work is queued and executed by background workers, so the API stays responsive, jobs survive a deploy, and long-running work can be inspected while it runs rather than only after it fails.