The company’s funding round, co-led by Premji Invest, Wing Venture Capital, and Seligman Ventures, brings its total capital to $73 million. Clockwork.io focuses on a critical pain point: the high rate of failure in massive GPU clusters, where a single malfunctioning component can trigger a cascade of downtime. By acting as a software layer between hardware and AI workloads, the system ensures training and inference tasks continue even when individual GPUs or network links fail.
In section Releases
Clockwork.io Secures $31M to Solve AI Infrastructure Downtime
As large-scale AI training becomes increasingly fragile, Clockwork.io has raised $31 million to scale its fault-tolerance software. The platform, which prevents massive GPU-hour losses by rerouting traffic and migrating workloads during hardware failures, is currently being deployed by enterprise leaders including LinkedIn and Together AI.

New updates to the company's TorchPass solution allow for platform-level snapshots and fast background checkpoints. These features enable teams to save the state of a distributed job without requiring modifications to the original training code. LinkedIn, which has already implemented the software’s LinkPass technology, reports saving tens of thousands of GPU-hours each month by eliminating the need to restart jobs after minor network disruptions. This shift toward treating infrastructure failure as an expected condition rather than an exception is becoming a core requirement for hyperscalers and cloud providers managing intensive AI deployments.
Comments (0)
No comments yet. Be the first!