Amazon ECS adds automated GPU and instance recovery
Amazon ECS now automatically detects and repairs faulty GPUs and instances used for running containerized applications. The update means ECS takes over more of the resilience responsibilities traditionally handled by infrastructure operators, such as applying patches and remediating hardware failures without user intervention. AWS Fargate receives similar improvements. The changes aim to reduce operational overhead and outages from hardware issues in production.
- Automated repair for failing GPUs and instances in ECS
- Applies OS and driver patches without downtime
- Reduces need for manual failure detection and response
- AWS Fargate receives similar recovery capabilities
Sources covering this
In this story
More in Enterprise
California launches behavioral health data dashboard with Databricks
California has rolled out a public dashboard that centralizes behavioral health data from all 58 counties, built on Databricks.
Alteryx integrates Live Query with Google BigQuery
Alteryx has rolled out its Live Query feature for Google Cloud’s BigQuery, letting companies process and analyze unstructured data like…
Kubernetes containers drop cgroup v1 as edge AI grows
Kubernetes is phasing out support for cgroup v1, an older Linux kernel interface, as part of a shift toward handling more AI workloads…
TextExpander now offers a free plan with core features
TextExpander, the text shortcut tool, just introduced a free plan.