Click for detailed status
News of the year 2025¶
21. Oct 2025 - LEE power test¶
As part of the regular electrical inspections there will be a power test in the LEE data center on 21/22 October 2025. Due to this, the GPU nodes in Euler that are hosting the RTX 2080 Ti GPUs will be offline from 21 October 6am until 22 October 12pm (noon).
This affects jobs in the "gpu.*" queues (please note that the "gpuhe.*" and "gpupr.*" queues are not affected by this). The "gpu.*" batch queues will be progressively inactivated in the days and hours prior to the power test.
Please note that if you submit a job to the "gpu.*" (not "gpuhe.*" or "gpupr.*") queues before the power test, then the runtime that you request will determine if the job can still run before the power test or if the job will be kept in the queue until the power test is over.
We are sorry for the inconvenience
12. Dec 2025 - Power grid incident¶
Due to a power grid incident that affected the LCA data center last night around 10:00 p.m., more than 300 nodes on Euler shut down unexpectedly. We are in close contact with CSCS to investigate the cause and to bring the affected nodes back online as quickly as possible.
Jobs that were running on those nodes terminated (likely with a NODE_FAIL status). If you notice failed jobs with NODE_FAIL status, then they were likely affected by the incident. Please restart them, such that they can run on different nodes in Euler.
We are sorry for the inconvenience.
Updates will be published on this page.
Timeline:
- 2025-12-12 09:30: Half of the nodes that unexpectedly powered off are back online. We are checking the remaining nodes.
- 2025-12-12 16:00: All nodes that were affected by the incident are back online.
18. Dec 2025 - Euler VII closed¶
Due to an ongoing network incident, Euler VII (eu-a2p-* nodes) is closed and not dispatching any further jobs until the incident is resolved.
The issue has triggered node crashes and multiple jobs might have their status set to NODE_FAILURE (NF).
We apologize for the inconvenience.
Timeline:
- 19.12.2025 13:00: The network issue has been identified and resolved. All nodes have been put back into production.
19. Dec 2025 - Network incident¶
Due to a network issue, some nodes in Euler are losing their connection and causing jobs to fail.
We have stopped all queues (Slurm partitions are down) until this new issue is resolved.
We apologize for the inconvenience.
Timline:
- 20.12.2025 00:15: Shorter jobs can run over the weekend while continue to investigate the issue.
- 22.12.2025 10:00: Partitions remain closed while we are investigating the issue.
- 22.12.2025 16:00: The networking issue has been stabilized and will be further investigated in January. In the meantime the majority of Euler compute jobs can continue to run jobs over the holidays. The queues (partitions) are gradually being opened, starting with the shorter ones.
25. Dec 2025 - Power incident¶
Due to an incident with the electrical installation in one of the ETH datacenters at 21:05 on 25.12.2025, all of the nodes there (eu-lo-*) lost power and will remain unavailable until further notice. All jobs running there were lost.
This incident affects
- all regular GPU nodes (
gpu.*hpartitions, RTX 2080Ti GPUs) - some high-end (Titan RTX GPus) and
- most V100 professional GPUs.
Since all GPU jobs in the gpu.*h partitions are affected, we will allow some other GPU nodes to run these jobs. If your jobs explicitly request the 2080 Ti GPUs, then they will remain pending until the nodes are brought back online.
Timline:
- 26.12.2025 15:00: Power to all affected nodes has been restored.