Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .github/actions/spelling/allow.txt
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@ AMPI
Alpstein
Ambertools
Apertus
APUs
autoscaling
auditability
balfrin
Expand Down Expand Up @@ -52,6 +53,7 @@ Fock
Foket
GAPW
GBit
GCDs
GGA
GPFS
GPG
Expand Down Expand Up @@ -110,6 +112,7 @@ NVSHMEM
NVLINK
Nordend
netcdf
OAMs
OpenFabrics
OAuth
OIDC
Expand Down
2 changes: 1 addition & 1 deletion docs/access/firecrest.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,7 +35,7 @@ FirecREST is available for all three major [Alps platforms][ref-alps-platforms],
| [HPC Platform][ref-platform-hpcp] | https://api.cscs.ch/hpc/firecrest/v2 | [Daint][ref-cluster-daint], [Eiger][ref-cluster-eiger] |
| [ML Platform][ref-platform-mlp] | https://api.cscs.ch/ml/firecrest/v2 | [Bristen][ref-cluster-bristen], [Clariden][ref-cluster-clariden] |
| [C&W Platform][ref-platform-cwp] | https://api.cscs.ch/cw/firecrest/v2 | [Santis][ref-cluster-santis] |
| Beverin | https://api.cscs.ch/beverin/firecrest/v2 | Beverin |
| [Beverin][ref-cluster-beverin] | https://api.cscs.ch/beverin/firecrest/v2 | [Beverin][ref-cluster-beverin] |

## Accessing FirecREST

Expand Down
14 changes: 10 additions & 4 deletions docs/alps/hardware.md
Original file line number Diff line number Diff line change
Expand Up @@ -135,15 +135,21 @@ The Grizzly Peak blades contain two nodes, where each node has:
[](){#ref-alps-mi200-node}
### AMD MI250x GPU Nodes

!!! todo
Each HPE Cray EX235A (code-name Bard Peak) blade contains two nodes, where each node is made of:

Bard Peak
* 1 socket AMD Epyc 7A53 "Trento" 64-Core processor with 2 threads per core
* 512 GB DDR4 memory
* 4 AMD MI250X liquid-cooled OAMs for a total of 8 Graphics Compute Dies (GCDs) and 64 GB HBM2e memory per die
* 4 NICs -- one per OAM

[](){#ref-alps-mi300-node}
### AMD MI300A GPU Nodes

![](../images/alps/mi300-schematic.svg)

!!! todo
Each HPE Cray EX255A (code-name Parry Peak) blade contains two nodes, where each node is made of:

Parry Peak
* 4 MI300A Accelerated Processing Units (APUs)
* 24 Zen4 CPU cores per APU with 2 threads per core, for a total of 192 threads per node
* 512 GB unified physical HBM3e memory
* 4 NICs -- one per MI300A APU
121 changes: 121 additions & 0 deletions docs/clusters/beverin.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,121 @@
[](){#ref-cluster-beverin}
# Beverin

Beverin is an AMD-based GPU cluster for the purpose of testing and porting applications.

## Cluster Specification

### Login node

Beverin has one [MI250X][ref-alps-mi200-node] login node.

### Compute Nodes

Beverin consists of 13 [MI250X][ref-alps-mi200-node] and 128 [MI300A][ref-alps-mi300-node] compute nodes.

| node type | number of nodes | total CPU sockets | total GPUs |
|-----------|--------| ----------------- | ---------- |
| [mi200][ref-alps-mi200-node] | 13 | 13 | 104 |
| [mi300][ref-alps-mi300-node] | 128 | 512 | 512 |

See [Running jobs on Beverin][ref-cluster-beverin-running-jobs] for details on the associated slurm partitions.

### Storage and file systems

Beverin uses the [HPCP filesystems and storage policies][ref-hpcp-storage].

## Getting started

### Logging into Beverin

To connect to Beverin via SSH, first refer to the [ssh guide][ref-ssh].

!!! example "`~/.ssh/config`"
Add the following to your [SSH configuration][ref-ssh-config] to enable you to directly connect to Beverin using `ssh beverin`.
```
Host beverin
HostName beverin.vc.cscs.ch
ProxyJump ela
User cscsusername
IdentityFile ~/.ssh/cscs-key
IdentitiesOnly yes
```

### Software

[](){#ref-cluster-beverin-uenv}

Beverin provides uenv to deliver programming environments and application software. Please refer to the [uenv documentation][ref-uenv] for detailed information on how to use the uenv tools on the system.
Comment thread
sgozel marked this conversation as resolved.

<div class="grid cards" markdown>

- :fontawesome-solid-layer-group: __Programming Environments__

Provide compilers, MPI, Python, common libraries and tools used to build your own applications.

* [prgenv-gnu][ref-uenv-prgenv-gnu]
* [linalg][ref-uenv-linalg]

</div>

In addition to the base `GNU` and `linalg` programming environments, a few other dedicated environments are provided on a best effort basis:

<div class="grid cards" markdown>

- :fontawesome-solid-layer-group: __Climate and Weather Applications__

Provide software stacks for climate and weather workflows on Beverin.

* [ICON][ref-software-icon]

</div>

<div class="grid cards" markdown>

- :fontawesome-solid-layer-group: __Scientific Applications__

Provide scientific applications.

* [Quantumespresso][ref-uenv-quantumespresso]

</div>

<div class="grid cards" markdown>

- :fontawesome-solid-layer-group: __Tools__

Provide tools like

* [Linaro Forge][ref-uenv-linaro]
</div>

See the [uenv quick start guide][ref-uenv-quickstart] to find all programming environments available on Beverin.


[](){#ref-cluster-beverin-containers}
#### Containers

Beverin supports container workloads using the [Container Engine][ref-container-engine].

To build images, see the [guide to building container images on Alps][ref-build-containers].

[](){#ref-cluster-beverin-running-jobs}
## Running Jobs on Beverin

### Slurm

Beverin uses [Slurm][ref-slurm] as the workload manager, which is used to launch and monitor distributed workloads.

There are two Slurm partitions on the system, corresponding to the two node types:

* `mi200` (default)
* `mi300`

| name | nodes | max nodes per job | time limit |
| -- | -- | -- | -- |
| `mi200` | 13 | 13 | 24 hours |
| `mi300` | 128 | 128 | 24 hours |

### FirecREST

Beverin can also be accessed using [FirecREST][ref-firecrest] at the `https://api.cscs.ch/ml/firecrest/v2` API endpoint.
11 changes: 7 additions & 4 deletions docs/clusters/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,11 +46,14 @@ The following clusters are part of the platforms that are fully operated by CSCS
## Other systems

<div class="grid cards" markdown>
- :fontawesome-solid-mountain: __Porting and Development__
- :fontawesome-solid-mountain: __Testing, Porting and Development__

Besso is a small system used by some partners for development and porting with AMD and NVIDIA GPUs.
Beverin is a medium-sized system used for testing, developing and porting applications to AMD GPUs.

[:octicons-arrow-right-24: Beverin][ref-cluster-beverin]

Besso is a small system used by some partners for development and porting with AMD and NVIDIA GPUs.

[:octicons-arrow-right-24: Besso][ref-cluster-besso]
</div>


</div>
1 change: 1 addition & 0 deletions zensical.toml
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,7 @@ nav = [
{"Clusters" = [
"clusters/index.md",
{"Besso" = "clusters/besso.md"},
{"Beverin" = "clusters/beverin.md"},
{"Bristen" = "clusters/bristen.md"},
{"Clariden" = "clusters/clariden.md"},
{"Daint" = "clusters/daint.md"},
Expand Down
Loading