apptainer-mpi-on-slurm
Run MPI applications from Apptainer containers inside Slurm allocations.
- Schedulers
- slurm
- Tools
- apptainer, sbatch, srun, mpirun
Repository release and tag provenance are verified; comparative evidence and independent promotion remain explicit gates.
All skills are shown. Filter state is reflected in the page URL.
No skills match the current filters.
| Skill | Summary | Risk | Maturity | Schedulers | Categories | Tags |
|---|---|---|---|---|---|---|
Apptainer MPI On Slurm
apptainer-mpi-on-slurm
|
Run MPI applications from Apptainer containers inside Slurm allocations. | medium | seed | slurm | containers mpi scheduler performance | apptainer singularity container mpi slurm srun |
Apptainer Run Container
apptainer-run-container
|
Run Apptainer containers safely on shared HPC systems. | medium | seed | slurm | containers scheduler | apptainer singularity container slurm gpu |
BLAS OpenMP Thread Control
blas-openmp-thread-control
|
Align BLAS, OpenMP, and language thread pools with HPC CPU allocations. | low | seed | agnostic | performance software debugging | openmp blas mkl openblas numexpr threads |
BLAST On Slurm
blast-on-slurm
|
Run local BLAST+ searches on Slurm with bounded smoke-test defaults. | medium | seed | slurm | workflow scheduler data | blast blastn blastp makeblastdb fasta bioinformatics |
Checkpoint Restart Workflow
checkpoint-restart-workflow
|
Structure long HPC jobs so they can resume after time limits or preemption. | medium | seed | slurm | scheduler debugging | checkpoint restart slurm preemption timeout state |
Checksum Manifest Create
checksum-manifest-create
|
Create checksum manifests for transfer validation and reproducibility. | low | seed | agnostic | data | checksum sha256 manifest transfer validation archive |
Cluster Usage Report Readonly
cluster-usage-report-readonly
|
Collect read-only Slurm usage evidence for facility support reports. | low | seed | slurm | admin scheduler | slurm sacct sreport squeue sinfo usage |
CMake HPC Build Preflight
cmake-hpc-build-preflight
|
Plan reproducible CMake configure, build, test, and install steps on HPC systems. | medium | seed | agnostic | software debugging mpi | cmake build compiler mpi ctest install-prefix |
Compiler MPI Matrix
compiler-mpi-matrix
|
Check compiler, MPI wrapper, and module compatibility before building HPC codes. | low | seed | agnostic | software mpi debugging | compiler mpi modules mpicc wrapper compatibility |
Conda Mamba On HPC
conda-mamba-on-hpc
|
Create Conda or Mamba environments while protecting shared HPC filesystems. | medium | seed | agnostic | software | conda mamba micromamba environment filesystem reproducibility |
Container Build For HPC
container-build-for-hpc
|
Plan and build Apptainer-compatible images for shared HPC systems. | medium | seed | agnostic | containers software | apptainer singularity container sif definition oci |
CP2K On Slurm
cp2k-on-slurm
|
Run CP2K calculations on Slurm with MPI/OpenMP layout and restart planning. | medium | seed | slurm | mpi gpu scheduler performance | cp2k dft quantum-chemistry molecular-dynamics mpi openmp |
CWL On Slurm
cwl-on-slurm
|
Run small CWL workflows inside Slurm allocations with cwltool. | medium | seed | slurm | workflow scheduler data | cwl cwltool workflow slurm bioinformatics reproducibility |
Darshan I/O Profile Analysis
darshan-io-profile-analysis
|
Analyze Darshan logs for HPC I/O behavior and bottleneck evidence. | low | seed | agnostic | performance data debugging mpi | darshan io profiling mpi-io hdf5 pnetcdf |
Dask Jobqueue On Slurm
dask-jobqueue-on-slurm
|
Launch Dask workers through Slurm with dask-jobqueue and bounded dry-run defaults. | medium | seed | slurm | software workflow scheduler | dask dask-jobqueue distributed python slurm workers |
Dataset Staging To Scratch
dataset-staging-to-scratch
|
Stage inputs to scratch, run work, and collect outputs from Slurm jobs. | medium | seed | slurm | data scheduler | scratch staging slurm rsync outputs workflow |
DeepSpeed On Slurm
deepspeed-on-slurm
|
Plan and smoke test DeepSpeed launches on Slurm GPU allocations. | medium | seed | slurm | gpu scheduler | deepspeed pytorch slurm zero checkpoint gpu |
EasyBuild Install Software
easybuild-install-software
|
Install scientific software with EasyBuild and generated modules. | medium | seed | agnostic | software | easybuild modules software easyconfig lmod |
File Descriptor Limit Triage
file-descriptor-limit-triage
|
Collect read-only evidence for too-many-open-files failures. | low | seed | agnostic | debugging data workflow | emfile ulimit nofile file-descriptor lsof dataloader |
GATK Workflow On HPC
gatk-workflow-on-hpc
|
Plan and run GATK variant-calling workflows on shared HPC systems. | medium | seed | slurm | workflow scheduler data | gatk genomics haplotypecaller java slurm variant-calling |
Globus Transfer Dataset
globus-transfer-dataset
|
Stage large datasets with Globus transfer and verification steps. | medium | seed | agnostic | data | globus transfer dataset checksum staging |
GPU Memory Triage
gpu-memory-triage
|
Distinguish GPU allocation, framework, and model memory failures. | low | seed | slurm | debugging gpu | gpu memory oom cuda rocm pytorch |
GPU Sanity Check
gpu-sanity-check
|
Verify GPU allocation, runtime visibility, and basic framework access. | medium | seed | slurm | gpu scheduler debugging | gpu cuda rocm nvidia amd slurm |
Grid Engine Submit Job
grid-engine-submit-job
|
Generate safe Grid Engine qsub scripts for common HPC job shapes. | medium | seed | grid-engine sge uge | scheduler | grid-engine sge uge qsub qstat array |
GROMACS On Slurm
gromacs-on-slurm
|
Run GROMACS jobs on Slurm with MPI, OpenMP, GPU, and checkpoint planning. | medium | seed | slurm | gpu mpi scheduler performance | gromacs molecular-dynamics mpi openmp gpu checkpoint |
HTCondor Submit Job
htcondor-submit-job
|
Generate safe HTCondor submit descriptions for common high-throughput job shapes. | medium | seed | htcondor | scheduler workflow | htcondor condor-submit condor-q high-throughput cpu gpu |
Hugging Face Accelerate On Slurm
huggingface-accelerate-on-slurm
|
Plan and smoke test Hugging Face Accelerate launches on Slurm. | medium | seed | slurm | software gpu scheduler | huggingface accelerate transformers pytorch slurm gpu |
Hybrid MPI/OpenMP On Slurm
hybrid-mpi-openmp-slurm
|
Plan and verify hybrid MPI/OpenMP task and thread layouts on Slurm. | medium | seed | slurm | mpi scheduler performance debugging | mpi openmp slurm cpus-per-task omp-num-threads cpu-bind |
Interactive Session
interactive-session
|
Start short interactive compute sessions for debugging and notebooks. | medium | seed | slurm | interactive scheduler | slurm salloc srun jupyter debugging |
IOR MDTest Storage Smoke
ior-mdtest-storage-smoke
|
Collect small IOR and MDTest storage benchmark evidence on Slurm. | medium | seed | slurm | performance data mpi scheduler | ior mdtest storage filesystem io mpi |
JAX Distributed On Slurm
jax-distributed-on-slurm
|
Plan and smoke test distributed JAX jobs on Slurm GPU allocations. | medium | seed | slurm | software gpu scheduler | jax xla python distributed slurm gpu |
Job Failure Triage
job-failure-triage
|
Diagnose common HPC job failures from scheduler and log evidence. | low | seed | slurm | debugging scheduler | slurm failure oom timeout logs triage |
Julia On Slurm
julia-on-slurm
|
Run Julia scripts on Slurm with explicit depot, project, and thread settings. | medium | seed | slurm | software scheduler | julia slurm depot packages threads batch |
Jupyter On Slurm
jupyter-on-slurm
|
Launch Jupyter notebooks inside short Slurm compute allocations. | medium | seed | slurm | interactive scheduler | jupyter notebook lab slurm tunnel interactive |
LAMMPS On Slurm
lammps-on-slurm
|
Run LAMMPS molecular dynamics jobs on Slurm with MPI, GPU, and restart planning. | medium | seed | slurm | mpi scheduler performance | lammps molecular-dynamics mpi slurm restart gpu |
Large File Archive Prepare
large-file-archive-prepare
|
Prepare large HPC datasets for archival, publication, or handoff. | medium | seed | agnostic | data | archive tar manifest publication dataset metadata |
License-Aware Slurm Job
license-aware-slurm-job
|
Plan Slurm jobs that request and verify tracked software license resources. | medium | seed | slurm | scheduler software debugging | slurm licenses sbatch scontrol lmutil licensed-software |
LSF Submit Job
lsf-submit-job
|
Generate safe IBM LSF bsub scripts for common HPC job shapes. | medium | seed | lsf | scheduler | lsf bsub bjobs cpu gpu mpi |
Lustre Striping Layout Planning
lustre-striping-layout-planning
|
Inspect and plan Lustre stripe layouts for data-intensive HPC workloads. | medium | seed | agnostic | data performance debugging | lustre striping lfs filesystem io layout |
MATLAB Batch On Slurm
matlab-batch-on-slurm
|
Run non-interactive MATLAB workloads on Slurm with explicit logs and license notes. | medium | seed | slurm | software scheduler | matlab slurm license batch logs m-files |
Module Environment Debug
module-environment-debug
|
Diagnose module, compiler, MPI, and library path conflicts. | low | seed | agnostic | software debugging | modules lmod environment compiler mpi ld_library_path |
Module Tree Health Check
module-tree-health-check
|
Collect read-only evidence about visible HPC module tree health. | low | seed | agnostic | software debugging admin | modules lmod environment-modules modulepath compiler mpi |
MPI Fabric Diagnostics
mpi-fabric-diagnostics
|
Collect MPI transport and fabric evidence for multi-node Slurm jobs. | medium | seed | slurm | mpi performance debugging scheduler | mpi slurm ucx libfabric ofi openmpi |
MPI Hello And Benchmark
mpi-hello-and-benchmark
|
Compile and run MPI sanity checks across allocated nodes. | medium | seed | slurm | mpi scheduler debugging | mpi slurm srun mpicc benchmark |
MPI Rank Binding Diagnostics
mpi-rank-binding-diagnostics
|
Collect MPI rank placement and CPU binding evidence from Slurm jobs. | medium | seed | slurm | mpi scheduler performance debugging | mpi slurm srun cpu-bind affinity rank-placement |
mpi4py On Slurm
mpi4py-on-slurm
|
Run mpi4py Python programs on Slurm with matching MPI and Python environments. | medium | seed | slurm | software mpi scheduler | python mpi4py mpi slurm srun mpiexec |
NAMD On Slurm
namd-on-slurm
|
Run NAMD molecular dynamics jobs on Slurm with CPU/GPU and restart planning. | medium | seed | slurm | mpi gpu scheduler performance | namd molecular-dynamics charm++ slurm gpu restart |
NCCL Diagnostics
nccl-diagnostics
|
Collect NCCL communication evidence for multi-GPU and multi-node jobs. | medium | seed | slurm | gpu debugging scheduler | nccl gpu distributed slurm all-reduce network |
Nextflow On Slurm
nextflow-on-slurm
|
Configure Nextflow pipelines to run through the Slurm executor. | medium | seed | slurm | workflow scheduler | nextflow slurm workflow nf-core |
nf-core On Slurm
nf-core-on-slurm
|
Run nf-core Nextflow pipelines on Slurm with conservative HPC defaults. | medium | seed | slurm | workflow scheduler data | nf-core nextflow slurm bioinformatics singularity samplesheet |
Node Health Readonly Triage
node-health-readonly-triage
|
Collect read-only Slurm node evidence for support triage. | low | seed | slurm | admin scheduler debugging | slurm sinfo scontrol squeue sacct node-health |
Node Local Scratch Staging
node-local-scratch-staging
|
Stage data through node-local scratch with guarded cleanup. | medium | seed | slurm | data scheduler debugging | slurm scratch tmpdir node-local staging cleanup |
Object Storage Transfer
object-storage-transfer
|
Plan object-storage transfers between HPC filesystems and cloud remotes. | medium | seed | agnostic | data | object-storage s3 rclone aws cloud transfer |
Open OnDemand Batch Connect
open-ondemand-batch-connect
|
Prepare reviewable Open OnDemand Batch Connect app templates. | medium | seed | slurm | interactive scheduler software education | openondemand batch-connect portal slurm interactive yaml |
OpenFOAM On Slurm
openfoam-on-slurm
|
Run OpenFOAM CFD cases on Slurm with decomposition, MPI launch, and reconstruction planning. | medium | seed | slurm | mpi scheduler performance | openfoam cfd mpi slurm decomposition reconstruction |
OpenMP Thread Affinity
openmp-thread-affinity
|
Align OpenMP threads with Slurm CPU allocations and affinity settings. | medium | seed | slurm | performance scheduler | openmp threads affinity slurm cpu-bind omp-num-threads |
Parallel HDF5 NetCDF Preflight
parallel-hdf5-netcdf-preflight
|
Check parallel HDF5 and NetCDF MPI-IO build and runtime assumptions. | medium | seed | slurm | software mpi data performance | hdf5 netcdf mpi mpi-io parallel-io smoke-test |
Parsl On Slurm
parsl-on-slurm
|
Run small Parsl workflows on Slurm with explicit provider and executor limits. | medium | seed | slurm | software workflow scheduler | parsl python workflow slurm htex slurmprovider |
PBS Submit Job
pbs-submit-job
|
Generate safe PBS or OpenPBS batch scripts for common HPC job shapes. | medium | seed | pbs openpbs pbs-pro | scheduler | pbs openpbs pbs-pro qsub qstat cpu |
Performance Profile Basic
performance-profile-basic
|
Collect first-pass performance evidence for an HPC workload. | low | seed | agnostic | performance debugging | profiling time memory io gpu perf |
Python Virtualenv On HPC
python-virtualenv-on-hpc
|
Create lightweight Python virtual environments with explicit HPC module assumptions. | low | seed | agnostic | software | python virtualenv venv pip modules reproducibility |
PyTorch DDP On Slurm
pytorch-ddp-on-slurm
|
Launch and verify PyTorch distributed data parallel jobs on Slurm. | medium | seed | slurm | gpu scheduler | pytorch ddp distributed slurm gpu nccl |
Quantum ESPRESSO On Slurm
quantum-espresso-on-slurm
|
Run Quantum ESPRESSO PWscf jobs on Slurm with MPI sizing and restart planning. | medium | seed | slurm | mpi scheduler performance | quantum-espresso qe pw.x pwscf dft mpi |
Quota And Filesystem Triage
quota-and-filesystem-triage
|
Diagnose quota, inode, and filesystem-space failures from user-visible evidence. | low | seed | agnostic | debugging data | quota inode filesystem storage enospc permission |
Ray On Slurm
ray-on-slurm
|
Launch resource-bounded Ray clusters inside Slurm allocations. | medium | seed | slurm | software gpu scheduler | ray python distributed slurm ai gpu |
Reproducible Run Capture
reproducible-run-capture
|
Capture command, environment, provenance, and logs for reproducible HPC runs. | low | seed | agnostic | debugging software | reproducibility provenance environment logs git checksum |
Rscript On Slurm
rscript-on-slurm
|
Run R scripts on Slurm with explicit package-library and output controls. | medium | seed | slurm | software scheduler | r rscript slurm cran rlibs packages |
RStudio On Slurm
rstudio-on-slurm
|
Launch policy-aware RStudio or Posit sessions from Slurm allocations. | medium | seed | slurm | interactive scheduler software | rstudio posit r slurm interactive tunnel |
Rsync Data Transfer
rsync-data-transfer
|
Transfer datasets with rsync dry-runs, resumable options, and validation hooks. | medium | seed | agnostic | data | rsync transfer checksum resume data logs |
Scientific Simulation Workflows
scientific-simulation-workflows
|
Plan, check, and triage common scientific simulation workloads on HPC systems. | medium | seed | slurm pbs openpbs pbs-pro lsf | workflow software scheduler debugging performance | simulation scientific-computing molecular-dynamics cfd weather materials |
Scratch Storage Management
scratch-storage-management
|
Inspect scratch, project, and working-directory usage before HPC jobs. | low | seed | agnostic | data debugging | scratch storage quota filesystem cleanup staging |
Shared Project Permissions Triage
shared-project-permissions-triage
|
Collect read-only evidence for shared project directory permission failures. | low | seed | agnostic | debugging data admin | permissions acl group setgid scratch project |
Slurm Array Retry Plan
slurm-array-retry-plan
|
Plan safe retries for failed Slurm array tasks. | medium | seed | slurm | scheduler debugging workflow | slurm array retry sacct sbatch failed-tasks |
Slurm Efficiency Report
slurm-efficiency-report
|
Summarize completed Slurm job efficiency from accounting data. | low | seed | slurm | scheduler performance debugging | slurm seff sacct efficiency cpu memory |
Slurm GPU Binding Diagnostics
slurm-gpu-binding-diagnostics
|
Collect per-task GPU binding and visibility evidence from Slurm jobs. | medium | seed | slurm | gpu scheduler debugging performance | slurm gpu cuda-visible-devices gpu-bind gres nvidia-smi |
Slurm Job Array Patterns
slurm-job-array-patterns
|
Run parameter sweeps and many independent tasks with Slurm job arrays. | medium | seed | slurm | scheduler | slurm array manifest parameter-sweep throughput rerun |
Slurm Job Dependency Chain
slurm-job-dependency-chain
|
Chain multi-stage Slurm jobs with explicit dependencies. | medium | seed | slurm | scheduler workflow | slurm sbatch dependency afterok afterany singleton |
Slurm Maintenance Reservation Triage
slurm-maintenance-reservation-triage
|
Collect read-only Slurm evidence for maintenance windows and reservations. | low | seed | slurm | scheduler debugging admin | slurm maintenance reservation squeue sinfo scontrol |
Slurm Monitor Job
slurm-monitor-job
|
Inspect Slurm job state, accounting records, and output paths. | low | seed | slurm | scheduler debugging | slurm squeue sacct scontrol logs |
Slurm Node Failure Triage
slurm-node-failure-triage
|
Collect read-only evidence for Slurm node-related job failures. | low | seed | slurm | scheduler debugging admin | slurm node-fail boot-fail sacct scontrol sinfo |
Slurm OOM Memory Triage
slurm-oom-memory-triage
|
Collect Slurm memory evidence for out-of-memory jobs. | low | seed | slurm | debugging scheduler performance | slurm oom memory sacct maxrss reqmem |
Slurm Output Log Triage
slurm-output-log-triage
|
Collect read-only evidence for missing or confusing Slurm output logs. | low | seed | slurm | scheduler debugging data | slurm stdout stderr logs output sacct |
Slurm Pending Reason Triage
slurm-pending-reason-triage
|
Explain why Slurm jobs are pending using read-only scheduler signals. | low | seed | slurm | scheduler debugging | slurm pending squeue scontrol sprio priority |
Slurm Preemption And Requeue
slurm-preemption-requeue
|
Handle Slurm preemption signals and guarded requeue workflows. | medium | seed | slurm | scheduler debugging | slurm preemption requeue signal checkpoint restart |
Slurm QOS Account Limit Triage
slurm-qos-account-limit-triage
|
Collect read-only evidence for Slurm account, QOS, and fairshare limits. | low | seed | slurm | scheduler debugging admin | slurm qos account association fairshare limits |
Slurm Resource Estimator
slurm-resource-estimator
|
Estimate future Slurm resource requests from accounting history. | low | seed | slurm | scheduler performance | slurm sacct resources memory walltime |
Slurm Submit Job
slurm-submit-job
|
Generate safe Slurm batch scripts for common HPC job shapes. | medium | seed | slurm | scheduler | slurm sbatch cpu gpu mpi array |
Slurm Time Limit Triage
slurm-time-limit-triage
|
Collect read-only evidence for Slurm time-limit failures. | low | seed | slurm | scheduler debugging workflow | slurm timeout walltime timelimit sacct scontrol |
Snakemake On Slurm
snakemake-on-slurm
|
Configure Snakemake workflows to submit jobs through Slurm. | medium | seed | slurm | workflow scheduler | snakemake slurm workflow profile |
Spack Environment Create
spack-environment-create
|
Create reproducible Spack environments for HPC software stacks. | medium | seed | agnostic | software | spack environment modules software reproducibility |
Streamlit On Slurm
streamlit-on-slurm
|
Run policy-aware Streamlit apps from short Slurm compute allocations. | medium | seed | slurm | interactive scheduler software | streamlit python web-app slurm interactive tunnel |
TensorBoard On Slurm
tensorboard-on-slurm
|
Run policy-aware TensorBoard monitors from short Slurm allocations. | medium | seed | slurm | interactive scheduler software debugging | tensorboard tensorflow pytorch training slurm monitoring |
TensorFlow Multiworker On Slurm
tensorflow-multiworker-on-slurm
|
Plan and smoke test TensorFlow MultiWorkerMirroredStrategy jobs on Slurm. | medium | seed | slurm | software gpu scheduler | tensorflow tfconfig multiworker keras slurm gpu |
Training Cluster Reset Checklist
training-cluster-reset-checklist
|
Prepare and review HPC training environments before and after workshops. | medium | seed | slurm | education admin scheduler | training workshop slurm preflight checklist onboarding |
VS Code Tunnel On Slurm
vscode-tunnel-on-slurm
|
Run VS Code Remote Tunnels from short Slurm compute allocations. | medium | seed | slurm | interactive scheduler software | vscode remote-tunnels ide slurm interactive code-cli |
WDL On Slurm
wdl-on-slurm
|
Run small WDL workflows inside Slurm allocations with miniwdl. | medium | seed | slurm | workflow scheduler data | wdl miniwdl workflow slurm bioinformatics reproducibility |
WRF On Slurm
wrf-on-slurm
|
Run WRF real-data jobs on Slurm with MPI sizing, real.exe staging, restart, and I/O planning. | medium | seed | slurm | mpi scheduler performance | wrf weather climate mpi slurm restart |
apptainer-mpi-on-slurm
Run MPI applications from Apptainer containers inside Slurm allocations.
apptainer-run-container
Run Apptainer containers safely on shared HPC systems.
blas-openmp-thread-control
Align BLAS, OpenMP, and language thread pools with HPC CPU allocations.
blast-on-slurm
Run local BLAST+ searches on Slurm with bounded smoke-test defaults.
checkpoint-restart-workflow
Structure long HPC jobs so they can resume after time limits or preemption.
checksum-manifest-create
Create checksum manifests for transfer validation and reproducibility.
cluster-usage-report-readonly
Collect read-only Slurm usage evidence for facility support reports.
cmake-hpc-build-preflight
Plan reproducible CMake configure, build, test, and install steps on HPC systems.
compiler-mpi-matrix
Check compiler, MPI wrapper, and module compatibility before building HPC codes.
conda-mamba-on-hpc
Create Conda or Mamba environments while protecting shared HPC filesystems.
container-build-for-hpc
Plan and build Apptainer-compatible images for shared HPC systems.
cp2k-on-slurm
Run CP2K calculations on Slurm with MPI/OpenMP layout and restart planning.
cwl-on-slurm
Run small CWL workflows inside Slurm allocations with cwltool.
darshan-io-profile-analysis
Analyze Darshan logs for HPC I/O behavior and bottleneck evidence.
dask-jobqueue-on-slurm
Launch Dask workers through Slurm with dask-jobqueue and bounded dry-run defaults.
dataset-staging-to-scratch
Stage inputs to scratch, run work, and collect outputs from Slurm jobs.
deepspeed-on-slurm
Plan and smoke test DeepSpeed launches on Slurm GPU allocations.
easybuild-install-software
Install scientific software with EasyBuild and generated modules.
file-descriptor-limit-triage
Collect read-only evidence for too-many-open-files failures.
gatk-workflow-on-hpc
Plan and run GATK variant-calling workflows on shared HPC systems.
globus-transfer-dataset
Stage large datasets with Globus transfer and verification steps.
gpu-memory-triage
Distinguish GPU allocation, framework, and model memory failures.
gpu-sanity-check
Verify GPU allocation, runtime visibility, and basic framework access.
grid-engine-submit-job
Generate safe Grid Engine qsub scripts for common HPC job shapes.
gromacs-on-slurm
Run GROMACS jobs on Slurm with MPI, OpenMP, GPU, and checkpoint planning.
htcondor-submit-job
Generate safe HTCondor submit descriptions for common high-throughput job shapes.
huggingface-accelerate-on-slurm
Plan and smoke test Hugging Face Accelerate launches on Slurm.
hybrid-mpi-openmp-slurm
Plan and verify hybrid MPI/OpenMP task and thread layouts on Slurm.
interactive-session
Start short interactive compute sessions for debugging and notebooks.
ior-mdtest-storage-smoke
Collect small IOR and MDTest storage benchmark evidence on Slurm.
jax-distributed-on-slurm
Plan and smoke test distributed JAX jobs on Slurm GPU allocations.
job-failure-triage
Diagnose common HPC job failures from scheduler and log evidence.
julia-on-slurm
Run Julia scripts on Slurm with explicit depot, project, and thread settings.
jupyter-on-slurm
Launch Jupyter notebooks inside short Slurm compute allocations.
lammps-on-slurm
Run LAMMPS molecular dynamics jobs on Slurm with MPI, GPU, and restart planning.
large-file-archive-prepare
Prepare large HPC datasets for archival, publication, or handoff.
license-aware-slurm-job
Plan Slurm jobs that request and verify tracked software license resources.
lsf-submit-job
Generate safe IBM LSF bsub scripts for common HPC job shapes.
lustre-striping-layout-planning
Inspect and plan Lustre stripe layouts for data-intensive HPC workloads.
matlab-batch-on-slurm
Run non-interactive MATLAB workloads on Slurm with explicit logs and license notes.
module-environment-debug
Diagnose module, compiler, MPI, and library path conflicts.
module-tree-health-check
Collect read-only evidence about visible HPC module tree health.
mpi-fabric-diagnostics
Collect MPI transport and fabric evidence for multi-node Slurm jobs.
mpi-hello-and-benchmark
Compile and run MPI sanity checks across allocated nodes.
mpi-rank-binding-diagnostics
Collect MPI rank placement and CPU binding evidence from Slurm jobs.
mpi4py-on-slurm
Run mpi4py Python programs on Slurm with matching MPI and Python environments.
namd-on-slurm
Run NAMD molecular dynamics jobs on Slurm with CPU/GPU and restart planning.
nccl-diagnostics
Collect NCCL communication evidence for multi-GPU and multi-node jobs.
nextflow-on-slurm
Configure Nextflow pipelines to run through the Slurm executor.
nf-core-on-slurm
Run nf-core Nextflow pipelines on Slurm with conservative HPC defaults.
node-health-readonly-triage
Collect read-only Slurm node evidence for support triage.
node-local-scratch-staging
Stage data through node-local scratch with guarded cleanup.
object-storage-transfer
Plan object-storage transfers between HPC filesystems and cloud remotes.
open-ondemand-batch-connect
Prepare reviewable Open OnDemand Batch Connect app templates.
openfoam-on-slurm
Run OpenFOAM CFD cases on Slurm with decomposition, MPI launch, and reconstruction planning.
openmp-thread-affinity
Align OpenMP threads with Slurm CPU allocations and affinity settings.
parallel-hdf5-netcdf-preflight
Check parallel HDF5 and NetCDF MPI-IO build and runtime assumptions.
parsl-on-slurm
Run small Parsl workflows on Slurm with explicit provider and executor limits.
pbs-submit-job
Generate safe PBS or OpenPBS batch scripts for common HPC job shapes.
performance-profile-basic
Collect first-pass performance evidence for an HPC workload.
python-virtualenv-on-hpc
Create lightweight Python virtual environments with explicit HPC module assumptions.
pytorch-ddp-on-slurm
Launch and verify PyTorch distributed data parallel jobs on Slurm.
quantum-espresso-on-slurm
Run Quantum ESPRESSO PWscf jobs on Slurm with MPI sizing and restart planning.
quota-and-filesystem-triage
Diagnose quota, inode, and filesystem-space failures from user-visible evidence.
ray-on-slurm
Launch resource-bounded Ray clusters inside Slurm allocations.
reproducible-run-capture
Capture command, environment, provenance, and logs for reproducible HPC runs.
rscript-on-slurm
Run R scripts on Slurm with explicit package-library and output controls.
rstudio-on-slurm
Launch policy-aware RStudio or Posit sessions from Slurm allocations.
rsync-data-transfer
Transfer datasets with rsync dry-runs, resumable options, and validation hooks.
scientific-simulation-workflows
Plan, check, and triage common scientific simulation workloads on HPC systems.
scratch-storage-management
Inspect scratch, project, and working-directory usage before HPC jobs.
shared-project-permissions-triage
Collect read-only evidence for shared project directory permission failures.
slurm-array-retry-plan
Plan safe retries for failed Slurm array tasks.
slurm-efficiency-report
Summarize completed Slurm job efficiency from accounting data.
slurm-gpu-binding-diagnostics
Collect per-task GPU binding and visibility evidence from Slurm jobs.
slurm-job-array-patterns
Run parameter sweeps and many independent tasks with Slurm job arrays.
slurm-job-dependency-chain
Chain multi-stage Slurm jobs with explicit dependencies.
slurm-maintenance-reservation-triage
Collect read-only Slurm evidence for maintenance windows and reservations.
slurm-monitor-job
Inspect Slurm job state, accounting records, and output paths.
slurm-node-failure-triage
Collect read-only evidence for Slurm node-related job failures.
slurm-oom-memory-triage
Collect Slurm memory evidence for out-of-memory jobs.
slurm-output-log-triage
Collect read-only evidence for missing or confusing Slurm output logs.
slurm-pending-reason-triage
Explain why Slurm jobs are pending using read-only scheduler signals.
slurm-preemption-requeue
Handle Slurm preemption signals and guarded requeue workflows.
slurm-qos-account-limit-triage
Collect read-only evidence for Slurm account, QOS, and fairshare limits.
slurm-resource-estimator
Estimate future Slurm resource requests from accounting history.
slurm-submit-job
Generate safe Slurm batch scripts for common HPC job shapes.
slurm-time-limit-triage
Collect read-only evidence for Slurm time-limit failures.
snakemake-on-slurm
Configure Snakemake workflows to submit jobs through Slurm.
spack-environment-create
Create reproducible Spack environments for HPC software stacks.
streamlit-on-slurm
Run policy-aware Streamlit apps from short Slurm compute allocations.
tensorboard-on-slurm
Run policy-aware TensorBoard monitors from short Slurm allocations.
tensorflow-multiworker-on-slurm
Plan and smoke test TensorFlow MultiWorkerMirroredStrategy jobs on Slurm.
training-cluster-reset-checklist
Prepare and review HPC training environments before and after workshops.
vscode-tunnel-on-slurm
Run VS Code Remote Tunnels from short Slurm compute allocations.
wdl-on-slurm
Run small WDL workflows inside Slurm allocations with miniwdl.
wrf-on-slurm
Run WRF real-data jobs on Slurm with MPI sizing, real.exe staging, restart, and I/O planning.
Adopt portable skills, adapt them to public site policy, and improve them through community evidence.
Start with the issue type that matches the smallest public-safe change.
| Collection | Summary | Status | Skills | Audience |
|---|---|---|---|---|
| ai-hpc | Skills for launching, validating, monitoring, demoing, and troubleshooting distributed AI workloads, CPU thread pools, file descriptor pressure, GPU binding, and accelerator visibility on Slurm-backed HPC systems. | draft | 20 | AI/HPC users, machine learning researchers, research software engineers, HPC support teams |
| bioinformatics-workflows | Domain skills for running reviewed nf-core, GATK, and BLAST bioinformatics workflows on Slurm-backed HPC systems. | draft | 7 | bioinformatics teams, core facilities, genomics platform engineers |
| containers | Skills for building, validating, running MPI containers, and staging data for containerized HPC workloads. | draft | 7 | research software engineers, container users, HPC support teams |
| core-hpc | Starter skills for Slurm jobs, arrays, array retry planning, dependency chains, pending reason and maintenance triage, QOS/account limit evidence, OOM, time-limit, node-failure, and output-log triage, efficiency review, file descriptor triage, license-aware jobs, restartable and requeue-safe workflows, notebooks, RStudio, IDE tunnels, OpenMP placement, debugging, node-local scratch staging, shared project permissions, and storage triage. | draft | 28 | new HPC users, research software engineers, support teams |
| data-movement | Skills for staging, transferring, validating, profiling, layout planning, benchmarking, permissions and file descriptor triage, and managing research data across HPC filesystems, node-local scratch, and object storage. | draft | 14 | data stewards, research groups, facility support teams |
| facility-ops | Read-only operational skills for usage reporting, pending reason and maintenance triage, QOS/account limit evidence, OOM memory, time-limit, and node-failure triage, file descriptor pressure, node triage, shared project permissions, and module tree health. | draft | 11 | HPC support teams, facility maintainers, research computing operators |
| gpu-mpi-performance | Skills for validating GPU allocations, GPU binding, GPU memory failures, TensorBoard training monitors, Ray, JAX, Hugging Face Accelerate, TensorFlow multi-worker training, NCCL communication, DeepSpeed and PyTorch DDP launches, MPI fabric evidence, MPI rank binding, hybrid MPI/OpenMP layouts, BLAS/OpenMP thread pools, CMake build preflight, parallel HDF5/NetCDF preflight, Darshan I/O profile analysis, Lustre striping layout planning, containerized MPI, and mpi4py launches, OpenMP placement, Slurm efficiency review, storage smoke benchmarks, and first-pass performance evidence. | draft | 27 | AI/HPC users, simulation teams, performance engineers |
| scheduler-basics | Starter skills for submitting and comparing basic jobs across Slurm, PBS-style, LSF, HTCondor, and Grid Engine schedulers, including array retry, common failure, output-log triage, memory triage, time-limit triage, and node-failure triage. | draft | 12 | new HPC users, training instructors, support teams, sites with mixed schedulers |
| simulation-workflows | Domain skills for scientific simulation planning, MPI/GPU-heavy simulation, MPI fabric evidence, hybrid MPI/OpenMP layouts, rank binding diagnostics, CMake build preflight, parallel HDF5/NetCDF preflight, Darshan I/O profile analysis, Lustre striping layout planning, electronic-structure, CFD, weather, restart, profiling, and storage-smoke workflows on HPC systems. | draft | 21 | simulation teams, computational scientists, performance engineers |
| software-stacks | Skills for debugging modules, checking module tree health, compiler/MPI compatibility, CMake build preflight, parallel HDF5/NetCDF preflight, BLAS/OpenMP thread pools, licensed software jobs, Python, TensorBoard, Streamlit, Open OnDemand templates, Ray, Dask, Parsl, JAX, Hugging Face Accelerate, TensorFlow, mpi4py, R, RStudio, Julia, MATLAB, Conda environments, IDE tunnels, containers, containerized MPI, and reproducible HPC software stacks. | draft | 30 | research software engineers, HPC support teams, tool maintainers |
| training-onboarding | Skills for teaching new HPC users, including Slurm jobs, maintenance and reservation triage, array retry planning, output-log triage, OOM memory, time-limit, and node-failure triage, file descriptor limits, BLAS/OpenMP thread pools, node-local scratch staging, shared project permissions, requeue-safe restart behavior, license-aware software use, notebooks, Open OnDemand templates, TensorBoard monitors, Streamlit apps, RStudio, IDE tunnels, Python, R, Julia, and MATLAB workloads, and workshop environments. | draft | 29 | instructors, new HPC users, training cluster maintainers |
| workflow-engines | Skills for launching portable workflow engines, CWL/WDL runs, Dask and Parsl worker pools, file descriptor and time-limit triage, lightweight Slurm dependency chains, and safe array retry planning. | draft | 11 | pipeline authors, bioinformatics teams, workflow platform maintainers |
| Adapter | Summary | Status | Scheduler | Partitions |
|---|---|---|---|---|
| example-campus-cluster | A non-production adapter showing how a site can map generic skills to local HPC policy. | example | slurm | debug, cpu, gpu |
| nersc-perlmutter-public | A public-doc-backed draft adapter for mapping portable skills to NERSC Perlmutter conventions. | draft | slurm | regular, debug, interactive, preempt, shared |