Skip to main content

BIT 495/595: Accessing High Performance Computing

Overview:

  • Practical introduction to high-performance computing for life-science and biotechnology research
  • Getting on and getting around a cluster: accounts, login nodes, file systems, and data transfer
  • Submitting, monitoring, and scaling analyses with a job scheduler
  • Running research workflows that do not fit on a laptop — genomics, imaging, simulation, and machine learning

Lecture Topics:

  • What HPC is, and when a problem actually needs it
  • Cluster architecture: nodes, cores, memory, storage tiers, and queues
  • The command line and remote access essentials
  • Job scripts and schedulers: requesting the right resources
  • Software environments: modules, containers, and reproducibility
  • Parallelism in practice: array jobs, multithreading, and GPUs
  • Moving and managing research-scale data
  • Benchmarking, troubleshooting, and good citizenship on shared systems

Hands-On Topics:

  • Connecting to the cluster and navigating the file system
  • Writing and submitting a first batch job
  • Running a bioinformatics pipeline at scale as an array job
  • Building and running a containerized analysis environment
  • GPU-accelerated analysis of a research dataset
  • Capstone: port a workflow of your own onto the cluster