Kempner Computing Handbook

Kempner Computing Handbook#

Welcome to the Kempner Institute Computing Handbook, a comprehensive resource designed to empower researchers and students with the knowledge and tools necessary to leverage High-Performance Computing (HPC) for advanced computational research. This guide covers everything from the basics of getting started on the Kempner AI cluster, understanding its architecture, and navigating its environment, to more advanced topics such as job scheduling with SLURM, optimizing computational workflows, and harnessing the power of GPU computing. Through detailed sections on development and runtime environments, scalability, data management, and performance monitoring, users are equipped to efficiently manage resources, develop and run sophisticated applications, and analyze performance to ensure optimal outcomes. Whether you are new to HPC or looking to enhance your computational research projects, this guide provides the foundational knowledge and practical insights to effectively utilize the HPC resources available at the Kempner Institute.

Harvard’s Kempner AI cluster ranks 32nd on the Green500 and 85th on the TOP500 in November 2024 listing, making it one of the fastest and most energy-efficient supercomputers for advancing AI and neuroscience research.

_images/Main-Art.AI-Clustering.jpg
The Kempner Institute AI cluster’s graphics processing units are networked together to allow for incredibly fast parallel processing. Credit: Harvard University.

Read more about the Kempner AI cluster’s ranking and performance in the Harvard Gazette article: Kempner AI cluster named one of world’s fastest ‘green’ supercomputers

Explore the Handbook#

Use the sections below to jump into the handbook, or browse the full table of contents in the sidebar.

High Performance Computing

Get started on the Kempner AI cluster: SLURM scheduling, environments, storage, and data transfer.

Kempner AI Cluster
Software Engineering for Research

Version control, design principles, documentation, testing, packaging, and reproducible research.

Collaborative Code Development
AI Tools and Workflows

Reference workflows for training, distributed inference, and working with models on the cluster.

Data Discovery and Tokenization
Neuro AI Workflows

Domain workflows for neuroscience, including spike sorting on HPC.

Spike Sorting
AI Scaling and Engineering

Parallel and GPU computing, efficiency, profiling, and experiment management at scale.

Scalability
Security and Compliance

Guidance for handling data and using computing resources responsibly and securely.

Security and Compliance
Open Source Hub

Open-source tools and projects from the Kempner engineering team.

Open Source Hub
Support

Where to get help, plus answers to frequently asked questions.

Support and Troubleshooting
Workshops

Hands-on training materials from Kempner workshops and tutorials.

About Workshops @ Kempner