Chapter 03 Software and System Management
Nazmul Rahat
- •PaperRank:
- •
Chapter-3
Software and System Management OF
SUPERCOMPUTER
3.1 Software and System Management of Supercomputer:
9
3.1.1 Supercomputer Operating Systems:
Since the end of the 20th century, supercomputer operating systems have
undergone major transformations, as fundamental changes have taken place
in supercomputer architecture. While early operating systems were custom
tailored to each supercomputer to gain speed, the trend has been to move
away from in-house operating systems to the adaptation of generic software
such as Linux.
Figure A.4: The supercomputer center at NASA Ames.
Given that modern massively parallel supercomputers typically separate
computations from other services by using multiple types of nodes, they
usually run different operating systems on different nodes, e.g. using a small
and efficient lightweight kernel such as CNK or CNL on compute nodes, but a
larger system such as a Linux-derivative on server and I/O nodes.
While in a traditional multi-user computer system job scheduling is in effect a
tasking problem for processing and peripheral resources, in a massively
parallel system, the job management system needs to manage the allocation
of both computational and communication resources, as well as gracefully
dealing with inevitable hardware failures when tens of thousands of
processors are present.
Although most modern supercomputers use the Linux operating system, each
manufacturer has made its own specific changes to the Linux-derivative they
use, and no industry standard exists, partly due to the fact that the differences
in hardware architectures require changes to optimize the operating system to
each hardware design.
10
3.1.2 Software Tools and Message Passing:
The parallel architectures of supercomputers often dictate the use of special
programming techniques to exploit their speed. Software tools for distributed
processing include standard APIs such as MPI and PVM, VTL, and open
source-based software solutions such as Beowulf.
In the most common scenario, environments such as PVM and MPI for
loosely connected clusters and Open MP for tightly coordinated shared
memory machines are used. Significant effort is required to optimize an
algorithm for the interconnect characteristics of the machine it will be run on;
the aim is to prevent any of the CPUs from wasting time waiting on data from
other nodes. GPGPUs have hundreds of processor cores and are
programmed using programming models such as CUDA.
Figure A.5: Technicians working on a cluster consisting of many computers working
together by sending messages over a network.
3.2 Distributed Supercomputing:
11
3.2.1 Opportunistic Approaches:
Opportunistic Supercomputing is a form of networked grid computing whereby
a “super virtual computer” of many loosely coupled volunteer computing
machines performs very large computing tasks.
Grid computing has been applied to a number of large-scale embarrassingly
parallel problems that require supercomputing performance scales. However,
basic grid and cloud computing approaches that rely on volunteer computing
can't handle traditional supercomputing tasks such as fluid dynamic
simulations.
Figure A.6: Example architecture of a grid computing system connecting many
personal computers over the internet.
The fastest grid computing system is the distributed computing project
Folding@home. F@h reported 8.1 petaflops of x86 processing power as of
March 2012. Of this, 5.8 petaflops are contributed by clients running on
various GPUs, 1.7 petaflops come from PlayStation 3 systems, and the rest
from various CPU systems.
The BOINC platform hosts a number of distributed computing projects. As of
May 2011, BOINC recorded a processing power of over 5.5 petaflops through
over 480,000 active computers on the network[61] The most active project
(measured by computational power), MilkyWay@home, reports processing
power of over 700 teraflops through over 33,000 active computers.
As of May 2011, GIMPS's distributed Mersenne Prime search currently
achieves about 60 teraflops through over 25,000 registered computers. The
12
Internet PrimeNet Server supports GIMPS's grid computing approach, one of
the earliest and most successful grid computing projects, since 1997.
3.2.3 Quasi-Opportunistic Approaches:
Quasi-opportunistic Supercomputing is a form of distributed computing
whereby the “super virtual computer” of a large number of networked
geographically disperse computers performs huge processing power
demanding computing tasks. Quasi-opportunistic supercomputing aims to
provide a higher quality of service than opportunistic grid computing by
achieving more control over the assignment of tasks to distributed resources
and the use of intelligence about the availability and reliability of individual
systems within the supercomputing network. However, quasi-opportunistic
distributed execution of demanding parallel computing software in grids
should be achieved through implementation of grid-wise allocation
agreements, co-allocation subsystems, communication topology-aware
allocation mechanisms, fault tolerant message passing libraries and data pre-
conditioning.
Figure A.7: Representation of an atmospheric model with differential
equations that require supercomputing capabilities.
13

