<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>Jamie Mair — Blog</title>
    <link>https://jamiemair.github.io/blog</link>
    <description>Articles on machine learning, reinforcement learning and high performance computing.</description>
    <language>en-gb</language>
    <atom:link href="https://jamiemair.github.io/rss.xml" rel="self" type="application/rss+xml" />
    <lastBuildDate>Fri, 30 Sep 2022 00:00:00 GMT</lastBuildDate>
    <item>
      <title>Using Julia on the HPC</title>
      <link>https://jamiemair.github.io/blog/using-julia-on-the-hpc</link>
      <guid isPermaLink="true">https://jamiemair.github.io/blog/using-julia-on-the-hpc</guid>
      <pubDate>Fri, 30 Sep 2022 00:00:00 GMT</pubDate>
      <dc:creator>Jamie Mair</dc:creator>
      <description>Exploring the speed and ease of Julia for HPC workloads, from a single core through threads and GPUs to a whole cluster.</description>
      <category>julia</category>
      <category>hpc</category>
      <category>gpu</category>
      <category>performance</category>
      <content:encoded><![CDATA[<p><em>This post was first published as a guest post on the University of
Nottingham’s Digital Research blog, which is no longer online (an <a href="https://web.archive.org/web/20251211100513/https://blogs.nottingham.ac.uk/digitalresearch/2022/09/30/using-julia-on-the-hpc/" rel="nofollow">archived copy</a> survives). It accompanied a talk at the 2022 UoN HPC conference. I have
lightly revised it, redrawn the figures as interactive plots from the
original benchmark data, and added an explanation of where Julia’s speed
advantage over C++ comes from. The code is in <a href="https://github.com/JamieMair/julia-for-research-with-hpc" rel="nofollow">this repository</a>.</em></p> <p>In this post, I explore the speed and efficiency of the programming language
Julia on a university high-performance computing (HPC) cluster.</p> <h2 id="what-is-julia"><a aria-hidden="true" tabindex="-1" href="https://jamiemair.github.io/blog/using-julia-on-the-hpc#what-is-julia"><span class="icon icon-link"></span></a>What is Julia?</h2> <p><a href="https://julialang.org/" rel="nofollow">Julia</a> is a modern, general-purpose programming
language, designed for high-performance scientific programming. According to a <a href="https://julialang.org/blog/2012/02/why-we-created-julia/" rel="nofollow">blog post</a> written by
its creators, Julia was borne out of the desire for a language which is as fast
as C, yet as easy to use as Python, MATLAB and R. Personally, I use Julia
because it is incredibly easy to prototype new code for my research, while
being flexible enough to execute that code on my home machine, my GPU or an
entire cluster. And all of this with minimal effort on my part!</p> <p>In this post, I discuss what’s possible in Julia for very little investment.
You will see how to write reusable code and take advantage of every level of
parallelism, from the hardware parallelism of SIMD, to executing native Julia
code on the GPU and scaling it across a cluster. The ease of use and
flexibility of this parallelism provides a very low bar to entry.</p> <h2 id="julia-is-fast"><a aria-hidden="true" tabindex="-1" href="https://jamiemair.github.io/blog/using-julia-on-the-hpc#julia-is-fast"><span class="icon icon-link"></span></a>Julia is fast</h2> <p>Julia has many advantages over other languages, but the one that will interest
the HPC community most is speed. Julia claims to be as fast as C, so let’s put
that to the test.</p> <p>I will use the example of simulating a Monte Carlo process, specifically a <a href="https://en.wikipedia.org/wiki/Random_walk" rel="nofollow">random walk</a>. This process is
simple, but the basics are used in a wide variety of fields and are common in
HPC workloads. The code for running a random walk is:</p> <div class="code-block"><div class="code-header"><span class="language-label">julia</span> </div> <pre class="shiki shiki-themes light-plus dark-plus" style="background-color:#FFFFFF;--shiki-dark-bg:#1E1E1E;color:#000000;--shiki-dark:#D4D4D4" tabindex="0"><code><span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">using</span><span style="color:#000000;--shiki-dark:#D4D4D4"> Random</span></span>
<span class="line"></span>
<span class="line"><span style="color:#0000FF;--shiki-dark:#569CD6">function</span><span style="color:#795E26;--shiki-dark:#DCDCAA"> simple_monte_carlo</span><span style="color:#000000;--shiki-dark:#D4D4D4">(n, T)</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">    x = </span><span style="color:#795E26;--shiki-dark:#DCDCAA">zeros</span><span style="color:#000000;--shiki-dark:#D4D4D4">(n)</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">    for</span><span style="color:#000000;--shiki-dark:#D4D4D4"> i in </span><span style="color:#795E26;--shiki-dark:#DCDCAA">eachindex</span><span style="color:#000000;--shiki-dark:#D4D4D4">(x)</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">        for</span><span style="color:#000000;--shiki-dark:#D4D4D4"> t in </span><span style="color:#098658;--shiki-dark:#B5CEA8">1</span><span style="color:#000000;--shiki-dark:#D4D4D4">:T</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">            x[i] += Random.</span><span style="color:#795E26;--shiki-dark:#DCDCAA">randn</span><span style="color:#000000;--shiki-dark:#D4D4D4">()</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">        end</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">    end</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">    return</span><span style="color:#000000;--shiki-dark:#D4D4D4"> x</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">end</span></span></code></pre></div> <p>This function performs <span class="math math-inline"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>n</mi></mrow><annotation encoding="application/x-tex">n</annotation></semantics></math></span></span></span> random walks, all of length <span class="math math-inline"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>T</mi></mrow><annotation encoding="application/x-tex">T</annotation></semantics></math></span></span></span>. We can write some
roughly equivalent code in C++, to see the empirical differences in
performance:</p> <div class="code-block"><div class="code-header"><span class="language-label">cpp</span> </div> <pre class="shiki shiki-themes light-plus dark-plus" style="background-color:#FFFFFF;--shiki-dark-bg:#1E1E1E;color:#000000;--shiki-dark:#D4D4D4" tabindex="0"><code><span class="line"><span style="color:#0000FF;--shiki-dark:#569CD6">double</span><span style="color:#0000FF;--shiki-dark:#569CD6"> *</span><span style="color:#795E26;--shiki-dark:#DCDCAA">random_walk</span><span style="color:#000000;--shiki-dark:#D4D4D4">(</span><span style="color:#0000FF;--shiki-dark:#569CD6">size_t</span><span style="color:#001080;--shiki-dark:#9CDCFE"> n</span><span style="color:#000000;--shiki-dark:#D4D4D4">, </span><span style="color:#0000FF;--shiki-dark:#569CD6">size_t</span><span style="color:#001080;--shiki-dark:#9CDCFE"> T</span><span style="color:#000000;--shiki-dark:#D4D4D4">, </span><span style="color:#267F99;--shiki-dark:#4EC9B0">std</span><span style="color:#000000;--shiki-dark:#D4D4D4">::</span><span style="color:#267F99;--shiki-dark:#4EC9B0">normal_distribution</span><span style="color:#000000;--shiki-dark:#D4D4D4">&#x3C;</span><span style="color:#0000FF;--shiki-dark:#569CD6">double</span><span style="color:#000000;--shiki-dark:#D4D4D4">> </span><span style="color:#001080;--shiki-dark:#9CDCFE">randn</span><span style="color:#000000;--shiki-dark:#D4D4D4">, </span><span style="color:#267F99;--shiki-dark:#4EC9B0">std</span><span style="color:#000000;--shiki-dark:#D4D4D4">::</span><span style="color:#267F99;--shiki-dark:#4EC9B0">mt19937_64</span><span style="color:#001080;--shiki-dark:#9CDCFE"> rng</span><span style="color:#000000;--shiki-dark:#D4D4D4">)</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">&#123;</span></span>
<span class="line"><span style="color:#0000FF;--shiki-dark:#569CD6">    double</span><span style="color:#000000;--shiki-dark:#D4D4D4"> *x = </span><span style="color:#AF00DB;--shiki-dark:#C586C0">new</span><span style="color:#0000FF;--shiki-dark:#569CD6"> double</span><span style="color:#000000;--shiki-dark:#D4D4D4">[n];</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">    for</span><span style="color:#000000;--shiki-dark:#D4D4D4"> (</span><span style="color:#0000FF;--shiki-dark:#569CD6">size_t</span><span style="color:#000000;--shiki-dark:#D4D4D4"> i = </span><span style="color:#098658;--shiki-dark:#B5CEA8">0</span><span style="color:#000000;--shiki-dark:#D4D4D4">; i &#x3C; n; i++)</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">    &#123;</span></span>
<span class="line"><span style="color:#0000FF;--shiki-dark:#569CD6">        double</span><span style="color:#267F99;--shiki-dark:#4EC9B0"> x_t</span><span style="color:#000000;--shiki-dark:#D4D4D4"> = </span><span style="color:#098658;--shiki-dark:#B5CEA8">0.0</span><span style="color:#000000;--shiki-dark:#D4D4D4">;</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">        for</span><span style="color:#000000;--shiki-dark:#D4D4D4"> (</span><span style="color:#0000FF;--shiki-dark:#569CD6">size_t</span><span style="color:#000000;--shiki-dark:#D4D4D4"> t = </span><span style="color:#098658;--shiki-dark:#B5CEA8">0</span><span style="color:#000000;--shiki-dark:#D4D4D4">; t &#x3C; T; t++)</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">        &#123;</span></span>
<span class="line"><span style="color:#267F99;--shiki-dark:#4EC9B0">            x_t</span><span style="color:#000000;--shiki-dark:#D4D4D4"> += </span><span style="color:#795E26;--shiki-dark:#DCDCAA">randn</span><span style="color:#000000;--shiki-dark:#D4D4D4">(rng);</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">        &#125;</span></span>
<span class="line"><span style="color:#001080;--shiki-dark:#9CDCFE">        x</span><span style="color:#000000;--shiki-dark:#D4D4D4">[i] = </span><span style="color:#267F99;--shiki-dark:#4EC9B0">x_t</span><span style="color:#000000;--shiki-dark:#D4D4D4">;</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">    &#125;</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">    return</span><span style="color:#000000;--shiki-dark:#D4D4D4"> x;</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">&#125;</span></span></code></pre></div> <p>This C++ implementation is not optimal (I am no C++ programmer), but it is
indicative of what a typical PhD student might write for their simulations.
With <span class="math math-inline"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>T</mi><mo>=</mo><mn>100</mn></mrow><annotation encoding="application/x-tex">T = 100</annotation></semantics></math></span></span></span>, the Julia version is between 4.5 and 7.5 times faster,
depending on the size of <span class="math math-inline"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>n</mi></mrow><annotation encoding="application/x-tex">n</annotation></semantics></math></span></span></span>, as shown in <a href="https://jamiemair.github.io/blog/using-julia-on-the-hpc#fig:julia-hpc-cpp">Figure 1</a>.</p> <div class="code-exec-output"><figure id="fig:julia-hpc-cpp"><p><picture><source srcset="https://jamiemair.github.io/_app/immutable/assets/c0ad83c5467c3c77-122dce8d-dark.I6Rb0h0d.png" media="(prefers-color-scheme: dark)" /><img src="https://jamiemair.github.io/_app/immutable/assets/c0ad83c5467c3c77-122dce8d-light.QAusdhy3.png" alt="A static image of the interactive figure" width="768" height="450" /></picture></p><p><a href="https://jamiemair.github.io/blog/using-julia-on-the-hpc">View the interactive figure on the website.</a></p> <figcaption><strong>Figure 1:</strong> Time to simulate n random walks of length T = 100, in serial Julia and in C++. Lower is better.</figcaption></figure></div> <h3 id="why-is-julia-faster-here"><a aria-hidden="true" tabindex="-1" href="https://jamiemair.github.io/blog/using-julia-on-the-hpc#why-is-julia-faster-here"><span class="icon icon-link"></span></a>Why is Julia faster here?</h3> <p>I don’t think that Julia is inherently faster than C or C++, and the main
reason for the difference here is not the compiler but the <strong>random number
generator</strong>. Almost all of the work in a random walk is generating random
numbers, so the choice of generator dominates the run time.</p> <p>The C++ code uses <code>std::mt19937_64</code>, the 64-bit Mersenne Twister, which dates
from the late 1990s. It is a good-quality generator, but it carries a large
state of 312 64-bit words (about 2.5 KB), and every 312 numbers it stops to
regenerate the whole of that state. Since version 1.7, Julia’s default
generator has instead been <a href="https://prng.di.unimi.it/" rel="nofollow">Xoshiro256++</a>, a much more modern design. Its state
is only four 64-bit words, which fit comfortably in registers, and each number
takes only a handful of shifts, rotations, XORs and additions. Even one number
at a time, that is considerably cheaper. Better still, this simple, branch-free
arithmetic maps directly onto SIMD instructions, which operate on several
values at once, so Julia can advance several Xoshiro streams in parallel within
a single core when filling arrays with random numbers. Julia also gives every
task its own Xoshiro state, which, as we will see, lets threaded code draw
random numbers without any locking.</p> <p>A C++ developer could close most of this gap by swapping in a faster generator.
The point I want to make is that Julia makes good choices like this by default,
so you get C/C++ levels of performance for minimal effort, while keeping your
code simple and trusting Julia’s just-in-time compiler (built on LLVM, like
Clang) to produce fast machine code. Not all Julia code will run quickly, but
the <a href="https://docs.julialang.org/en/v1/manual/performance-tips/" rel="nofollow">performance tips</a> in the manual will get you to these C-like speeds.</p> <h2 id="multi-threading-in-julia"><a aria-hidden="true" tabindex="-1" href="https://jamiemair.github.io/blog/using-julia-on-the-hpc#multi-threading-in-julia"><span class="icon icon-link"></span></a>Multi-threading in Julia</h2> <p>Almost every modern CPU has multiple cores, each providing additional compute
power. Julia provides lightweight multi-threading natively. The language offers
many ways of parallelising your code, but here we’ll just use a macro to make
the process easy. A macro is code that writes other code. To execute a loop in
parallel, we just add the <code>@threads</code> macro from the standard <code>Threads</code> library:</p> <div class="code-block"><div class="code-header"><span class="language-label">julia</span> </div> <pre class="shiki shiki-themes light-plus dark-plus" style="background-color:#FFFFFF;--shiki-dark-bg:#1E1E1E;color:#000000;--shiki-dark:#D4D4D4" tabindex="0"><code><span class="line"><span style="color:#0000FF;--shiki-dark:#569CD6">function</span><span style="color:#795E26;--shiki-dark:#DCDCAA"> simple_monte_carlo_threaded</span><span style="color:#000000;--shiki-dark:#D4D4D4">(n, T)</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">    x = </span><span style="color:#795E26;--shiki-dark:#DCDCAA">zeros</span><span style="color:#000000;--shiki-dark:#D4D4D4">(n)</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">    Threads.</span><span style="color:#795E26;--shiki-dark:#DCDCAA">@threads</span><span style="color:#AF00DB;--shiki-dark:#C586C0"> for</span><span style="color:#000000;--shiki-dark:#D4D4D4"> i in </span><span style="color:#795E26;--shiki-dark:#DCDCAA">eachindex</span><span style="color:#000000;--shiki-dark:#D4D4D4">(x)</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">        for</span><span style="color:#000000;--shiki-dark:#D4D4D4"> t in </span><span style="color:#098658;--shiki-dark:#B5CEA8">1</span><span style="color:#000000;--shiki-dark:#D4D4D4">:T</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">            x[i] += Random.</span><span style="color:#795E26;--shiki-dark:#DCDCAA">randn</span><span style="color:#000000;--shiki-dark:#D4D4D4">()</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">        end</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">    end</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">    return</span><span style="color:#000000;--shiki-dark:#D4D4D4"> x</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">end</span></span></code></pre></div> <p>This is essentially the same source code as above, with a single change in
front of the <code>for</code> loop. Julia must be started with more than one thread for
this to help, for example with <code>julia --threads=8</code>. We can compare its speed-up
over the serial version in <a href="https://jamiemair.github.io/blog/using-julia-on-the-hpc#fig:julia-hpc-threaded">Figure 2</a>.</p> <div class="code-exec-output"><figure id="fig:julia-hpc-threaded"><p><picture><source srcset="https://jamiemair.github.io/_app/immutable/assets/758be84286b413e0-122dce8d-dark.DOUnyp3D.png" media="(prefers-color-scheme: dark)" /><img src="https://jamiemair.github.io/_app/immutable/assets/758be84286b413e0-122dce8d-light.tCCw4QPj.png" alt="A static image of the interactive figure" width="768" height="450" /></picture></p><p><a href="https://jamiemair.github.io/blog/using-julia-on-the-hpc">View the interactive figure on the website.</a></p> <figcaption><strong>Figure 2:</strong> Speed-up over serial Julia of the C++ version and the threaded Julia version on 8 threads. Higher is better.</figcaption></figure></div> <p>This is pretty good performance out of the box. For small <span class="math math-inline"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>n</mi></mrow><annotation encoding="application/x-tex">n</annotation></semantics></math></span></span></span>, the overhead of
starting the threads dominates and the threaded version is even slower than the
serial one, but for large <span class="math math-inline"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>n</mi></mrow><annotation encoding="application/x-tex">n</annotation></semantics></math></span></span></span> it approaches the ideal speed-up of 8 times.
There are also techniques to <a href="https://github.com/JamieMair/julia-for-research-with-hpc/blob/00681690fb4dd4a5b83c2563616658be594a916d/main.jmd#L163" rel="nofollow">reduce this overhead</a>.</p> <h2 id="moving-to-the-gpu"><a aria-hidden="true" tabindex="-1" href="https://jamiemair.github.io/blog/using-julia-on-the-hpc#moving-to-the-gpu"><span class="icon icon-link"></span></a>Moving to the GPU</h2> <p>The modern GPU excels at array-based operations at scale, such as standard BLAS
operations (matrix multiplications, and so on), which make up the backbone of a
lot of numerical and scientific computing. Writing code for a GPU is
notoriously difficult, and writing custom kernels in C presents a high barrier
to entry for beginners.</p> <p>The good news is that Julia has a successful package, called <a href="https://cuda.juliagpu.org/stable/" rel="nofollow">CUDA.jl</a>, which compiles native Julia code
directly to GPU code. Let’s take it for a spin!</p> <p>Up to now, we have only used Julia’s standard libraries, but there is a rich
ecosystem of packages available. Adding a package is as simple as running:</p> <div class="code-block"><div class="code-header"><span class="language-label">julia</span> </div> <pre class="shiki shiki-themes light-plus dark-plus" style="background-color:#FFFFFF;--shiki-dark-bg:#1E1E1E;color:#000000;--shiki-dark:#D4D4D4" tabindex="0"><code><span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">using</span><span style="color:#000000;--shiki-dark:#D4D4D4"> Pkg</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">Pkg.</span><span style="color:#795E26;--shiki-dark:#DCDCAA">add</span><span style="color:#000000;--shiki-dark:#D4D4D4">(</span><span style="color:#A31515;--shiki-dark:#CE9178">"CUDA"</span><span style="color:#000000;--shiki-dark:#D4D4D4">)</span></span></code></pre></div> <p>This uses Julia’s built-in package manager to download CUDA.jl. What makes this
painless is that the first time you use CUDA.jl, it downloads the CUDA
libraries it needs, instead of you having to install CUDA manually.</p> <p>GPUs are very different devices to CPUs, with a different programming model.
The main difference is that indexing individual elements of an array is very
slow on a GPU, so CUDA.jl raises an error when you try to do it. Instead, we can
use Julia’s excellent “broadcasting” notation, which lets us write concise,
“vectorised” code. This notation also lets us write <em>non-allocating</em> code,
which reuses memory, and that can make a huge difference to performance. Here’s
the same random walk, rewritten with array operations that avoid allocating
memory:</p> <div class="code-block"><div class="code-header"><span class="language-label">julia</span> </div> <pre class="shiki shiki-themes light-plus dark-plus" style="background-color:#FFFFFF;--shiki-dark-bg:#1E1E1E;color:#000000;--shiki-dark:#D4D4D4" tabindex="0"><code><span class="line"><span style="color:#0000FF;--shiki-dark:#569CD6">function</span><span style="color:#795E26;--shiki-dark:#DCDCAA"> random_walk!</span><span style="color:#000000;--shiki-dark:#D4D4D4">(x, T)</span></span>
<span class="line"><span style="color:#795E26;--shiki-dark:#DCDCAA">    fill!</span><span style="color:#000000;--shiki-dark:#D4D4D4">(x, </span><span style="color:#795E26;--shiki-dark:#DCDCAA">zero</span><span style="color:#000000;--shiki-dark:#D4D4D4">(</span><span style="color:#795E26;--shiki-dark:#DCDCAA">eltype</span><span style="color:#000000;--shiki-dark:#D4D4D4">(x))) </span><span style="color:#008000;--shiki-dark:#6A9955"># Set all elements of the x array to zero</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">    cache = </span><span style="color:#795E26;--shiki-dark:#DCDCAA">similar</span><span style="color:#000000;--shiki-dark:#D4D4D4">(x) </span><span style="color:#008000;--shiki-dark:#6A9955"># Create a cache array of the same size and type as x</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">    for</span><span style="color:#000000;--shiki-dark:#D4D4D4"> t in </span><span style="color:#098658;--shiki-dark:#B5CEA8">1</span><span style="color:#000000;--shiki-dark:#D4D4D4">:T</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">        Random.</span><span style="color:#795E26;--shiki-dark:#DCDCAA">randn!</span><span style="color:#000000;--shiki-dark:#D4D4D4">(cache) </span><span style="color:#008000;--shiki-dark:#6A9955"># Populate the cached memory with random numbers</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">        x .+= cache</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">    end</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">    return</span><span style="color:#000000;--shiki-dark:#D4D4D4"> x</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">end</span></span>
<span class="line"></span>
<span class="line"><span style="color:#0000FF;--shiki-dark:#569CD6">function</span><span style="color:#795E26;--shiki-dark:#DCDCAA"> simple_monte_carlo_array</span><span style="color:#000000;--shiki-dark:#D4D4D4">(n, T)</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">    x = </span><span style="color:#795E26;--shiki-dark:#DCDCAA">zeros</span><span style="color:#000000;--shiki-dark:#D4D4D4">(n)</span></span>
<span class="line"><span style="color:#795E26;--shiki-dark:#DCDCAA">    random_walk!</span><span style="color:#000000;--shiki-dark:#D4D4D4">(x, T)</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">    return</span><span style="color:#000000;--shiki-dark:#D4D4D4"> x</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">end</span></span></code></pre></div> <p><em>The <code>!</code> at the end of a function’s name is a convention which tells us that
the function mutates its arguments.</em></p> <p>This code runs on the CPU with ordinary arrays, and is very similar to how you
would write it in MATLAB or NumPy. We can add its speed-up to the comparison in <a href="https://jamiemair.github.io/blog/using-julia-on-the-hpc#fig:julia-hpc-array">Figure 3</a>.</p> <div class="code-exec-output"><figure id="fig:julia-hpc-array"><p><picture><source srcset="https://jamiemair.github.io/_app/immutable/assets/f5b8dc547607398e-122dce8d-dark.B8syGkok.png" media="(prefers-color-scheme: dark)" /><img src="https://jamiemair.github.io/_app/immutable/assets/f5b8dc547607398e-122dce8d-light.B551yzKL.png" alt="A static image of the interactive figure" width="768" height="450" /></picture></p><p><a href="https://jamiemair.github.io/blog/using-julia-on-the-hpc">View the interactive figure on the website.</a></p> <figcaption><strong>Figure 3:</strong> Speed-up over serial Julia, adding the single-threaded array version. Higher is better.</figcaption></figure></div> <p>Even on a single core, this is around 1.7 times faster than the serial version
for large <span class="math math-inline"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>n</mi></mrow><annotation encoding="application/x-tex">n</annotation></semantics></math></span></span></span>, as the whole-array operations let the compiler make use of
hardware SIMD instructions. What is really powerful about this approach is that
we can reuse this code directly to run on the GPU:</p> <div class="code-block"><div class="code-header"><span class="language-label">julia</span> </div> <pre class="shiki shiki-themes light-plus dark-plus" style="background-color:#FFFFFF;--shiki-dark-bg:#1E1E1E;color:#000000;--shiki-dark:#D4D4D4" tabindex="0"><code><span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">using</span><span style="color:#000000;--shiki-dark:#D4D4D4"> CUDA</span></span>
<span class="line"></span>
<span class="line"><span style="color:#0000FF;--shiki-dark:#569CD6">function</span><span style="color:#795E26;--shiki-dark:#DCDCAA"> simple_monte_carlo_gpu</span><span style="color:#000000;--shiki-dark:#D4D4D4">(n, T)</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">    x = CUDA.</span><span style="color:#795E26;--shiki-dark:#DCDCAA">zeros</span><span style="color:#000000;--shiki-dark:#D4D4D4">(n) </span><span style="color:#008000;--shiki-dark:#6A9955"># Creates a GPU array of zeros</span></span>
<span class="line"><span style="color:#795E26;--shiki-dark:#DCDCAA">    random_walk!</span><span style="color:#000000;--shiki-dark:#D4D4D4">(x, T) </span><span style="color:#008000;--shiki-dark:#6A9955"># Same function as before, but a different type of &#96;x&#96;</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">    return</span><span style="color:#000000;--shiki-dark:#D4D4D4"> x</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">end</span></span></code></pre></div> <p>There are no changes to our <code>random_walk!</code> function! The only thing we changed
was the type of our array. Julia’s just-in-time compiler specialises code on
the types of its arguments, so it knows to perform all of the array operations
on <code>x</code> on the GPU. The performance is shown in <a href="https://jamiemair.github.io/blog/using-julia-on-the-hpc#fig:julia-hpc-gpu">Figure 4</a>.</p> <div class="code-exec-output"><figure id="fig:julia-hpc-gpu"><p><picture><source srcset="https://jamiemair.github.io/_app/immutable/assets/0886c200f84ef13e-122dce8d-dark.CisQEz6B.png" media="(prefers-color-scheme: dark)" /><img src="https://jamiemair.github.io/_app/immutable/assets/0886c200f84ef13e-122dce8d-light.Bi_pIx5u.png" alt="A static image of the interactive figure" width="768" height="450" /></picture></p><p><a href="https://jamiemair.github.io/blog/using-julia-on-the-hpc">View the interactive figure on the website.</a></p> <figcaption><strong>Figure 4:</strong> Speed-up over serial Julia, adding the GPU version. Higher is better.</figcaption></figure></div> <p>Launching work on the GPU has a fixed cost of around half a millisecond here,
so it is only worth using for large arrays. But once we need millions of
calculations, the GPU is around 90 times faster than serial Julia on the CPU,
and still more than 10 times faster than all 8 threads. We got this huge
increase in performance by changing only the type of the input to our function,
reusing our original code.</p> <p>We have only scratched the surface of GPU programming here, as you can always
move on to writing your own kernels. For a more detailed review, see the <a href="https://cuda.juliagpu.org/stable/" rel="nofollow">CUDA.jl documentation</a>. If you would rather
not tie your code to NVIDIA hardware, <a href="https://github.com/JuliaGPU/KernelAbstractions.jl" rel="nofollow">KernelAbstractions.jl</a> lets
the same kernel run on NVIDIA, AMD, Intel and Apple GPUs, as well as the CPU.</p> <h2 id="scaling-up-to-a-cluster"><a aria-hidden="true" tabindex="-1" href="https://jamiemair.github.io/blog/using-julia-on-the-hpc#scaling-up-to-a-cluster"><span class="icon icon-link"></span></a>Scaling up to a cluster</h2> <p>This wouldn’t be an HPC post without using a cluster. While a single node with
multi-threading goes a long way, Julia also supports multiprocessing natively,
to scale across many nodes, with the <a href="https://docs.julialang.org/en/v1/stdlib/Distributed/" rel="nofollow">Distributed</a> standard
library, and supports schedulers like SLURM with <a href="https://github.com/JuliaParallel/ClusterManagers.jl" rel="nofollow">ClusterManagers.jl</a>. To
start eight workers through SLURM, all we need to run is:</p> <div class="code-block"><div class="code-header"><span class="language-label">julia</span> </div> <pre class="shiki shiki-themes light-plus dark-plus" style="background-color:#FFFFFF;--shiki-dark-bg:#1E1E1E;color:#000000;--shiki-dark:#D4D4D4" tabindex="0"><code><span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">using</span><span style="color:#000000;--shiki-dark:#D4D4D4"> Distributed</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">using</span><span style="color:#000000;--shiki-dark:#D4D4D4"> ClusterManagers</span></span>
<span class="line"></span>
<span class="line"><span style="color:#795E26;--shiki-dark:#DCDCAA">addprocs</span><span style="color:#000000;--shiki-dark:#D4D4D4">(</span><span style="color:#795E26;--shiki-dark:#DCDCAA">SlurmManager</span><span style="color:#000000;--shiki-dark:#D4D4D4">(</span><span style="color:#098658;--shiki-dark:#B5CEA8">8</span><span style="color:#000000;--shiki-dark:#D4D4D4">), exeflags=[</span><span style="color:#A31515;--shiki-dark:#CE9178">"--project"</span><span style="color:#000000;--shiki-dark:#D4D4D4">])</span></span></code></pre></div> <p>We could use the parallel primitives described in the Distributed
documentation, but let’s reuse our previous code and keep the array-based
programming model. The <a href="https://github.com/JuliaParallel/DistributedArrays.jl" rel="nofollow">DistributedArrays.jl</a> package provides the <code>DArray</code> type which, as the name suggests, splits the
memory of an array into blocks, one for each process:</p> <div class="code-block"><div class="code-header"><span class="language-label">julia</span> </div> <pre class="shiki shiki-themes light-plus dark-plus" style="background-color:#FFFFFF;--shiki-dark-bg:#1E1E1E;color:#000000;--shiki-dark:#D4D4D4" tabindex="0"><code><span class="line"><span style="color:#795E26;--shiki-dark:#DCDCAA">@everywhere</span><span style="color:#AF00DB;--shiki-dark:#C586C0"> using</span><span style="color:#000000;--shiki-dark:#D4D4D4"> DistributedArrays</span></span>
<span class="line"></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">x = </span><span style="color:#795E26;--shiki-dark:#DCDCAA">dzeros</span><span style="color:#000000;--shiki-dark:#D4D4D4">(n) </span><span style="color:#008000;--shiki-dark:#6A9955"># A DArray of zeros, split between the workers</span></span></code></pre></div> <p>We would hope to copy the approach we took for the GPU:</p> <div class="code-block"><div class="code-header"><span class="language-label">julia</span> </div> <pre class="shiki shiki-themes light-plus dark-plus" style="background-color:#FFFFFF;--shiki-dark-bg:#1E1E1E;color:#000000;--shiki-dark:#D4D4D4" tabindex="0"><code><span class="line"><span style="color:#0000FF;--shiki-dark:#569CD6">function</span><span style="color:#795E26;--shiki-dark:#DCDCAA"> simple_monte_carlo_dist</span><span style="color:#000000;--shiki-dark:#D4D4D4">(n, T)</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">    x = </span><span style="color:#795E26;--shiki-dark:#DCDCAA">dzeros</span><span style="color:#000000;--shiki-dark:#D4D4D4">(n) </span><span style="color:#008000;--shiki-dark:#6A9955"># Creates a Distributed Array (DArray) of zeros</span></span>
<span class="line"><span style="color:#795E26;--shiki-dark:#DCDCAA">    random_walk!</span><span style="color:#000000;--shiki-dark:#D4D4D4">(x, T) </span><span style="color:#008000;--shiki-dark:#6A9955"># Same function as before, but a different type of &#96;x&#96;</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">    return</span><span style="color:#000000;--shiki-dark:#D4D4D4"> x</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">end</span></span></code></pre></div> <p>Unfortunately, this does not work on its own, as <code>Random.randn!</code> has no method
for the <code>DArray</code> type. But since our function maps neatly onto each block of
the <code>DArray</code>, we can write a specific method of <code>random_walk!</code> for this type,
with a little help from the <a href="https://juliaparallel.org/DistributedArrays.jl/stable/#SPMD-Context" rel="nofollow">DistributedArrays.jl documentation</a>:</p> <div class="code-block"><div class="code-header"><span class="language-label">julia</span> </div> <pre class="shiki shiki-themes light-plus dark-plus" style="background-color:#FFFFFF;--shiki-dark-bg:#1E1E1E;color:#000000;--shiki-dark:#D4D4D4" tabindex="0"><code><span class="line"><span style="color:#795E26;--shiki-dark:#DCDCAA">@everywhere</span><span style="color:#0000FF;--shiki-dark:#569CD6"> function</span><span style="color:#795E26;--shiki-dark:#DCDCAA"> random_walk!</span><span style="color:#000000;--shiki-dark:#D4D4D4">(x::</span><span style="color:#267F99;--shiki-dark:#4EC9B0">DArray</span><span style="color:#000000;--shiki-dark:#D4D4D4">, T)</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">    if</span><span style="color:#795E26;--shiki-dark:#DCDCAA"> myid</span><span style="color:#000000;--shiki-dark:#D4D4D4">() == </span><span style="color:#098658;--shiki-dark:#B5CEA8">1</span><span style="color:#008000;--shiki-dark:#6A9955"> # The main process (not a worker)</span></span>
<span class="line"><span style="color:#000000;--shiki-dark:#D4D4D4">        SPMD.</span><span style="color:#795E26;--shiki-dark:#DCDCAA">spmd</span><span style="color:#000000;--shiki-dark:#D4D4D4">(random_walk!, x, T; pids=</span><span style="color:#795E26;--shiki-dark:#DCDCAA">workers</span><span style="color:#000000;--shiki-dark:#D4D4D4">())</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">    else</span></span>
<span class="line"><span style="color:#795E26;--shiki-dark:#DCDCAA">        random_walk!</span><span style="color:#000000;--shiki-dark:#D4D4D4">(</span><span style="color:#795E26;--shiki-dark:#DCDCAA">localpart</span><span style="color:#000000;--shiki-dark:#D4D4D4">(x), T)</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">    end</span></span>
<span class="line"><span style="color:#AF00DB;--shiki-dark:#C586C0">end</span></span></code></pre></div> <p>The <code>@everywhere</code> macro defines the function on every worker. Usually, you
would put all the code that the workers need in a separate file and load it
with a single <code>@everywhere include("shared_code.jl")</code>.</p> <p>This method tells Julia to call our original <code>random_walk!</code> on each worker’s
block of the array, reusing our code from before. The performance with 8
workers is added in <a href="https://jamiemair.github.io/blog/using-julia-on-the-hpc#fig:julia-hpc-dist">Figure 5</a>.</p> <div class="code-exec-output"><figure id="fig:julia-hpc-dist"><p><picture><source srcset="https://jamiemair.github.io/_app/immutable/assets/4a9264d023980f5e-122dce8d-dark.BWcosqIm.png" media="(prefers-color-scheme: dark)" /><img src="https://jamiemair.github.io/_app/immutable/assets/4a9264d023980f5e-122dce8d-light.hGno1yXZ.png" alt="A static image of the interactive figure" width="768" height="450" /></picture></p><p><a href="https://jamiemair.github.io/blog/using-julia-on-the-hpc">View the interactive figure on the website.</a></p> <figcaption><strong>Figure 5:</strong> Speed-up over serial Julia of every version, including the DArray version with 8 worker processes. Higher is better. Click an entry in the legend to hide or show it.</figcaption></figure></div> <p>Multiprocessing always has a much higher overhead than multi-threading, as
the processes must communicate by passing messages rather than sharing memory.
In exchange, we can scale to thousands of processes and use the memory and
resources of hundreds of different computers.</p> <h2 id="whats-next"><a aria-hidden="true" tabindex="-1" href="https://jamiemair.github.io/blog/using-julia-on-the-hpc#whats-next"><span class="icon icon-link"></span></a>What’s next?</h2> <p>Hopefully, I have shown you the range of parallelism that Julia makes
available, and convinced you that there is a low barrier to entry to get
started! All of the code in this post can be found in the <code>blog</code> folder of <a href="https://github.com/JamieMair/julia-for-research-with-hpc" rel="nofollow">this GitHub repository</a>.</p> <p>If you want to learn more about Julia, visit the <a href="https://julialang.org/" rel="nofollow">website</a>, and if you want to learn the language, start
with the <a href="https://docs.julialang.org/en/v1/manual/getting-started/" rel="nofollow">Julia manual</a>.</p> <p>I haven’t gone through many of the other reasons that people use Julia, such
as <a href="https://www.youtube.com/watch?v=kc9HwsxE1OY" rel="nofollow">multiple dispatch</a> or its
state-of-the-art support for <a href="https://docs.sciml.ai/DiffEqDocs/stable/" rel="nofollow">differential equations</a> and <a href="https://sciml.ai/" rel="nofollow">scientific machine learning</a>. If the packages you need are
only available in Python, <a href="https://github.com/JuliaPy/PythonCall.jl" rel="nofollow">PythonCall.jl</a> lets you call Python
code directly from Julia, giving you access to the entire Python ecosystem.</p> <p>Finally, if you want to ask questions, there is a very active community on the <a href="https://discourse.julialang.org/" rel="nofollow">Julia Discourse</a>.</p> <p>Happy programming!</p>]]></content:encoded>
    </item>
  </channel>
</rss>
