<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Scientific-Computing | Bas Nijholt</title><link>https://www.nijho.lt/tag/scientific-computing/</link><atom:link href="https://www.nijho.lt/tag/scientific-computing/index.xml" rel="self" type="application/rss+xml"/><description>Scientific-Computing</description><generator>Hugo</generator><language>en-us</language><copyright>© 2026</copyright><lastBuildDate>Mon, 23 Dec 2024 00:00:00 +0000</lastBuildDate><image><url>https://www.nijho.lt/media/icon_hu_408a8f6cef43e327.png</url><title>Scientific-Computing</title><link>https://www.nijho.lt/tag/scientific-computing/</link></image><item><title>pfapack 🔢</title><link>https://www.nijho.lt/post/advent-of-open-source/23-pfapack/</link><pubDate>Mon, 23 Dec 2024 00:00:00 +0000</pubDate><guid isPermaLink="false">/post/advent-of-open-source/23-pfapack/</guid><description>&lt;p&gt;(See my &lt;a href="../"&gt;intro post&lt;/a&gt;)&lt;/p&gt;&#10;&lt;p&gt;Ever needed to calculate the square root of the determinant of a skew-symmetric matrix? Probably not, but if you&amp;rsquo;re in quantum physics, you might have! This is called the Pfaffian. Today&amp;rsquo;s project, &lt;a href="https://github.com/basnijholt/pfapack" target="_blank" rel="noopener"&gt;&lt;code&gt;pfapack&lt;/code&gt;&lt;/a&gt;, is a Python package that provides efficient routines for computing it. The underlying algorithms and original code implementations were developed by my PhD co-promotor Michael Wimmer. I created the &lt;code&gt;pfapack&lt;/code&gt; package to make this powerful tool easily accessible.&lt;/p&gt;&#10;&lt;h2 id="-origin-story"&gt;📖 Origin Story &lt;a class="anchor" href="#-origin-story" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Working in quantum physics, I frequently need to compute the Pfaffian. Michael had not only developed the algorithms but also wrote highly optimized C, FORTRAN, and Python implementations, distributing them as a tar file &lt;a href="https://michaelwimmer.org/downloads.html" target="_blank" rel="noopener"&gt;on his website&lt;/a&gt;. To make it easier for researchers working in Python, I packaged his existing Python code, added wrappers for his C and FORTRAN implementations, and set up automated builds for PyPI and conda-forge to make it easily installable.&lt;/p&gt;&#10;&lt;h2 id="-technical-highlights"&gt;🔧 Technical Highlights &lt;a class="anchor" href="#-technical-highlights" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Python Wrapper for Efficient Code:&lt;/strong&gt; User-friendly Python interface to Wimmer&amp;rsquo;s optimized C and FORTRAN.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Multiple Algorithms:&lt;/strong&gt; Supports both the Parlett-Reid and Householder methods.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Optimized for Performance:&lt;/strong&gt; Leverages low-level optimizations for speed.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Conda and Pip Installable:&lt;/strong&gt; Easy installation via &lt;code&gt;conda-forge&lt;/code&gt; and PyPI.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Automated Wheel Builds:&lt;/strong&gt; Uses &lt;a href="https://github.com/pypa/cibuildwheel" target="_blank" rel="noopener"&gt;&lt;code&gt;pypa/cibuildwheel&lt;/code&gt;&lt;/a&gt; for streamlined, automated builds and releases.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Cross-Platform Compatibility:&lt;/strong&gt; Works on Linux, macOS, and Windows (with some caveats).&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-impact"&gt;📊 Impact &lt;a class="anchor" href="#-impact" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Makes specialized mathematical algorithms readily available to Python users.&lt;/li&gt;&#10;&lt;li&gt;Simplifies the use of highly optimized code.&lt;/li&gt;&#10;&lt;li&gt;15 GitHub stars.&lt;/li&gt;&#10;&lt;li&gt;Used in research projects that require efficient Pfaffian computations, particularly in quantum physics.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-challenges-and-solutions"&gt;🎯 Challenges and Solutions &lt;a class="anchor" href="#-challenges-and-solutions" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Bridging Languages:&lt;/strong&gt; Wrapping C and FORTRAN for Python integration.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Automated Builds:&lt;/strong&gt; Setting up automated builds for multiple platforms was complex.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Windows Compatibility:&lt;/strong&gt; As always Windows support is a PITA; it currently requires MSYS2 and MinGW64.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-lessons-learned"&gt;💡 Lessons Learned &lt;a class="anchor" href="#-lessons-learned" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Making specialized tools accessible broadens their impact.&lt;/li&gt;&#10;&lt;li&gt;Wrapping low-level code in Python can improve usability.&lt;/li&gt;&#10;&lt;li&gt;Automated builds are essential for maintainability.&lt;/li&gt;&#10;&lt;li&gt;&lt;a href="https://github.com/pypa/cibuildwheel" target="_blank" rel="noopener"&gt;&lt;code&gt;pypa/cibuildwheel&lt;/code&gt;&lt;/a&gt; is awesome for cross-platform builds.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-credits"&gt;🙏 Credits &lt;a class="anchor" href="#-credits" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;All the credits for the underlying algorithms and original code implementations go to Michael Wimmer. I simply made it easier to distribute and use in Python.&lt;/p&gt;&#10;&lt;p&gt;Want to compute Pfaffians efficiently in Python? Check out &lt;a href="https://github.com/basnijholt/pfapack" target="_blank" rel="noopener"&gt;&lt;code&gt;pfapack&lt;/code&gt; on GitHub&lt;/a&gt; or install it via &lt;code&gt;pip install pfapack&lt;/code&gt; or &lt;code&gt;conda install -c conda-forge pfapack&lt;/code&gt;!&lt;/p&gt;&#10;&lt;p&gt;#OpenSource #Python #ScientificComputing #Math #Pfaffian #FORTRAN #C&lt;/p&gt;&#10;</description></item><item><title>pipefunc 🕸️</title><link>https://www.nijho.lt/post/advent-of-open-source/20-pipefunc/</link><pubDate>Fri, 20 Dec 2024 00:00:00 +0000</pubDate><guid isPermaLink="false">/post/advent-of-open-source/20-pipefunc/</guid><description>&lt;p&gt;(See my &lt;a href="../"&gt;intro post&lt;/a&gt;)&lt;/p&gt;&#10;&lt;p&gt;Of all the projects I&amp;rsquo;m sharing this month, this one that started as a passion project excites me the most!&lt;/p&gt;&#10;&lt;h2 id="-origin-story"&gt;📖 Origin Story &lt;a class="anchor" href="#-origin-story" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Sometimes the best projects are born from the most unexpected moments. More than 1½ year ago, while on parental leave with my twin boys, I found myself with small pockets of time during naps to work on a problem that had been bothering me for years: the tedious bookkeeping required in complex computational workflows. Every scientific computation project seemed to reinvent the same patterns - managing function dependencies, parameter sweeps, result caching, parallelization, and a lot of boilerplate to combine the resulting data. I wanted something that would let scientists focus on their science, not on pipeline management.&lt;/p&gt;&#10;&lt;figure id="figure-example-of-a-pipeline-generated-by-pipefunc"&gt;&lt;img src="https://www.nijho.lt/post/advent-of-open-source/20-pipefunc/image_2_hu_91068e582234f921.webp" srcset="https://www.nijho.lt/post/advent-of-open-source/20-pipefunc/image_2_hu_75ce53994a0e90e4.webp 400w, https://www.nijho.lt/post/advent-of-open-source/20-pipefunc/image_2_hu_91068e582234f921.webp 760w, https://www.nijho.lt/post/advent-of-open-source/20-pipefunc/image_2_hu_220beece6a403360.webp 1200w" sizes="auto, (max-width: 760px) 100vw, 720px" width="760" height="383" alt="A screenshot of a code editor showing Python code that uses the &amp;#39;pipefunc&amp;#39; library to define and visualize a computational pipeline. On the right, a graph generated by &amp;#39;pipeline.visualize()&amp;#39; is displayed, showing nodes and connections representing the pipeline&amp;#39;s structure." loading="lazy" data-zoomable&gt;&lt;figcaption&gt;Example of a pipeline generated by PipeFunc&lt;/figcaption&gt;&lt;/figure&gt;&#10;&lt;h2 id="-technical-highlights"&gt;🔧 Technical Highlights &lt;a class="anchor" href="#-technical-highlights" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Automatic DAG construction via very simple and lightweight syntax&lt;/li&gt;&#10;&lt;li&gt;N-dimensional parameter sweeps with automatic parallelization&lt;/li&gt;&#10;&lt;li&gt;Visual pipeline representation using NetworkX&lt;/li&gt;&#10;&lt;li&gt;Resource profiling (CPU, memory, time)&lt;/li&gt;&#10;&lt;li&gt;Type validation between pipeline stages&lt;/li&gt;&#10;&lt;li&gt;Ultra-fast: only 15 µs overhead per function&lt;/li&gt;&#10;&lt;li&gt;Flexible caching strategies (memory, disk, cloud)&lt;/li&gt;&#10;&lt;li&gt;Integration with scientific computing tools:&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;a href="https://networkx.org/" target="_blank" rel="noopener"&gt;NetworkX&lt;/a&gt; for graph algorithms&lt;/li&gt;&#10;&lt;li&gt;&lt;a href="https://numpy.org/" target="_blank" rel="noopener"&gt;NumPy&lt;/a&gt; for multi-dimensional arrays&lt;/li&gt;&#10;&lt;li&gt;&lt;a href="https://docs.xarray.dev/" target="_blank" rel="noopener"&gt;Xarray&lt;/a&gt; for labeled arrays&lt;/li&gt;&#10;&lt;li&gt;&lt;a href="https://zarr.dev/" target="_blank" rel="noopener"&gt;Zarr&lt;/a&gt; for distributed storage&lt;/li&gt;&#10;&lt;li&gt;&lt;a href="https://adaptive.readthedocs.io/" target="_blank" rel="noopener"&gt;Adaptive&lt;/a&gt; for parallel computing&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-impact"&gt;📊 Impact &lt;a class="anchor" href="#-impact" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;230 GitHub stars&lt;/li&gt;&#10;&lt;li&gt;700+ tests with 100% coverage&lt;/li&gt;&#10;&lt;li&gt;Fully typed codebase&lt;/li&gt;&#10;&lt;li&gt;Comprehensive documentation&lt;/li&gt;&#10;&lt;li&gt;Tested on real workflow on SLURM cluster&lt;/li&gt;&#10;&lt;li&gt;Covered on &lt;a href="https://pycoders.com/issues/650" target="_blank" rel="noopener"&gt;Pycoder&amp;rsquo;s Weekly&lt;/a&gt; with &amp;gt;100k subscribers&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-challenges-and-solutions"&gt;🎯 Challenges and Solutions &lt;a class="anchor" href="#-challenges-and-solutions" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Balancing simplicity with power&lt;/li&gt;&#10;&lt;li&gt;Making complex workflows intuitive&lt;/li&gt;&#10;&lt;li&gt;Handling distributed computing edge cases&lt;/li&gt;&#10;&lt;li&gt;Ensuring type safety across the pipeline&lt;/li&gt;&#10;&lt;li&gt;Optimizing performance without sacrificing features&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-lessons-learned"&gt;💡 Lessons Learned &lt;a class="anchor" href="#-lessons-learned" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;Sometimes the best time to code is during baby naps&lt;/li&gt;&#10;&lt;li&gt;Complex problems can have elegant solutions&lt;/li&gt;&#10;&lt;li&gt;Scientific computing needs better tooling&lt;/li&gt;&#10;&lt;li&gt;Good abstractions make hard things easy&lt;/li&gt;&#10;&lt;li&gt;Type hints and tests prevent headaches&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;h2 id="-future-plans"&gt;🔮 Future Plans &lt;a class="anchor" href="#-future-plans" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;The journey is far from over. Plans include:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Enhanced cloud computing support&lt;/li&gt;&#10;&lt;li&gt;More interactive visualization options&lt;/li&gt;&#10;&lt;li&gt;Interactive pipeline debugging tools&lt;/li&gt;&#10;&lt;li&gt;Expanded parameter sweep capabilities&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Want to simplify your computational workflows? Check out &lt;a href="https://github.com/pipefunc/pipefunc" target="_blank" rel="noopener"&gt;pipefunc on GitHub&lt;/a&gt; or read the &lt;a href="https://pipefunc.readthedocs.io/" target="_blank" rel="noopener"&gt;documentation&lt;/a&gt;!&lt;/p&gt;&#10;&lt;p&gt;#OpenSource #Python #DataScience #ScientificComputing #Programming&lt;/p&gt;&#10;</description></item><item><title>Adaptive Scheduler 🚀</title><link>https://www.nijho.lt/post/advent-of-open-source/17-adaptive-scheduler/</link><pubDate>Tue, 17 Dec 2024 00:00:00 +0000</pubDate><guid isPermaLink="false">/post/advent-of-open-source/17-adaptive-scheduler/</guid><description>&lt;p&gt;(See my &lt;a href="../"&gt;intro post&lt;/a&gt;)&lt;/p&gt;&#10;&lt;p&gt;Ever tried to efficiently run 100,000 cores from a Jupyter notebook? No? Well, I have, and it led to some interesting discoveries about the limits of parallel computing tools.&lt;/p&gt;&#10;&lt;h2 id="-origin-story"&gt;📖 Origin Story &lt;a class="anchor" href="#-origin-story" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;During my 2018 internship at Microsoft Quantum, we faced a unique challenge: our &lt;a href="https://github.com/python-adaptive/adaptive" target="_blank" rel="noopener"&gt;adaptive sampling algorithms&lt;/a&gt; needed to run on massive computing clusters (50,000+ cores), but existing tools couldn&amp;rsquo;t handle it efficiently.&lt;/p&gt;&#10;&lt;p&gt;The problem was interesting: traditional parallel computing tools like Dask rely on a central scheduler process. Imagine 100,000 cores, each finishing a task every 10 seconds. That means every 10 microseconds, a core needs new work! With typical scheduler overhead being 1-50 milliseconds, workers would spend 99% of their time waiting rather than computing.&lt;/p&gt;&#10;&lt;p&gt;Usually, you&amp;rsquo;d solve this by batching work ahead of time. But with adaptive sampling (&lt;a href="../16-adaptive"&gt;yesterday&amp;rsquo;s project&lt;/a&gt;), that&amp;rsquo;s impossible - you need the results of previous calculations to know what to compute next. It was like trying to conduct an orchestra where each musician needs personal instructions every few seconds!&lt;/p&gt;&#10;&lt;h2 id="-technical-highlights"&gt;🔧 Technical Highlights &lt;a class="anchor" href="#-technical-highlights" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;We created a solution that:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Acts as a meta-scheduler, launching multiple sub-schedulers (&lt;a href="https://www.dask.org/" target="_blank" rel="noopener"&gt;Dask&lt;/a&gt;, &lt;a href="https://ipyparallel.readthedocs.io/en/latest/" target="_blank" rel="noopener"&gt;ipyparallel&lt;/a&gt;, &lt;a href="https://mpi4py.readthedocs.io/en/stable/" target="_blank" rel="noopener"&gt;mpi4py&lt;/a&gt;, etc.)&lt;/li&gt;&#10;&lt;li&gt;Minimizes communication between scheduling layers&lt;/li&gt;&#10;&lt;li&gt;Handles job management through a simple database&lt;/li&gt;&#10;&lt;li&gt;Supports automatic fault tolerance and data saving&lt;/li&gt;&#10;&lt;li&gt;Works with any cluster scheduler (SLURM, PBS)&lt;/li&gt;&#10;&lt;li&gt;Provides real-time progress tracking in Jupyter notebooks&lt;/li&gt;&#10;&lt;li&gt;Scales to 50,000+ cores without breaking a sweat&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-impact"&gt;📊 Impact &lt;a class="anchor" href="#-impact" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Used extensively for real quantum device simulations&lt;/li&gt;&#10;&lt;li&gt;26 GitHub stars (but don&amp;rsquo;t let that fool you - it&amp;rsquo;s highly specialized!)&lt;/li&gt;&#10;&lt;li&gt;Enables interactive supercomputing from a simple Jupyter notebook&lt;/li&gt;&#10;&lt;li&gt;Likely ran for many millions of core-hours across various projects&lt;/li&gt;&#10;&lt;li&gt;Actively maintained with 61 releases since 2018, proving its ongoing value&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-challenges-we-solved"&gt;🎯 Challenges We Solved &lt;a class="anchor" href="#-challenges-we-solved" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Scheduler overhead killing performance at scale&lt;/li&gt;&#10;&lt;li&gt;Data locality and fault tolerance&lt;/li&gt;&#10;&lt;li&gt;Real-time monitoring of massive parallel jobs&lt;/li&gt;&#10;&lt;li&gt;Making supercomputing accessible through notebooks&lt;/li&gt;&#10;&lt;li&gt;Automatic job recovery after crashes (e.g., when using spot VMs or Heisenbugs)&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-lessons-learned"&gt;💡 Lessons Learned &lt;a class="anchor" href="#-lessons-learned" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;Sometimes you need to think outside the box&lt;/li&gt;&#10;&lt;li&gt;The best tools hide complexity behind simple interfaces&lt;/li&gt;&#10;&lt;li&gt;Real-time feedback is crucial, even for massive computations&lt;/li&gt;&#10;&lt;li&gt;When scaling up 1000x, you often need a completely different approach&lt;/li&gt;&#10;&lt;li&gt;Making things &amp;ldquo;just work&amp;rdquo; from a notebook is worth the effort&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;Want to run massive parallel computations efficiently? Check out &lt;a href="https://github.com/basnijholt/adaptive-scheduler" target="_blank" rel="noopener"&gt;Adaptive Scheduler on GitHub&lt;/a&gt;!&lt;/p&gt;&#10;&lt;p&gt;#OpenSource #Python #HPC #QuantumComputing #ParallelComputing&lt;/p&gt;&#10;</description></item><item><title>Adaptive 📈</title><link>https://www.nijho.lt/post/advent-of-open-source/16-adaptive/</link><pubDate>Mon, 16 Dec 2024 00:00:00 +0000</pubDate><guid isPermaLink="false">/post/advent-of-open-source/16-adaptive/</guid><description>&lt;p&gt;(See my &lt;a href="../"&gt;intro post&lt;/a&gt;)&lt;/p&gt;&#10;&lt;p&gt;During my PhD, I was running quantum transport simulations on our 800-core computational cluster.&#10;Traditional parameter sweeps using dense grids were hitting the cluster&amp;rsquo;s limits - wasting precious compute time on uninteresting regions while potentially missing important features.&#10;Most parameter spaces we explored were &amp;ldquo;boring&amp;rdquo; with occasional interesting regions that required higher sampling density.&lt;/p&gt;&#10;&lt;h2 id="-origin-story"&gt;📖 Origin Story &lt;a class="anchor" href="#-origin-story" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Initially, in 2016, I had a collection of one-off adaptive sampling scripts.&#10;Together with Joseph Weston and the support of my PhD advisor Anton Akhmerov, we turned it into a full-fledged Python package &amp;ldquo;&lt;a href="https://github.com/python-adaptive/adaptive" target="_blank" rel="noopener"&gt;Adaptive&lt;/a&gt;&amp;rdquo;.&#10;What began as an internal tool to make our cluster usage more efficient slowly gained traction in the wider scientific computing community.&lt;/p&gt;&#10;&lt;h2 id="-technical-highlights"&gt;🔧 Technical Highlights &lt;a class="anchor" href="#-technical-highlights" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Multiple adaptive sampling strategies (1D, 2D, ND, averaging stochastic functions)&lt;/li&gt;&#10;&lt;li&gt;Parallel execution support (local, MPI clusters, distributed systems)&lt;/li&gt;&#10;&lt;li&gt;Live plotting in Jupyter notebooks with real-time feedback&lt;/li&gt;&#10;&lt;li&gt;Customizable loss functions for different sampling priorities&lt;/li&gt;&#10;&lt;li&gt;Minimal dependencies&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-impact--applications"&gt;📊 Impact &amp;amp; Applications &lt;a class="anchor" href="#-impact--applications" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;1164 GitHub stars and thousands of daily downloads&lt;/li&gt;&#10;&lt;li&gt;Cited in 20+ scientific publications&lt;/li&gt;&#10;&lt;li&gt;Dependency of many packages, most notably: Orange Quantum System&amp;rsquo;s Quantify software for controlling real quantum computers!&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Used in cutting-edge research across fields:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Quantum Computing&lt;/strong&gt;: Simulating and controlling superconducting qubits and trapped ion systems&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Quantum Physics&lt;/strong&gt;: Exploring topological quantum systems and novel materials&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Materials Science&lt;/strong&gt;: Studying graphene and other 2D materials&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Engineering&lt;/strong&gt;:&#10;&lt;ul&gt;&#10;&lt;li&gt;Optimizing radar system performance&lt;/li&gt;&#10;&lt;li&gt;Fluid dynamics simulations&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Chemistry&lt;/strong&gt;: Molecular modeling and vibrational spectroscopy&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Most Unexpected Use&lt;/strong&gt;: Found in a paper on predicting energy usage of wartime bungalows!&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;This range of applications shows how a tool built for quantum transport simulations can find surprising uses across disciplines, from fundamental physics to everyday buildings.&lt;/p&gt;&#10;&lt;h2 id="-challenges-and-solutions"&gt;🎯 Challenges and Solutions &lt;a class="anchor" href="#-challenges-and-solutions" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Making parallel execution truly scalable while maintaining sampling intelligence&lt;/li&gt;&#10;&lt;li&gt;Creating intuitive visualization tools for complex parameter spaces&lt;/li&gt;&#10;&lt;li&gt;Handling stochastic functions and noisy data&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-lessons-learned"&gt;💡 Lessons Learned &lt;a class="anchor" href="#-lessons-learned" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;Building a user base takes active effort - giving talks and engaging online made all the difference&lt;/li&gt;&#10;&lt;li&gt;Good documentation converts interested users into regular contributors&lt;/li&gt;&#10;&lt;li&gt;This was my first open source package and learned all about packaging, testing, and CI/CD&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;Want smarter parameter space exploration? Check out &lt;a href="https://github.com/python-adaptive/adaptive" target="_blank" rel="noopener"&gt;Adaptive on GitHub&lt;/a&gt; or try our &lt;a href="https://adaptive.readthedocs.io/en/latest/tutorial/tutorial.html" target="_blank" rel="noopener"&gt;interactive tutorial&lt;/a&gt;!&lt;/p&gt;&#10;&lt;p&gt;#OpenSource #Python #ScientificComputing #ParallelComputing #Programming&lt;/p&gt;&#10;</description></item><item><title>conda-forge Contributions 🐍</title><link>https://www.nijho.lt/post/advent-of-open-source/11-conda-forge-contributions/</link><pubDate>Wed, 11 Dec 2024 00:00:00 +0000</pubDate><guid isPermaLink="false">/post/advent-of-open-source/11-conda-forge-contributions/</guid><description>&lt;p&gt;(See my &lt;a href="../"&gt;intro post&lt;/a&gt;)&lt;/p&gt;&#10;&lt;p&gt;Unlike other projects in this advent calendar, I want to celebrate being part of something much bigger: the &lt;a href="https://conda-forge.org/" target="_blank" rel="noopener"&gt;conda-forge&lt;/a&gt; community, where thousands of volunteers maintain the packages that power scientific Python.&lt;/p&gt;&#10;&lt;h2 id="-origin-story"&gt;📖 Origin Story &lt;a class="anchor" href="#-origin-story" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;I started with conda-forge in 2016 during my PhD work with Kwant, a quantum transport simulator. Installing it was a nightmare - it required system-level numerical libraries like MUMPS and Scotch which required a special config file to link. When I discovered conda-forge and its mission to make scientific software installation painless, I knew this was the solution!&lt;/p&gt;&#10;&lt;figure id="figure-the-humble-work-of-conda-forge-maintainers"&gt;&lt;img src="https://www.nijho.lt/post/advent-of-open-source/11-conda-forge-contributions/meme_hu_227391ecf0bd252.webp" srcset="https://www.nijho.lt/post/advent-of-open-source/11-conda-forge-contributions/meme_hu_9b170f53ebe4dc1b.webp 400w, https://www.nijho.lt/post/advent-of-open-source/11-conda-forge-contributions/meme_hu_227391ecf0bd252.webp 760w, https://www.nijho.lt/post/advent-of-open-source/11-conda-forge-contributions/meme_hu_95b0c28d6df1a41a.webp 978w" sizes="auto, (max-width: 760px) 100vw, 720px" width="760" height="586" alt="Meme showing a farmer standing in a field, with the text &amp;#39;MAINTAINING CONDA-FORGE PACKAGES FOR THE SCIENTIFIC COMMUNITY&amp;#39; above and &amp;#39;It ain&amp;#39;t much, but it&amp;#39;s honest work&amp;#39; below." loading="lazy" data-zoomable&gt;&lt;figcaption&gt;The humble work of conda-forge maintainers&lt;/figcaption&gt;&lt;/figure&gt;&#10;&lt;h2 id="-technical-highlights"&gt;🔧 Technical Highlights &lt;a class="anchor" href="#-technical-highlights" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Over the past 8 years, I&amp;rsquo;ve maintained over 40 build recipes, including:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Major scientific packages like &lt;a href="https://vtk.org/" target="_blank" rel="noopener"&gt;VTK&lt;/a&gt;, &lt;a href="https://mumps-solver.org/" target="_blank" rel="noopener"&gt;MUMPS&lt;/a&gt;, and &lt;a href="https://www.labri.fr/perso/pelegrin/scotch/" target="_blank" rel="noopener"&gt;Scotch&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;Quantum physics tools like &lt;a href="https://quantumai.google/qsim" target="_blank" rel="noopener"&gt;Qsim&lt;/a&gt;, &lt;a href="https://github.com/microsoft/Qcodes" target="_blank" rel="noopener"&gt;Qcodes&lt;/a&gt;, and &lt;a href="https://kwant-project.org/" target="_blank" rel="noopener"&gt;Kwant&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;Visualization libraries like &lt;a href="https://holoviews.org/" target="_blank" rel="noopener"&gt;HoloViews&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;Plus all the packages I&amp;rsquo;ve authored myself&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The real magic of conda-forge is its infrastructure:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Automated builds across Linux, macOS, and Windows&lt;/li&gt;&#10;&lt;li&gt;Strict dependency version management&lt;/li&gt;&#10;&lt;li&gt;Comprehensive CI/CD pipelines&lt;/li&gt;&#10;&lt;li&gt;Community-driven quality control&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-impact"&gt;📊 Impact &lt;a class="anchor" href="#-impact" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;What makes conda-forge truly special is how initial contributions blossom into collaborative efforts. I&amp;rsquo;ve experienced this firsthand: after creating the initial MUMPS feedstock, 25 other contributors joined in, each bringing their unique expertise.&lt;/p&gt;&#10;&lt;p&gt;A recent experience also illustrates this: I added CUDA support to both &lt;a href="https://github.com/conda-forge/qsimcirq-feedstock/pull/3" target="_blank" rel="noopener"&gt;Google&amp;rsquo;s Qsim&lt;/a&gt; and &lt;a href="https://github.com/conda-forge/qiskit-aer-feedstock/pull/19" target="_blank" rel="noopener"&gt;IBM&amp;rsquo;s Cirq&lt;/a&gt; packages. Shortly after, &lt;a href="https://github.com/leofang" target="_blank" rel="noopener"&gt;Leo Fang&lt;/a&gt;, an engineer from NVIDIA who knows CUDA far better than I do, stepped up and made significant improvements. This is the beauty of open source - experts naturally gravitate to where they can make the biggest impact.&lt;/p&gt;&#10;&lt;p&gt;This collaborative approach solves a problem I&amp;rsquo;ve encountered at several companies: the tendency to maintain complicated build systems internally with custom hacks. Instead of each organization reinventing the wheel, conda-forge provides a shared platform where:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Maintenance burden is distributed&lt;/li&gt;&#10;&lt;li&gt;Expert knowledge flows freely&lt;/li&gt;&#10;&lt;li&gt;Build practices are standardized&lt;/li&gt;&#10;&lt;li&gt;Infrastructure is public and reusable&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-challenges-and-solutions"&gt;🎯 Challenges and Solutions &lt;a class="anchor" href="#-challenges-and-solutions" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Keeping up with upstream changes&lt;/li&gt;&#10;&lt;li&gt;Managing complex dependency trees&lt;/li&gt;&#10;&lt;li&gt;Cross-platform compatibility&lt;/li&gt;&#10;&lt;li&gt;Coordinating with upstream maintainers&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-lessons-learned"&gt;💡 Lessons Learned &lt;a class="anchor" href="#-lessons-learned" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;Community effort beats solo work for infrastructure&lt;/li&gt;&#10;&lt;li&gt;Automation is crucial for reliability&lt;/li&gt;&#10;&lt;li&gt;Small contributions compound over time&lt;/li&gt;&#10;&lt;li&gt;Being part of something bigger is rewarding&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;Want to contribute to scientific Python? Check out &lt;a href="https://conda-forge.org/" target="_blank" rel="noopener"&gt;conda-forge&lt;/a&gt; and the &lt;a href="https://github.com/conda-forge/staged-recipes" target="_blank" rel="noopener"&gt;staged-recipes repository&lt;/a&gt;!&lt;/p&gt;&#10;&lt;p&gt;#OpenSource #Python #Scientific #Programming #Community&lt;/p&gt;&#10;</description></item><item><title>numthreads 🧵</title><link>https://www.nijho.lt/post/advent-of-open-source/07-numthreads/</link><pubDate>Sat, 07 Dec 2024 00:00:00 +0000</pubDate><guid isPermaLink="false">/post/advent-of-open-source/07-numthreads/</guid><description>&lt;p&gt;(See my &lt;a href="../"&gt;intro post&lt;/a&gt;)&lt;/p&gt;&#10;&lt;p&gt;Sometimes the smallest tools solve the most persistent problems.&#10;Today&amp;rsquo;s project is about taming automatic parallelization in scientific computing - a deceptively simple challenge that has cost countless CPU hours.&lt;/p&gt;&#10;&lt;h2 id="-origin-story"&gt;📖 Origin Story &lt;a class="anchor" href="#-origin-story" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;While working on high-performance computing clusters, I frequently encountered a counterintuitive problem: code running slower on multiple CPU cores than on a single one.&#10;Many scientific libraries (NumPy, SciPy) automatically parallelize operations, which sounds great but can actually harm performance when you&amp;rsquo;re already parallelizing at a higher level.&#10;After repeatedly explaining to colleagues why they needed to set various environment variables to disable this behavior, I created &lt;a href="https://github.com/basnijholt/numthreads" target="_blank" rel="noopener"&gt;&lt;code&gt;numthreads&lt;/code&gt;&lt;/a&gt; to solve it once and for all.&lt;/p&gt;&#10;&lt;figure id="figure-overenthusiastic-threading"&gt;&lt;img src="https://www.nijho.lt/post/advent-of-open-source/07-numthreads/meme_hu_4d36ab41a66a48cb.webp" srcset="https://www.nijho.lt/post/advent-of-open-source/07-numthreads/meme_hu_8178332437a0cb36.webp 297w, https://www.nijho.lt/post/advent-of-open-source/07-numthreads/meme_hu_4d36ab41a66a48cb.webp 565w, https://www.nijho.lt/post/advent-of-open-source/07-numthreads/meme_hu_f55150e31c7bd3b9.webp 891w" sizes="auto, (max-width: 760px) 100vw, 720px" width="565" height="760" alt="Meme showing a person labeled &amp;#39;YOUR CODE&amp;#39; about to be hugged by a pink creature labeled &amp;#39;NUMTHREADS.SET(1)&amp;#39;, representing setting the number of threads to 1. In the background, a yellow circle labeled &amp;#39;CPU PERFORMANCE&amp;#39; is visible." loading="lazy" data-zoomable&gt;&lt;figcaption&gt;Overenthusiastic threading&lt;/figcaption&gt;&lt;/figure&gt;&#10;&lt;h2 id="-technical-highlights"&gt;🔧 Technical Highlights &lt;a class="anchor" href="#-technical-highlights" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Controls thread count for major numerical libraries:&#10;&lt;ul&gt;&#10;&lt;li&gt;OpenBLAS&lt;/li&gt;&#10;&lt;li&gt;Intel&amp;rsquo;s Math Kernel Library (MKL)&lt;/li&gt;&#10;&lt;li&gt;OpenMP&lt;/li&gt;&#10;&lt;li&gt;NumExpr&lt;/li&gt;&#10;&lt;li&gt;Accelerate&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;Zero dependencies&lt;/li&gt;&#10;&lt;li&gt;Tiny (≤7KB) yet solves a real problem&lt;/li&gt;&#10;&lt;li&gt;Both CLI and Python API&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-impact"&gt;📊 Impact &lt;a class="anchor" href="#-impact" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Over my scientific computing career, I&amp;rsquo;ve encountered this threading issue dozens of times. Since creating this package a year ago, I&amp;rsquo;ve helped at least 5 different researchers speed up their workflows, sometimes by orders of magnitude. The beauty lies in its simplicity - the fix is often as simple as:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;eval&lt;/span&gt; &lt;span class="k"&gt;$(&lt;/span&gt;numthreads 1&lt;span class="k"&gt;)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="-lessons-learned"&gt;💡 Lessons Learned &lt;a class="anchor" href="#-lessons-learned" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;The simplest solutions are often the best&lt;/li&gt;&#10;&lt;li&gt;More cores doesn&amp;rsquo;t always mean faster code&lt;/li&gt;&#10;&lt;li&gt;If you find yourself repeating the same instructions, write a tool&lt;/li&gt;&#10;&lt;li&gt;Even a 7KB package can save hours of computing time&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;Want better control over your numerical computations? Check out &lt;a href="https://github.com/basnijholt/numthreads" target="_blank" rel="noopener"&gt;numthreads on GitHub&lt;/a&gt;!&lt;/p&gt;&#10;&lt;p&gt;#OpenSource #Python #HPC #ScientificComputing #Programming&lt;/p&gt;&#10;</description></item><item><title>Open Science Publications 🔬</title><link>https://www.nijho.lt/post/advent-of-open-source/06-open-science-publications/</link><pubDate>Fri, 06 Dec 2024 00:00:00 +0000</pubDate><guid isPermaLink="false">/post/advent-of-open-source/06-open-science-publications/</guid><description>&lt;p&gt;(See my &lt;a href="../"&gt;intro post&lt;/a&gt;)&lt;/p&gt;&#10;&lt;p&gt;Today&amp;rsquo;s post is about something I have strong opinions on: making scientific research reproducible. While many researchers talk about open science, actually making your work reproducible for anyone requires significant effort.&lt;/p&gt;&#10;&lt;h2 id="-origin-story"&gt;📖 Origin Story &lt;a class="anchor" href="#-origin-story" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;During my Ph.D. and subsequent research, I noticed a frustrating pattern in scientific publications: claims of &amp;ldquo;code available upon request&amp;rdquo; that often led nowhere, or code snippets that wouldn&amp;rsquo;t run without substantial modification. I made an effort that my publications would be different – each includes complete, runnable code that reproduces every figure and result.&lt;/p&gt;&#10;&lt;h2 id="-technical-highlights"&gt;🔧 Technical Highlights &lt;a class="anchor" href="#-technical-highlights" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Multiple repositories covering quantum physics experiments and simulations:&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;a href="https://github.com/basnijholt/orbitalfield" target="_blank" rel="noopener"&gt;orbitalfield&lt;/a&gt; - &amp;ldquo;Orbital effect of magnetic field on the Majorana phase diagram&amp;rdquo; (94 citations)&lt;/li&gt;&#10;&lt;li&gt;&lt;a href="https://github.com/basnijholt/supercurrent-majorana-nanowire" target="_blank" rel="noopener"&gt;supercurrent-majorana-nanowire&lt;/a&gt; - &amp;ldquo;Supercurrent interference in few-mode nanowire Josephson junctions&amp;rdquo; (78 citations)&lt;/li&gt;&#10;&lt;li&gt;&lt;a href="https://github.com/basnijholt/zigzag-majoranas" target="_blank" rel="noopener"&gt;zigzag-majoranas&lt;/a&gt; - &amp;ldquo;Enhanced proximity effect in zigzag-shaped Majorana Josephson junctions&amp;rdquo; (56 citations)&lt;/li&gt;&#10;&lt;li&gt;&lt;a href="https://github.com/microsoft/azure-quantum-tgp" target="_blank" rel="noopener"&gt;azure-quantum-tgp&lt;/a&gt; - &amp;ldquo;Protocol to identify a topological superconducting phase in a three-terminal device&amp;rdquo; (52 citations)&lt;/li&gt;&#10;&lt;li&gt;&lt;a href="https://github.com/basnijholt/spin-orbit-nanowires" target="_blank" rel="noopener"&gt;spin-orbit-nanowires&lt;/a&gt; - &amp;ldquo;Spin-orbit protection of induced superconductivity in Majorana nanowires&amp;rdquo; (81 citations)&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-evolution-of-best-practices"&gt;🔄 Evolution of Best Practices &lt;a class="anchor" href="#-evolution-of-best-practices" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;What I consider best practices has evolved over time. Today, I would:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Use &lt;a href="https://pixi.sh/" target="_blank" rel="noopener"&gt;Pixi&lt;/a&gt; for universal lock files across operating systems and programming languages&lt;/li&gt;&#10;&lt;li&gt;Provide a minimal Docker container that just installs the pixi lock file&lt;/li&gt;&#10;&lt;li&gt;Create self-documenting Jupyter notebooks that reproduce every result&lt;/li&gt;&#10;&lt;li&gt;Include clear figure-to-code mapping for paper reproducibility&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-challenges-and-solutions"&gt;🎯 Challenges and Solutions &lt;a class="anchor" href="#-challenges-and-solutions" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Balancing code cleanliness with research deadlines&lt;/li&gt;&#10;&lt;li&gt;Managing large datasets efficiently&lt;/li&gt;&#10;&lt;li&gt;Ensuring long-term reproducibility&lt;/li&gt;&#10;&lt;li&gt;Making complex physics simulations accessible&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="-lessons-learned"&gt;💡 Lessons Learned &lt;a class="anchor" href="#-lessons-learned" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;Clean code takes time but saves more in the long run&lt;/li&gt;&#10;&lt;li&gt;Documentation is as crucial as the code itself&lt;/li&gt;&#10;&lt;li&gt;Lock files are essential for true reproducibility&lt;/li&gt;&#10;&lt;li&gt;Making code public improves its quality&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;Want to explore quantum physics simulations? Check out the repositories above, each linked to their corresponding papers with full reproduction instructions.&lt;/p&gt;&#10;&lt;p&gt;#OpenSource #OpenScience #Physics #QuantumComputing #Python&lt;/p&gt;&#10;</description></item><item><title>Crafting unique Ph.D. thesis covers with Python 🎨</title><link>https://www.nijho.lt/post/thesis-cover/</link><pubDate>Sat, 04 Apr 2020 13:40:00 +0000</pubDate><guid isPermaLink="false">/post/thesis-cover/</guid><description>&lt;p&gt;&lt;em&gt;Originally posted on &lt;a href="https://quantumtinkerer.tudelft.nl/blog/thesis-cover/" target="_blank" rel="noopener"&gt;quantumtinkerer.tudelft.nl/blog/thesis-cover&lt;/a&gt; on 2020-04-04 13:40.&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;I will hand out around a hundred copies of my thesis to my defense committee, colleagues, and my family and friends.&#10;Let&amp;rsquo;s be honest, most people will probably not get further than attempting to read the summary and appreciating the cover.&#10;Four years of work has gone into generating the content of the thesis, so I figured, at least some thought has to go into the design of its cover.&#10;Unfortunately, I am by no means an expert on a graphical design or even competent enough to attempt to use any kind of graphic design software.&#10;But luckily for me, I do consider myself an expert in Python &lt;i class="fab fa-python"&gt;&lt;/i&gt;, and why not make the thesis design a fun process?&lt;/p&gt;&#10;&lt;h2 id="initial-design-adaptive-sampling-and-data-visualization"&gt;Initial Design: Adaptive Sampling and Data Visualization &lt;a class="anchor" href="#initial-design-adaptive-sampling-and-data-visualization" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;My vague plan was to somehow incorporate a plot resulting from a simulation run with &lt;a href="https://adaptive.readthedocs.io/" target="_blank" rel="noopener"&gt;Adaptive&lt;/a&gt; [&lt;a href="https://github.com/python-adaptive/adaptive" target="_blank" rel="noopener"&gt;source&lt;/a&gt; &lt;i class="fab fa-github"&gt;&lt;/i&gt;], the software I wrote during my Ph.D. to adaptively sample functions in parallel and my proudest work.&#10;I even had a potential simulation result in mind that I would use for it, specifically, one that has been on the background of my good friend’s phone for some years and one that has earned me quite some internet points on &lt;a href="https://www.reddit.com/r/dataisbeautiful/comments/eaap3p/python_quantum_research_art_saving_1000s_hours_of/" target="_blank" rel="noopener"&gt;Reddit&lt;/a&gt;.&#10;This plot (below) is one of the results of Chapter 4 of &lt;a href="http://files.nijho.lt/thesis.pdf" target="_blank" rel="noopener"&gt;my thesis&lt;/a&gt; [&lt;a href="https://github.com/basnijholt/thesis" target="_blank" rel="noopener"&gt;source&lt;/a&gt; &lt;i class="fab fa-github"&gt;&lt;/i&gt;] and &lt;a href="https://arxiv.org/abs/1807.01940%ef%bf%bc" target="_blank" rel="noopener"&gt;this paper&lt;/a&gt;.&lt;/p&gt;&#10;&lt;video style="width: 100%;max-width: 500px;" controls&gt;&#10; &lt;source src="https://user-images.githubusercontent.com/6897215/233811442-ce9bc475-af15-4b89-9578-67925419b769.mp4" type="video/mp4"&gt;&#10; Your browser does not support the video tag.&#10;&lt;/video&gt;&#10;&lt;p&gt;I went back to the code that generates the plot and started hacking away in Matplotlib.&#10;Quickly, I realized that not only the plot above was going to look cool, but that many other data sets that I had also resulted in beautiful patterns.&#10;I hit the paradox of choice: I had a total of &lt;em&gt;3671&lt;/em&gt; unique data sets (all for Majorana devices with either different parameter sweeps or different physical effects that were included in the model) and no idea how I was going to make a selection.&lt;/p&gt;&#10;&lt;p&gt;This is an early attempt (and rather ugly in hindsight)&lt;/p&gt;&#10;&lt;figure&gt;&lt;img src="https://user-images.githubusercontent.com/6897215/233811538-dc3b5b86-6f6f-4bd6-840c-e466918ce151.jpeg" alt="early-attempt" loading="lazy" data-zoomable&gt;&lt;/figure&gt;&lt;p&gt;I decided to postpone the choice of which data to use and just focus on how to plot the data first.&#10;I experimented with different colormaps, fading colormaps that gradually reveal the underlying triangulation better, and more.&#10;On a weekend night, after telling &lt;a href="https://github.com/sbalk" target="_blank" rel="noopener"&gt;a friend&lt;/a&gt; I was with about my plans, we (as usually happens) started to program some of the ideas we came up with.&#10;We settled on a design where the triangulation would slowly fade from the top to bottom and where each triangle simultaneously transitioned from a solid color (the average of the values at its vertices) to densely interpolated values.&#10;See &lt;a href="https://github.com/basnijholt/thesis-cover/blob/master/generate-covers.ipynb" target="_blank" rel="noopener"&gt;this Jupyter notebook&lt;/a&gt; which generated the final version of the cover.&#10;After we also chose a nice colormap (by simply creating the cover for each possible colormap,) the following was suggested: “You know what would be even cooler?! To make every cover unique.”&#10;Ten minutes later, I had started 3671 jobs on our computational cluster.&#10;When the files slowly started to appear on my computer (I have my local files synced with the cluster), we could not have been more excited.&#10;There were simply so many cool looking alternatives.&#10;Too many to manually select now, to be honest&amp;mdash;a problem for later I thought.&lt;/p&gt;&#10;&lt;h2 id="overcoming-the-paradox-of-choice-clustering-and-neural-networks"&gt;Overcoming the Paradox of Choice: Clustering and Neural Networks &lt;a class="anchor" href="#overcoming-the-paradox-of-choice-clustering-and-neural-networks" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Four months later, with the rest of my thesis completely done and about a day left on the deadline to submit my thesis, I found myself facing the selection problem I had successfully ignored for so long.&#10;There must be a smarter way to select the covers, I thought.&#10;Many of them look roughly similar (and some of them are even quite boring or not ecstatically pleasing), so it does not make sense to scroll through all of them.&#10;I could cluster them by similarity and then I would just have to choose the nicely looking clusters while being able to instantly discard many.&lt;/p&gt;&#10;&lt;p&gt;Finally, a good application to use neural networks!&lt;/p&gt;&#10;&lt;p&gt;I still had the 3671 covers in pdf format in some folder, so using the wonderful &lt;a href="https://www.gnu.org/software/parallel/" target="_blank" rel="noopener"&gt;GNU Parallel&lt;/a&gt;, I converted and rasterized all of them to jpeg with&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;find . -iname &lt;span class="se"&gt;\*&lt;/span&gt;pdf -print0 &lt;span class="p"&gt;|&lt;/span&gt; parallel -0 sips -s format jpeg &lt;span class="o"&gt;{}&lt;/span&gt; --out &lt;span class="o"&gt;{&lt;/span&gt;.&lt;span class="o"&gt;}&lt;/span&gt;.jpg&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Then I used a pre-trained convolutional neural network model (&lt;a href="https://keras.io/applications/#vgg16" target="_blank" rel="noopener"&gt;VGG16&lt;/a&gt;) to extract features in a lower-dimensional representation.&#10;I fed the feature arrays to a &lt;a href="https://en.wikipedia.org/wiki/K-means_clustering" target="_blank" rel="noopener"&gt;K-Means clustering algorithm&lt;/a&gt; that sorted them in 80 clusters.&#10;To procrastinate a little more, I created an interactive widget using &lt;a href="https://ipywidgets.readthedocs.io/" target="_blank" rel="noopener"&gt;&lt;code&gt;ipywidgets&lt;/code&gt;&lt;/a&gt; (see video below) to facilitate the selection process.&lt;/p&gt;&#10;&lt;video style="width: 100%;max-width: 800px;" controls&gt;&#10; &lt;source src="https://user-images.githubusercontent.com/6897215/233811575-98ea0be2-33f4-47f4-b3ee-f8e7bdf3c475.mp4" type="video/mp4"&gt;&#10; Your browser does not support the video tag.&#10;&lt;/video&gt;&#10;&lt;p&gt;These are the clusters that I selected in a grid&lt;/p&gt;&#10;&lt;figure&gt;&lt;img src="https://user-images.githubusercontent.com/6897215/233811546-36be2cec-eb8c-437a-88d1-595a8602a7ba.jpeg" alt="chosen-clusters-front-and-back" loading="lazy" data-zoomable&gt;&lt;/figure&gt;&lt;p&gt;The selection process was done using &lt;a href="https://github.com/basnijholt/thesis-cover/blob/master/thesis-cover-selection.ipynb" target="_blank" rel="noopener"&gt;this Jupyter notebook&lt;/a&gt;.&lt;/p&gt;&#10;&lt;p&gt;Finally, I copied the selected clusters into different folders and used the OS&amp;rsquo;s file explorer to manually select a few covers from each cluster.&lt;/p&gt;&#10;&lt;h2 id="final-touches-fine-tuning-colors-and-adding-edition-numbers"&gt;Final Touches: Fine-Tuning Colors and Adding Edition Numbers &lt;a class="anchor" href="#final-touches-fine-tuning-colors-and-adding-edition-numbers" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Then I printed an initial version of my thesis, however, I wasn&amp;rsquo;t really happy with the colors.&lt;/p&gt;&#10;&lt;p&gt;To choose the right color scheme I got help from a non-tech savvy friend who went to art school.&#10;Due to the COVID-19 crisis, we couldn&amp;rsquo;t sit side by side, so I created &lt;a href="https://mybinder.org/v2/gh/basnijholt/thesis-cover/master?filepath=color-picker.ipynb" target="_blank" rel="noopener"&gt;this Binder &amp;ldquo;app&amp;rdquo;&lt;/a&gt; with sliders.&lt;/p&gt;&#10;&lt;video style="width: 100%;max-width: 800px;" controls&gt;&#10; &lt;source src="https://user-images.githubusercontent.com/6897215/233811609-8fa74955-0c51-443e-9a87-115921530b7d.mp4" type="video/mp4"&gt;&#10; Your browser does not support the video tag.&#10;&lt;/video&gt;&#10;&lt;p&gt;Additionally, since all of them were different anyway, I gave them an identifier and added&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;edition n of 120&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;to the back.&lt;/p&gt;&#10;&lt;p&gt;These are a few examples of the final result (click on one of the images to open gallery view)&lt;/p&gt;&#10;&lt;div class="gallery"&gt;&#10; &lt;a href="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-alpha__data_learner_0061.pickle.jpg"&gt;&lt;img src="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-alpha__data_learner_0061.pickle_hu_8e2f6288e2299f4b.webp" width="350" height="242" alt="Image 1 of 12" loading="lazy" data-zoomable&gt;&lt;/a&gt;&#10; &lt;a href="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-alpha__data_learner_0087.pickle.jpg"&gt;&lt;img src="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-alpha__data_learner_0087.pickle_hu_cfe0a7647b8b4d1a.webp" width="350" height="242" alt="Image 2 of 12" loading="lazy" data-zoomable&gt;&lt;/a&gt;&#10; &lt;a href="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-alpha__data_learner_0094.pickle.jpg"&gt;&lt;img src="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-alpha__data_learner_0094.pickle_hu_4767f3839b1f22bc.webp" width="350" height="242" alt="Image 3 of 12" loading="lazy" data-zoomable&gt;&lt;/a&gt;&#10; &lt;a href="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-alpha__data_learner_0119.pickle.jpg"&gt;&lt;img src="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-alpha__data_learner_0119.pickle_hu_9a20ab2d5318b96f.webp" width="350" height="242" alt="Image 4 of 12" loading="lazy" data-zoomable&gt;&lt;/a&gt;&#10; &lt;a href="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-alpha__data_learner_0254.pickle.jpg"&gt;&lt;img src="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-alpha__data_learner_0254.pickle_hu_e670a31aca40cc2e.webp" width="350" height="242" alt="Image 5 of 12" loading="lazy" data-zoomable&gt;&lt;/a&gt;&#10; &lt;a href="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-alpha2__data_learner_0175.pickle.jpg"&gt;&lt;img src="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-alpha2__data_learner_0175.pickle_hu_236481dab96c15b5.webp" width="350" height="242" alt="Image 6 of 12" loading="lazy" data-zoomable&gt;&lt;/a&gt;&#10; &lt;a href="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-rotation-0-90-move-so-slowly__data_learner_0457.pickle.jpg"&gt;&lt;img src="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-rotation-0-90-move-SO-slowly__data_learner_0457.pickle_hu_cc81d6c6a833f7ce.webp" width="350" height="242" alt="Image 7 of 12" loading="lazy" data-zoomable&gt;&lt;/a&gt;&#10; &lt;a href="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-rotation-0-90-move-so-slowly__data_learner_0699.pickle.jpg"&gt;&lt;img src="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-rotation-0-90-move-SO-slowly__data_learner_0699.pickle_hu_4744ce916ab208d2.webp" width="350" height="242" alt="Image 8 of 12" loading="lazy" data-zoomable&gt;&lt;/a&gt;&#10; &lt;a href="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-rotation-0-90-move-so-slowly__data_learner_0811.pickle.jpg"&gt;&lt;img src="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__gradient-sweep-rotation-0-90-move-SO-slowly__data_learner_0811.pickle_hu_5f6e91d59b7c978e.webp" width="350" height="242" alt="Image 9 of 12" loading="lazy" data-zoomable&gt;&lt;/a&gt;&#10; &lt;a href="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__mu-sweep2__data_learner_0179.pickle.jpg"&gt;&lt;img src="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__mu-sweep2__data_learner_0179.pickle_hu_ad9a40d0f92c3c0f.webp" width="350" height="242" alt="Image 10 of 12" loading="lazy" data-zoomable&gt;&lt;/a&gt;&#10; &lt;a href="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__mu-sweep2__data_learner_0254.pickle.jpg"&gt;&lt;img src="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__mu-sweep2__data_learner_0254.pickle_hu_2e0e64805618b36d.webp" width="350" height="242" alt="Image 11 of 12" loading="lazy" data-zoomable&gt;&lt;/a&gt;&#10; &lt;a href="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__mu-sweep2__data_learner_0255.pickle.jpg"&gt;&lt;img src="https://www.nijho.lt/media/albums/basnijholt-thesis-cover/data__mu-sweep2__data_learner_0255.pickle_hu_e7bc31a50788a8cd.webp" width="350" height="242" alt="Image 12 of 12" loading="lazy" data-zoomable&gt;&lt;/a&gt;&#10;&lt;/div&gt;&#10;&#10;&lt;p&gt;A few days later, this was delivered to my doorstep&lt;/p&gt;&#10;&lt;p&gt;&lt;img src="https://user-images.githubusercontent.com/6897215/233811518-bd5426d6-05bc-4c7a-854a-4c373aa21fcf.jpg" alt="theses-everywhere" loading="lazy" data-zoomable&gt;&#10;&lt;img src="https://user-images.githubusercontent.com/6897215/233811525-645b67a4-a9ea-48ba-a391-4639ed92c19d.jpg" alt="printed" loading="lazy" data-zoomable&gt;&lt;/p&gt;&#10;&lt;p&gt;And finally I mailed them to my Ph.D. committee&lt;/p&gt;&#10;&lt;figure&gt;&lt;img src="https://user-images.githubusercontent.com/6897215/233811533-3b85ff8d-8359-453b-8aa1-ef653629a0f4.jpg" alt="mail" loading="lazy" data-zoomable&gt;&lt;/figure&gt;&lt;p&gt;Now it&amp;rsquo;s time to start preparing for my defense at May 11, 2020 at 15:00!&lt;/p&gt;&#10;&lt;h3 id="notes"&gt;Notes &lt;a class="anchor" href="#notes" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;The data and code to generate all printed covers is available at &lt;a href="https://github.com/basnijholt/thesis-cover" target="_blank" rel="noopener"&gt;basnijholt/thesis-cover&lt;/a&gt; &lt;i class="fab fa-github"&gt;&lt;/i&gt;&lt;/li&gt;&#10;&lt;li&gt;The simulation code that generates the data is available at &lt;a href="https://github.com/basnijholt/spin-orbit-nanowires" target="_blank" rel="noopener"&gt;basnijholt/spin-orbit-nanowires&lt;/a&gt; &lt;i class="fab fa-github"&gt;&lt;/i&gt;&lt;/li&gt;&#10;&lt;li&gt;The video of the Reddit post is generated using &lt;a href="https://gist.github.com/basnijholt/7aaeb7b9c506f830e6791d264a6cf467" target="_blank" rel="noopener"&gt;this code&lt;/a&gt; &lt;i class="fab fa-github"&gt;&lt;/i&gt;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;</description></item></channel></rss>