<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>emre şahin's digital garden 🍃 - Concurrency</title>
    <link>https://emresahin.net/tags/concurrency/</link>
    <description>Posts in the Concurrency tag</description>
    <language>en</language>
    <managingEditor>contact@emresahin.net (Emre Şahin)</managingEditor>
    <lastBuildDate>Tue, 29 Sep 2026 14:57:43 +0000</lastBuildDate>
    <atom:link href="https://emresahin.net/tags/concurrency/rss.xml" rel="self" type="application/rss+xml"/>
    <item>
      <title>Syncing Path Operations in Xvc</title>
      <published>2024-06-04T19:50:37+00:00</published>
      <updated>2024-06-04T19:50:37+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Tue, 04 Jun 2024 19:50:37 +0000</pubDate>
      <link>https://emresahin.net/syncing-path-operations-in-xvc/</link>
      <guid isPermaLink="true">https://emresahin.net/syncing-path-operations-in-xvc/</guid>
      <description>While writing a HOWTO post for the documentation, I found a bug where multiple carry-in commands were causing file system failures. When multiple threads were accessing the same cache directory, if one of them tried to set up the cache directory while another was still working on it, it caused a ...</description>
      <category>xvc</category>
      <category>Development</category>
      <category>Rust</category>
      <category>Concurrency</category>
      <category>File System</category>
      <category>Locking</category>
      <category>Path</category>
      <category>Multithreading</category>
      <category>Parallel</category>
      <category>OS</category>
      <content:encoded><![CDATA[<p>While writing a HOWTO post for the documentation, I found a bug where multiple
carry-in commands were causing file system failures. When multiple threads were
accessing the same cache directory, if one of them tried to set up the cache directory
while another was still working on it, it caused a permissions error.</p>
<p>Rust has <em>fearless concurrency</em> for memory access, but for the file system, there
seem to be no built-in locked access primitives.</p>
<p>I decided to write one. Parallel execution of file system operations is important
for Xvc. The error messages are annoying; having identical files in a repository is
common, and in those cases, these messages look like there is a problem. In
theory, when files are identical, having only one of them written to the cache
is not a problem, but there may be other issues preventing cache
access. We could just swallow the error and get away with it, but that’s not ideal.</p>
<p>Two options came to mind. One is modifying the list of cache files before
creating threads so that no two threads access the same cache file at the same
time. This is hard to implement and brings extra complexity to thread creation.
A one-in-a-thousand concern becomes an architectural burden.</p>
<p>The other solution is to lock paths while accessing them so two threads
working on the same cache path wait for each other. This is easier and requires
just dependency injection into the thread functions. It has the downside of
making file system operations slightly slower, as each path operation will now require
checking a mutex, but that seems of little concern for file system access, which
is already much slower than memory operations.</p>
<p>First, I tried to implement this with a <code>HashMap&lt;PathBuf, Mutex&lt;()&gt;&gt;</code>, but this
must also be passed to the function wrapped in <code>Arc&lt;Mutex&lt;HashMap&gt;&gt;</code>, which
made the “ceremony” of acquiring the lock for a single file much longer.</p>
<p>The issue is that you don’t want the <code>HashMap</code> itself to be a bottleneck. Multiple
threads shouldn’t wait for the <code>HashMap</code> to become available, as it’s not the
<code>HashMap</code> we want to lock, but the values inside it.</p>
<p>The solution is to return the lock value from a method of a struct. Something
like:</p>
<pre><code class="language-rust">pub struct PathSync {
    locks: Arc&lt;RwLock&lt;HashMap&lt;PathBuf, Arc&lt;Mutex&lt;()&gt;&gt;&gt;&gt;&gt;,
}</code></pre>
<p>Now the ceremony can be performed within the method, and threads working in
different directories won’t need to wait for the hash map to become available.</p>
<p>However, after implementing this, I realized I’d probably forget to lock a path
at some point. This is a general-purpose solution, and I should apply it to all
path operations when multiple threads are working. In most cases,
there are multiple paths to lock (cache_path, cache_dir, repository path), and if
I forget to lock one of them, a future user, some time, somewhere, will probably
see an error message.</p>
<p>So, I decided on a different approach: creating wrappers to run passed closures. This makes it
much more obvious that paths Xvc works on must be locked before operations.</p>
<pre><code class="language-rust">    pub fn with_sync_path(
        &amp;self,
        path: &amp;Path,
        mut f: impl FnMut(&amp;Path) -&gt; Result&lt;()&gt;,
    )</code></pre>
<p>This works by passing the path and a closure that operates on that path. It
first locks the path and then runs the closure. The locking mechanism allows
threads with different paths to run in parallel, but if they try to operate on the same path, they
will wait for each other.</p>
<p>The implementation is <a href="https://github.com/iesahin/xvc/blob/main/walker/src/sync.rs">here</a>.</p>]]></content:encoded>
    </item>
    <item>
      <title>LMAX Disruptor</title>
      <published>2024-01-25T09:01:54+00:00</published>
      <updated>2024-01-25T09:01:54+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Thu, 25 Jan 2024 09:01:54 +0000</pubDate>
      <link>https://emresahin.net/lmax-disruptor/</link>
      <guid isPermaLink="true">https://emresahin.net/lmax-disruptor/</guid>
      <description>LMAX Disruptor The Disruptor is a concurrency framework created by LMAX to achieve very low latency and high throughput on their Java trading platform. LMAX found that using queues between system stages introduced latency, leading them to focus on optimizing this area. The Disruptor uses a lock-f...</description>
      <category>Development</category>
      <category>Summary</category>
      <category>Concurrency</category>
      <category>Summary</category>
      <category>Libraries</category>
      <category>Disruptor</category>
      <category>Concurrency</category>
      <category>Java</category>
      <category>Performance</category>
      <content:encoded><![CDATA[<p><a href="https://lmax-exchange.github.io/disruptor/">LMAX Disruptor</a></p>
<ul>
<li>The Disruptor is a concurrency framework created by LMAX to achieve very low latency and high throughput on their Java trading platform.</li>
<li>LMAX found that using queues between system stages introduced latency, leading them to focus on optimizing this area.</li>
<li>The Disruptor uses a lock-free “ring buffer” approach to avoid cache misses and locks at the CPU level, which are very costly.</li>
<li>The Disruptor is intended as a general-purpose solution for concurrent programming, not just for financial applications.</li>
<li>Applying the Disruptor pattern is not as simple as replacing all queues with the ring buffer; the user guide provides further guidance.</li>
<li>Various blogs, articles, technical papers, and performance tests explain the Disruptor’s inner workings.</li>
<li>A presentation and discussion group are also available for learning about the Disruptor from its creators.</li>
<li>Martin Fowler has reviewed the Disruptor’s use at LMAX.</li>
<li>The Disruptor is significantly faster than the ArrayBlockingQueue, as shown in the latency histogram.</li>
<li>More performance results are available comparing the Disruptor to other approaches.</li>
</ul>]]></content:encoded>
    </item>
    <item>
      <title>bits 4</title>
      <published>2023-03-14T11:19:00+00:00</published>
      <updated>2023-03-14T11:19:00+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Tue, 14 Mar 2023 11:19:00 +0000</pubDate>
      <link>https://emresahin.net/bits-4/</link>
      <guid isPermaLink="true">https://emresahin.net/bits-4/</guid>
      <description>I want to use channels for Xvc pipelines. A thread will be created for each step and will receive dependency state changes via channels. Note that there may be thousands of dependencies (e.g., files for a step), and creating these channels beforehand is not feasible. The Crossbeam library has Sel...</description>
      <category>bits</category>
      <category>xvc</category>
      <category>Rust</category>
      <category>crossbeam</category>
      <category>select</category>
      <category>channels</category>
      <category>concurrency</category>
      <content:encoded><![CDATA[<p>I want to use channels for Xvc pipelines. A thread will be created for each
step and will receive dependency state changes via channels.</p>
<p>Note that there may be thousands of dependencies (e.g., files for a step), and
creating these channels beforehand is not feasible.</p>
<p>The Crossbeam library has
<a href="https://docs.rs/crossbeam/latest/crossbeam/channel/struct.Select.html#">Select</a>
which allows waiting for an arbitrary number of channels. Now, I need to figure out
how to structure the threads and update their states. A minor task! 😆</p>]]></content:encoded>
    </item>
    <item>
      <title>Software architecture as a tree</title>
      <published>2022-05-12T04:33:08+00:00</published>
      <updated>2022-05-12T04:33:08+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Thu, 12 May 2022 04:33:08 +0000</pubDate>
      <link>https://emresahin.net/Execution-and-memory-as-a-tree/</link>
      <guid isPermaLink="true">https://emresahin.net/Execution-and-memory-as-a-tree/</guid>
      <description>In Rust, there are several ways to manage memory. Memory has two regions: One is called the stack . Each new scope (between { and } ) adds variables to the stack. Variables on the stack are allocated at once when a new scope is introduced, and they are freed when the scope ends. This brings speed...</description>
      <category>Software Architecture</category>
      <category>Memory Management</category>
      <category>Rust</category>
      <category>Memory Management</category>
      <category>Stack</category>
      <category>Heap</category>
      <category>Tree</category>
      <category>Concurrency</category>
      <category>Software Design</category>
      <content:encoded><![CDATA[<p>In Rust, there are several ways to manage memory. Memory has two regions:
One is called the <em>stack</em>. Each new scope (between <code>{</code> and <code>}</code>) adds variables to
the stack. Variables on the stack are allocated at once when a new scope is
introduced, and they are freed when the scope ends. This brings speed.
Variables on the stack are also allocated contiguously, so they are <em>more
compatible</em> with the principle of locality.</p>
<pre><code>{
  var1
  var2
  {
    var3
    var4
  }
}
</code></pre>
<p>The inner scope variables can’t be accessed by the outer scope, but outer
scope variables can be accessed by the inner scope. When we are in the inner
scope, the stack contains both scopes’ variables. Stack-based memory
management can be called table d’hôte. It’s faster and simpler.</p>
<p>However, in order to allocate them at once, variable sizes must be known
beforehand. If the compiler doesn’t know the size of an array, it can’t
allocate the array on the stack. We need another method to use these <em>indefinite</em>
variables.</p>
<p>The other way of obtaining memory is using the <em>heap</em>. The heap is independent of
the scope. You can allocate and fill it when you need with a size determined
at runtime, share it with other scopes, and free it à la carte. Rust
provides several different ways to manage memory on the heap.</p>
<p>In my experience, making the program flow as scoped as possible without cross-cutting
memory management is a good sign. Ideally, a program should have a tree-like
flow of execution, and the number of heap allocations is minimized in this
case. However, most modern software systems are not designed this way, and
memory management is not considered during the requirements phase. When you
leave this to a garbage collector, even if it does its best, memory management often
remains haphazard.</p>
<p>Ideally, when you run the program, it should call a series of functions that
branch into further sub-functions:</p>
<pre class="mermaid">
flowchart LR

  main---&gt;fn1
  main---&gt;fn2
  main---&gt;fn3
  main---&gt;fn4
  fn1---&gt;fn1.1
  fn1---&gt;fn1.2
  fn1---&gt;fn1.3
  fn1---&gt;fn1.4
  fn2---&gt;fn2.1
  fn2---&gt;fn2.2
  fn2---&gt;fn2.3
  fn2---&gt;fn2.4

  fn1.1--&gt;fn1.1.1
  fn1.1--&gt;fn1.1.2
  fn1.1--&gt;fn1.1.3
  fn1.1--&gt;fn1.1.4
  fn1.1--&gt;fn1.1.5


</pre>

<p>If we adhere to this principle, building concurrent software becomes a lot
easier, too. If the functionality and memory of <code>fn1.1.1</code> and <code>fn2</code> are
completely separate, these can be run in parallel. Hence, our aim in
<em>software architecture</em> must be looking for ways to represent the problem at
hand as a tree structure, both physically (for memory) and temporally
(for execution).</p>]]></content:encoded>
    </item>
    <item>
      <title>Using a Single Threaded Functor in Multiple Threads with Futures in C++</title>
      <published>2013-01-13T14:00:00+00:00</published>
      <updated>2013-01-13T14:00:00+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Sun, 13 Jan 2013 14:00:00 +0000</pubDate>
      <link>https://emresahin.net/12236-21-1527/</link>
      <guid isPermaLink="true">https://emresahin.net/12236-21-1527/</guid>
      <description>Multithreaded programming requires a paradigm shift when it comes to the return values of functions. C++11 provides std::async to run functions asynchronously, but this is not available in older versions. My current project on word spotting in historical documents is fairly complete in functional...</description>
      <category>C++</category>
      <category>Programming</category>
      <category>multithreading</category>
      <category>boost</category>
      <category>c++11</category>
      <category>futures</category>
      <category>concurrency</category>
      <content:encoded><![CDATA[<p>Multithreaded programming requires a paradigm shift when it comes to the return
values of functions. C++11 provides
<a href="http://en.cppreference.com/w/cpp/thread/async"><code>std::async</code></a> to run functions
asynchronously, but this is not available in older versions.</p>
<p>My current project on word spotting in historical documents is fairly
complete in functionality, but I decided that searching for word images on
page images <em>concurrently</em> would be better for speeding it up. I’m already using
<a href="http://boost.org">Boost</a> for much of the functionality, and instead of
creating a dependency on the not-yet-mature C++11 support in various
compilers, I decided to use <code>boost::thread</code>.</p>
<p>Suppose we have a functor such as:</p>
<pre><code class="language-c++">class Search_t
{
   public:
       Search_t(Document d) { ... };
       SearchResult operator()(SearchItem i) { ... };
};
</code></pre>
<p>And we want to use this functor in multiple threads. We can’t simply do the following:</p>
<pre><code class="language-c++">std::vector&lt;SearchResult&gt; results;
Search_t search(document);
// search_items is a vector&lt;SearchItem&gt; and si is an iterator over it.
for (si = search_items.begin(); si != search_items.end(); ++si)
{
   boost::thread task(boost::bind(search, *si));
   results.push_back(task); // ERROR!
}
</code></pre>
<p>This is because <code>task</code> does not return a <code>SearchResult</code>.</p>
<p>Instead, we need to store results within the object and retrieve them
after they are generated.</p>
<p>I didn’t want to change the interface of <code>Search_t</code> because
multithreading should be optional, and other parts of the program may
depend on this interface. Instead, a wrapper class that runs these
threads with a similar interface seemed like a better solution.</p>
<pre><code class="language-c++">class SearchMT_t
{
   boost::shared_ptr&lt;std::vector&lt;boost::unique_future&lt;SearchResult&gt; &gt; &gt; futures_;

   public:
   SearchMT_t(Document d) :
   /* The most important assumption here is that Search_t does not alter
      the Document object's state in any way. Otherwise, we need to ensure that 
      document_ is accessed by only a single thread at a time using mutexes. */
   document_(d),
   futures_(new std::vector&lt;boost::unique_future&lt;SearchResult&gt; &gt;)
   {};

   void operator()(SearchItem si)
   {
       /* If you are sure that there won't be any race conditions
       between Search_t threads during search, you can move the following line
       to the constructor and use a single object for all searches. */
       Search_t search(document_);

       boost::packaged_task&lt;SearchResult&gt; search_task(std::bind(search, si));
       futures_-&gt;push_back(search_task.get_future());
       boost::thread task(boost::move(search_task));
   };

   std::vector&lt;SearchResult&gt; results()
   {
      std::vector&lt;SearchResult&gt; results;

      /* Wait for all threads to complete their work. */
      boost::wait_for_all(futures_-&gt;begin(), futures_-&gt;end());

      for(int i = 0; i &lt; futures_-&gt;size(); ++i)
      {
          results.push_back((*futures_)[i].get());
      }

      return results;
   }
};
</code></pre>
<p>This way, it becomes much more straightforward to use multithreading in
a loop:</p>
<pre><code class="language-c++">std::vector&lt;SearchResult&gt; results;
SearchMT_t search(document);
// search_items is a vector&lt;SearchItem&gt; and si is an iterator over it.
for (si = search_items.begin(); si != search_items.end(); ++si)
{
    search(*si);
}

results = search.results();
</code></pre>
<p>We kept the <code>Search_t</code> class intact and used a much simpler
approach in the loop.</p>]]></content:encoded>
    </item>
  </channel>
</rss>
