<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>emre şahin's digital garden 🍃 - xvc</title>
    <link>https://emresahin.net/tags/xvc/</link>
    <description>Posts in the xvc tag</description>
    <language>en</language>
    <managingEditor>contact@emresahin.net (Emre Şahin)</managingEditor>
    <lastBuildDate>Tue, 29 Sep 2026 14:57:43 +0000</lastBuildDate>
    <atom:link href="https://emresahin.net/tags/xvc/rss.xml" rel="self" type="application/rss+xml"/>
    <item>
      <title>devlog 30</title>
      <published>2025-05-11T18:17:52+00:00</published>
      <updated>2025-05-11T18:17:52+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Sun, 11 May 2025 18:17:52 +0000</pubDate>
      <link>https://emresahin.net/devlog-30/</link>
      <guid isPermaLink="true">https://emresahin.net/devlog-30/</guid>
      <description>🐢 Let’s discuss how to move Xvc forward—maybe we can write a post to Reddit and the Rust forum in the meantime. 🐇 I think the next step is rclone remotes. It will allow us to use all remote storages supported by Rclone, which is a nice feature. 🐢 Don’t you think we need to publish the current ver...</description>
      <category>devlog</category>
      <category>xvc</category>
      <category>rclone</category>
      <category>rsync</category>
      <category>ecs</category>
      <category>ecs index</category>
      <category>architecture</category>
      <category>storage</category>
      <category>doctor</category>
      <content:encoded><![CDATA[<p>🐢 Let’s discuss how to move Xvc forward—maybe we can write a post to Reddit and the Rust forum in the meantime.
🐇 I think the next step is rclone remotes. It will allow us to use all remote storages supported by Rclone, which is a nice feature.
🐢 Don’t you think we need to publish the current version to Reddit and the forum?
🐇 We can do that as well.
🦊 Adding rclone remote must be a straightforward task.
🐢 We need to understand rclone paths, but overall, yes. We’ll just need to get the remote name, like <code>drive://</code>, and a path, like <code>my-xvc-storage</code>, and build paths with these.
🐇 What are the commands?
🐢 We need to learn how to upload files from local to remote and how to download these files. We can also list the files and get files as well.
🐲 How about adding a <code>paths.txt</code> to folders in remotes to show which paths the files in <code>0.jpg</code> belong to? This will change the remote cache structure a bit. We will have a reverse index of files and they will be findable.
🐢 What’s the reason for this?
🐲 When I upload a file to Drive with only the content hash, I lose track of the actual path. This is not desirable. We can add a file to the directory, called <code>paths.txt</code>, to get the paths for a file.
🐢 This may prove to be a feat, though; adding these <code>XvcPaths</code> to a file requires a lookup.
🐇 Maybe a JSON file? It might be possible to look up a path with a JSON file, and it will be easier to parse.
🐲 I don’t think the issue is about parsing, though. We can just have a plain text file that lists the paths. It’s a text file, which is the most compatible across all storages.
🐢 Storages, you mean.
🐲 Ugh, yeah. If I have a file called <code>Alan Watts</code> but I only have the content, this file will be immensely useful.
🐢 This makes <code>XvcCachePath</code> and <code>XvcPath</code> coupled. Architecture-wise, it may not be a good thing, though.
🐇 Also, there may be common storages for multiple repositories.
🐲 Umm, that’s a good point. I don’t think the architecture will be much compromised, though. We already keep the file paths and their cache paths somewhere.
🐢 Cache paths are generated from the content, but any number of paths can point to a single path in the cache. If I have 1 million copies of the same file, will I add all these files to the <code>paths.txt</code> you mentioned?
🐲 That’s a good point too. We can have a limit, like 1,000 or something, not to make these files too big.
🐢 Instead of this, we can store the output of <code>xvc file list</code> at the storage root and allow looking up the files that way.
🐲 It has the same problem, though; if we have a million files, their list will be too large.
🐢 There can be a manual command, like <code>xvc file index --to storage</code>, that will show content hashes and paths of each file. We can also add URLs to files if possible.
🐲 No one will use it when it’s manual, though.
🐢 We can add functionality to update this index when we send a file, though.<br>🐲 So, after each send, we’ll update the index for the repository on that storage. Is that correct?
🐢 Not after each send. After each send session, maybe.
🐇 We can have an incremental way of updating the index, like we do in ECS?
🐢 It will be overkill for this functionality and add too much noise to the storage.
🐲 Let’s keep this discussion here, but I also want to have an index merge or index cleanup mechanism for the entity generator and the ECS.
🐢 We can have a “merge indices” functionality in ECS. That will remove all older entity-generator files and merge all store files.
🐇 Removing older entity files is easy, but what about merging the store files?
🐢 It’s easy too. We’ll just load all event logs from the directory, remove all other files, and save the event log to a file.
🐇 Will this be manual or automatic?
🐢 I think the first version can be manual, something like <code>xvc fsck merge-store-files</code> or something like that. We can notify the user if the number of files is &gt; 10,000 or something like that. I don’t think we need to make it automatic unless we measure the impact of these files. There is no point in trying to do it at every command.
🐇 Then we’ll have two new commands for the next version?
🐢 I think we can just add rclone remote now and release it, then make changes in the ECS for this new <code>xvc fsck</code> command.
🐇 Can the name be <code>doctor</code> or something? Or <code>util</code>? Or can we add a top-level <code>merge indices</code> command?
🐢 <code>xvc doctor</code> seems like a better alternative. We can have a <code>diagnose</code> subcommand as well to check for possible inconsistencies. <code>xvc doctor merge-store-files</code> is a better command.
🐲 Will we use <code>d</code> for this command?
🐢 No need to add a single-letter command for this, I believe. It shouldn’t be required to run frequently.
🐇 Hmm, ok. What do we need to know for rclone remote?
🐲 I noticed we don’t have the <code>xvc storage remove</code> command implemented yet. Maybe we can start from that.
🐢 Hmm, yeap. Let’s start by implementing that first. We can add the rclone command next.
🐇 Will we use a feature flag for rclone? It will run the command only with an external binary.
🐢 It’s better to have a feature flag. I think we can add a feature flag for rsync remote as well.
🐇 We can use the generic one to update the feature flag.
🐢 I think the only two items of information we need for rclone are the remote name and the remote directory. Will we make these required?</p>]]></content:encoded>
    </item>
    <item>
      <title>devlog 25</title>
      <published>2025-04-24T02:57:26+00:00</published>
      <updated>2025-04-24T02:57:26+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Thu, 24 Apr 2025 02:57:26 +0000</pubDate>
      <link>https://emresahin.net/devlog-25/</link>
      <guid isPermaLink="true">https://emresahin.net/devlog-25/</guid>
      <description>🐢 I want to update xvc.py to the latest version. 🐇 It should only be needed to update the dependency versions in Cargo.toml , right? 🐢 Let’s start with that. 🐇 We have interface changes regarding aliases; let’s start using uv for building. 🦊 Added requirements to pyproject.toml by running: uv add...</description>
      <category>XVC</category>
      <category>Python</category>
      <category>Development</category>
      <category>XVC</category>
      <category>Python</category>
      <category>uv</category>
      <category>Package Management</category>
      <category>Requirements</category>
      <content:encoded><![CDATA[<p>🐢 I want to update <code>xvc.py</code> to the latest version.</p>
<p>🐇 It should only be needed to update the dependency versions in <code>Cargo.toml</code>, right?</p>
<p>🐢 Let’s start with that.</p>
<p>🐇 We have interface changes regarding aliases; let’s start using <code>uv</code> for building.</p>
<p>🦊 Added requirements to <code>pyproject.toml</code> by running:</p>
<pre><code class="language-bash">uv add -r requirements.txt
</code></pre>]]></content:encoded>
    </item>
    <item>
      <title>devlog 24</title>
      <published>2025-04-24T02:54:54+00:00</published>
      <updated>2025-04-24T02:54:54+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Thu, 24 Apr 2025 02:54:54 +0000</pubDate>
      <link>https://emresahin.net/devlog-24/</link>
      <guid isPermaLink="true">https://emresahin.net/devlog-24/</guid>
      <description>🐇 Is there a way to implement Default for CLI structs? 🐢 There is a way if we start from the configuration, not just files. The default configuration is a TOML document. We can make it an XvcConfiguration struct and load and store it with confy . 🐇 We have a cascading set of configurations, but m...</description>
      <category>XVC</category>
      <category>Development</category>
      <category>XVC</category>
      <category>Rust</category>
      <category>Configuration</category>
      <category>confy</category>
      <category>CLI</category>
      <content:encoded><![CDATA[<p>🐇 Is there a way to implement <code>Default</code> for CLI structs?</p>
<p>🐢 There is a way if we start from the configuration, not just files. The default
configuration is a TOML document. We can make it an <code>XvcConfiguration</code> struct and
load and store it with <code>confy</code>.</p>
<p>🐇 We have a cascading set of configurations, but maybe we can start from the
struct and serialize/deserialize it on demand.</p>
<p>🐲 I don’t think we need to update <code>xvc-config</code> at the moment. It’s working, and
we don’t need to alter its inner workings in the near future.</p>
<p>🐢 I agree. We can use the <a href="https://docs.rs/config/"><code>config</code></a> crate when we need to update and have
enough time to work on this.</p>]]></content:encoded>
    </item>
    <item>
      <title>devlog 23</title>
      <published>2025-02-03T10:30:44+00:00</published>
      <updated>2025-02-03T10:30:44+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Mon, 03 Feb 2025 10:30:44 +0000</pubDate>
      <link>https://emresahin.net/devlog-23/</link>
      <guid isPermaLink="true">https://emresahin.net/devlog-23/</guid>
      <description>🐇 Let’s turn to discussing the JSON output changes. We can add another option to XvcOutputLine , like XvcOutputLine::Json(T: Serialize) , that will output the type using serde_json . It won’t introduce any other type. 🐢 Yes, but what will T be? Although store structures have Serde implementations...</description>
      <category>XVC</category>
      <category>Development</category>
      <category>XVC</category>
      <category>Rust</category>
      <category>Serde</category>
      <category>JSON</category>
      <category>Serialization</category>
      <category>xvc file list</category>
      <content:encoded><![CDATA[<p>🐇 Let’s turn to discussing the JSON output changes. We can add another option
to <code>XvcOutputLine</code>, like <code>XvcOutputLine::Json(T: Serialize)</code>, that will output
the type using <code>serde_json</code>. It won’t introduce any other type.</p>
<p>🐢 Yes, but what will <code>T</code> be? Although store structures have Serde
implementations, they are not particularly useful for this.</p>
<p>🐇 We can have output types, named like <code>XvcFileListOutput</code>, that will be
converted to strings with serialization.</p>
<p>🦊 We do something similar in <code>xvc pipeline export</code> and <code>import</code> commands. We
use
<a href="https://github.com/iesahin/xvc/blob/main/pipeline/src/pipeline/schema.rs#L41"><code>XvcPipelineSchema</code></a>
and <code>XvcStepSchema</code> just for the import and export commands. We’ll write similar
structs for all JSON output and will use Serde to convert these to strings.</p>
<p>🐢 Unlike the <code>import</code> and <code>export</code> commands, we have optional fields in the
output, though. I don’t want content digests to appear in JSON output if they
are not required.</p>
<p>🐇 Let’s search for optional fields in Serde.</p>
<p>🦊 There is a <a href="https://docs.rs/optional-field/latest/optional_field/attr.serde_optional_fields.html">crate for optional
fields</a>.</p>
<p>🐢 We don’t need another crate for this. Serde has the
<a href="https://serde.rs/attr-skip-serializing.html"><code>skip_serializing_if</code></a> attribute
for fields. We can add <code>Option::is_none</code> as a method to these to skip outputting <code>None</code>
fields. All those fields, in this case, will be optional.</p>
<p>🐇 This is fine. We already use structs to format the <code>xvc file list</code> output. We
can just use them to output JSON.</p>]]></content:encoded>
    </item>
    <item>
      <title>devlog 21</title>
      <published>2025-02-03T09:59:28+00:00</published>
      <updated>2025-02-03T09:59:28+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Mon, 03 Feb 2025 09:59:28 +0000</pubDate>
      <link>https://emresahin.net/devlog-21/</link>
      <guid isPermaLink="true">https://emresahin.net/devlog-21/</guid>
      <description>🐢 Now, the next version will have a --json output for xvc file list . We can start working on it or update the Readme file? 🐇 What about adding at least command completions for Nushell? 🐢 Let’s read a bit about clap_complete_nushell . 🦊 There seems to be a nu-complete command. Let’s check its doc...</description>
      <category>xvc</category>
      <category>Development</category>
      <category>Nushell</category>
      <category>clap_complete_nushell</category>
      <category>carapace</category>
      <category>dynamic completions</category>
      <category>JSON</category>
      <category>Lazygit</category>
      <category>Rust</category>
      <category>XVC</category>
      <category>completions</category>
      <content:encoded><![CDATA[<p>🐢 Now, the next version will have a <code>--json</code> output for <code>xvc file list</code>. We
can start working on it or update the Readme file?</p>
<p>🐇 What about adding at least command completions for Nushell?</p>
<p>🐢 Let’s read a bit about <code>clap_complete_nushell</code>.</p>
<p>🦊 There seems to be a <code>nu-complete</code> command. Let’s check its documentation.</p>
<p>🐇 Nothing was found, and Kagi doesn’t help much either.</p>
<p>🐢 There is a completions document for Nushell: https://www.nushell.sh/book/custom_completions.html</p>
<p>🐇 There is a tool called carapace to provide completions across shells.</p>
<p>🐢 Its <a href="https://carapace-sh.github.io/carapace/carapace.html">documentation</a> is
thin, and I’m not sure if it supports dynamic completions out of the box. I
believe instead of adding a carapace setup, we can just write a Nushell
completion script that will use JSON output from the commands and add some
(maybe hidden) utility commands to support it.</p>
<p>🐇 There are a set of example scripts in the Nushell repo:
https://github.com/nushell/nu_scripts/tree/main/custom-completions</p>
<p>🐢 The reason I want to write custom completions for Nushell is that it will be
an exercise for the scripting language. <a href="https://github.com/nushell/nu_scripts/blob/main/custom-completions/gh/gh-completions.nu"><code>gh</code>
completions</a>
are not as scary as a Bash script.</p>
<p>🐇 <a href="https://github.com/nushell/nu_scripts/blob/main/custom-completions/git/git-completions.nu"><code>git</code>
completions</a>
are a better example for XVC. They simply run <code>git</code> whenever necessary. We can
start from a static completions command and update this with dynamic
completions manually. It will teach a lot.</p>
<p>🐢 I <a href="https://github.com/iesahin/nu_scripts">forked</a> the <code>nu_scripts</code> repo and
will add XVC completions script there.</p>
<p>🐇 Then let’s begin by adding Nushell static completions. Shall we add a
command for this?</p>
<p>🦊 Reviving the <code>completion</code> command we removed in 0.6.13?</p>
<p>🐢 We shouldn’t list it. We can make a <code>_comp</code> subcommand for the time being
and generate and distribute completions in the repository. When
<code>clap_complete_nushell</code> has the feature parity to provide dynamic completions,
we can remove these commands.</p>
<p>🐇 What will we use this for other than generating completions?</p>
<p>🐢 Maybe dynamic completions can call this as well.</p>
<p>🐇 Added Nushell static completions to be output using <code>xvc _comp generate-nushell</code>. Let’s bump up the version to 0.6.15.</p>
<pre><code>cargo set-version 0.6.15-alpha.1
   Upgrading xvc from 0.6.14 to 0.6.15-alpha.1
...
</code></pre>
<p>🐢 I noticed we forgot a line in the CLI command handler that asserts <code>xvc_root_opt.is_some()</code>, and this fails when we run <code>xvc</code> outside of repositories. We need to release this version quickly.</p>
<p>🐇 Oops, now, ok, let’s write a static Nushell generator and just release quickly.</p>
<p>🦊 Generating completions with</p>
<pre><code>xvc comp generate-nushell
</code></pre>
<p>🐢 Completion command is run with <code>comp</code> instead of <code>_comp</code>. Should we rename it?</p>
<p>🐇 Renamed it to <code>_comp</code>. It’s not hidden, but at least we can be sure that it won’t be misunderstood as a common command.</p>
<p>🐢 Bumping the version again. Now let’s source the generated script and test it.</p>
<pre><code>cargo set-version 0.6.15-alpha.2
   Upgrading xvc from 0.6.15-alpha.1 to 0.6.15-alpha.2
...
</code></pre>
<p>🦊 Yep, it works. We now have completions for Nushell.</p>
<p>🐢 Let’s update the completions documentation.</p>
<p>🐇 Done. Now, let’s take a look at CI and see what fails.</p>
<pre><code class="language-nu">ghrl | first
╭──────────────┬─────────────────────────────────────────────────────────╮
│ conclusion   │ success                                                 │
│ displayTitle │ Add Nushell completions                                 │
│ headBranch   │ nushell-completions                                     │
│ url          │ https://github.com/iesahin/xvc/actions/runs/13070200765 │
╰──────────────┴─────────────────────────────────────────────────────────╯
</code></pre>
<p>🐢 It fails because of coverage, not the tests. <a href="https://github.com/iesahin/xvc/pull/266#issuecomment-2626794240">Codecov says</a> the new code isn’t tested.</p>
<p>🐇 The added <code>xvc _comp</code> command isn’t tested. We can add a test running those lines and testing if the command outputs a completion script.</p>
<p>🐢 We have a <a href="https://github.com/iesahin/xvc/blob/main/lib/tests/test_completions.rs#L18">test for completions</a>. We can add a test that runs the lines.</p>
<p>🐇 Added a test and bumping up the version.</p>
<pre><code class="language-nu">cargo set-version 0.6.15-alpha.3
   Upgrading xvc from 0.6.15-alpha.2 to 0.6.15-alpha.3
...
</code></pre>
<p>🦊 We can add some more coverage while waiting for the tests.</p>
<p>🐇 <a href="https://app.codecov.io/gh/iesahin/xvc/blob/main/logging%2Fsrc%2Flib.rs#L285"><code>XvcOutputLine</code> implementation</a> seems to have no tests. It’s weird because we use these everywhere.</p>
<p>🐢 I’m not sure we use this particular implementation; we just use <code>XvcOutputLine::Info(s)</code>, not <code>XvcOutputLine::info(s)</code> anywhere. We can delete these methods actually.</p>
<p>🐇 We’ll add JSON output via this particular struct. Can we refactor these to use formatting for JSON, for example? Or use these to output JSON?</p>
<p>🦊 We can add a formatter to <code>XvcOutputLine</code> to output structures.</p>
<p>🐢 The enum is now defined as:</p>
<pre><code class="language-rust">#[derive(Clone, Debug)]
pub enum XvcOutputLine {
    /// The output that we should be reporting to user
    Output(String),
    /// For informational messages
    Info(String),
    /// For debug output to show the internals of Xvc
    Debug(String),
    /// Warnings that are against some usual workflows
    Warn(String),
    /// Errors that interrupts a workflow but may be recoverable
    Error(String),
    /// Panics that interrupts the workflow and ends the program
    /// Note that this doesn't call panic! automatically
    Panic(String),
    /// Progress bar ticks.
    /// Self::Info is also used for Tick(1)
    Tick(usize),
}</code></pre>
<p>Here, these fields can also have a <code>formatter</code> that will render the string in a particular format. For example, the output can be</p>
<pre><code class="language-rust">XvcOutputLine::Output(XvcJsonFormatter, String)</code></pre>
<p>🐇 I’m not sure this is a good idea. <code>Output</code> already specifies this string as output. We can have a wrapper instead, like,</p>
<pre><code class="language-rust">struct XvcJsonOutput(Format&lt;XvcStructuredOutput&gt;, XvcStructuredOutput)</code></pre>
<p>and we can use the supplied format to render <code>XvcStructuredOutput</code> to an output line with <code>XvcOutputLine::Output</code>. If we don’t provide output as structured, it will be too much error-prone work to convert the current outputs to structured.</p>
<p>🦊 The transition will also be gradual. We may not need structured output for most of the commands. We can start with <code>xvc file list</code> and convert others as we go.</p>
<p>🐢 This is sensible. By the way, coverage still didn’t increase. There may be something going on with Codecov or running the test.</p>
<pre><code class="language-nu">ghrl | first
╭──────────────┬─────────────────────────────────────────────────────────╮
│ conclusion   │ success                                                 │
│ displayTitle │ Add Nushell completions                                 │
│ headBranch   │ nushell-completions                                     │
│ url          │ https://github.com/iesahin/xvc/actions/runs/13087431038 │
╰──────────────┴─────────────────────────────────────────────────────────╯
</code></pre>
<p>🐇 Let’s run the test:</p>
<pre><code class="language-sh">cargo test -p xvc --test test_completions
...
test test_completions ... ok

test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.80s
</code></pre>
<p>🐢 Can we make sure the output is a Nushell script and not an error message?</p>
<p>🐇 Let’s print it out.</p>
<p>🐢 It looks like when the <code>COMPLETE</code> environment variable is set, it never calls <code>_comp</code> subcommand and never calls those lines.</p>
<pre><code>cargo set-version 0.6.15-alpha.4
   Upgrading xvc from 0.6.15-alpha.3 to 0.6.15-alpha.4
...
</code></pre>
<p>🐢 Let’s make a release for 0.6.15. Coverage is OK now.</p>
<pre><code class="language-nu">cargo set-version 0.6.15
   Upgrading xvc from 0.6.15-alpha.4 to 0.6.15
...
</code></pre>
<pre><code class="language-nu">gh pr merge --squash --body $"(open CHANGELOG.md | lines | skip 2 | take 5)" --subject "Add static nushell completions"
</code></pre>
<p>🐇 Merged the PR.</p>
<p>🐢 Releases should appear in a few minutes.</p>
<p>🐇 We need to tag the merge commit for this.</p>
<p>🐢 Oh, yep. AFAIK Lazygit doesn’t have something for <code>git push --tags</code>. Let’s push from the CLI.</p>
<pre><code class="language-nu">git push --tags
You are on the main branch. Skipping CHANGELOG.md check.
To github.com:iesahin/xvc
 * [new tag]         v0.6.15 -&gt; v0.6.15
</code></pre>
<p>🦊 These commands, especially tables, are not rendered correctly on the web. We need to change the theme, I think.</p>]]></content:encoded>
    </item>
    <item>
      <title>devlog 9</title>
      <published>2025-01-04T13:03:16+00:00</published>
      <updated>2025-01-04T13:03:16+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Sat, 04 Jan 2025 13:03:16 +0000</pubDate>
      <link>https://emresahin.net/devlog-9/</link>
      <guid isPermaLink="true">https://emresahin.net/devlog-9/</guid>
      <description>🐢 Today I’m planning to add cross-compilation to the Xvc 0.6.13 branch to provide more platform support. 🐇 It looks like you first need to turn off this ghost text from the blink output. It makes writing insufferable. 🐢 Yep, let’s do that first. 🐇 Now, let’s restart Neovim. 🐢 I don’t know why bli...</description>
      <category>devlog</category>
      <category>Software Development</category>
      <category>blink-cmp</category>
      <category>github-cli</category>
      <category>github-actions</category>
      <category>xvc</category>
      <category>reflinks</category>
      <category>cross-compilation</category>
      <category>rust</category>
      <category>ci-cd</category>
      <content:encoded><![CDATA[<p>🐢 Today I’m planning to add cross-compilation to the Xvc 0.6.13 branch to provide more platform support.</p>
<p>🐇 It looks like you first need to turn off this ghost text from the <code>blink</code> output. It makes writing insufferable.</p>
<p>🐢 Yep, let’s do that first.</p>
<p>🐇 Now, let’s restart Neovim.</p>
<p>🐢 I don’t know why <code>blink.cmp</code> doesn’t prioritize the emojis I use. Maybe we can just use <code>#tor</code> and <code>#rab</code> for ourselves.</p>
<p>🐇 I think over time it will learn that the emojis I defined in the snippets file should have higher priority, but let’s skip this for now. What do we need to do to add cross-compilation?</p>
<p>🐢 Maybe we can just make the completion menu wait a bit longer. It shows up almost instantly, and I want it to wait for a few more milliseconds.</p>
<p>🐇 Okay, let’s look at the config.</p>
<p>🐢 The configuration file doesn’t seem to have a key for this. Let’s search: <code>blink.cmp</code>.</p>
<p>🐇 I think the culprit is typo resistance; it causes better options to be pushed down: <a href="https://cmp.saghen.dev/configuration/reference#fuzzy">https://cmp.saghen.dev/configuration/reference#fuzzy</a></p>
<p>🐢 Let’s turn that off.</p>
<p>🐇 There is an error in the configuration file. I don’t know why it’s failing.</p>
<p>🐢 I added emojis to Espanso and checked the error message. The configuration I copied from their docs seems to be broken. The message is:</p>
<pre><code class="language-text">...share/nvim/lazy/blink.cmp/lua/blink/cmp/config/utils.lua:14: fuzzy.max_items: unexpected field found in configuration
</code></pre>
<p>I’ll just delete that line.</p>
<p>🐇 It still doesn’t prioritize snippets, but we can look into this later. Espanso seems to be a better tool for this anyway.</p>
<p>🐢 Yep. Let’s look into cross-compilation support for Rust.</p>
<p>🐇 The well-known option is <code>cross.rs</code>: <a href="https://github.com/cross-rs/cross">https://github.com/cross-rs/cross</a></p>
<p>🐢 We can start with that. Installation is from the Git repository:</p>
<pre><code class="language-bash">$ cargo install cross --git https://github.com/cross-rs/cross
    Updating git repository `https://github.com/cross-rs/cross`
    Updating git submodule `https://github.com/cross-rs/cross-toolchains.git`
  Installing cross v0.2.5 (https://github.com/cross-rs/cross#4090beca)
...
   Installed package `cross v0.2.5 (https://github.com/cross-rs/cross#4090beca)` (executables `cross`, `cross-util`)
</code></pre>
<p>🐇 Now we have the <code>cross</code> and <code>cross-util</code> commands. Cross-compilation requires Podman on Linux or Docker on macOS. Do we have Docker?</p>
<p>🐢 It looks like we don’t. Maybe we can just set up a remote build using Podman or GitHub Actions. I saw a crate for that yesterday; I remember saving it somewhere but can’t find it now. Searching again seems easier, which says something about my archival and retrieval habits.</p>
<p>🐇 Maybe later you can add some vector search capabilities to your archive—semantic search.</p>
<p>🐢 Yep, <em>sometime</em> later.</p>
<p>🐇 Now let’s search for “adding rust cross compilation to github actions.”</p>
<p>🐢 I found a link to the action: <a href="https://github.com/marketplace/actions/build-rust-projects-with-cross">https://github.com/marketplace/actions/build-rust-projects-with-cross</a></p>
<p>Let’s look at the example:</p>
<pre><code class="language-yaml">jobs:
  release:
    name: Release - ${{ matrix.platform.os-name }}
    strategy:
      matrix:
        platform:
          - os-name: FreeBSD-x86_64
            runs-on: ubuntu-20.04
            target: x86_64-unknown-freebsd
            skip_tests: true

          - os-name: Linux-x86_64
            runs-on: ubuntu-20.04
            target: x86_64-unknown-linux-musl

          - os-name: Linux-aarch64
            runs-on: ubuntu-20.04
            target: aarch64-unknown-linux-musl

          - os-name: Linux-riscv64
            runs-on: ubuntu-20.04
            target: riscv64gc-unknown-linux-gnu

          - os-name: Windows-x86_64
            runs-on: windows-latest
            target: x86_64-pc-windows-msvc

          - os-name: macOS-x86_64
            runs-on: macOS-latest
            target: x86_64-apple-darwin

          # more targets here ...

    runs-on: ${{ matrix.platform.runs-on }}
    steps:
      - name: Checkout
        uses: actions/checkout@v3
      - name: Build binary
        uses: houseabsolute/actions-rust-cross@v0
        with:
           command: ${{ matrix.platform.command }}
          target: ${{ matrix.platform.target }}
          args: "--locked --release"
          strip: true
      - name: Publish artifacts and release
        uses: houseabsolute/actions-rust-release@v0
        with:
          executable-name: ubi
          target: ${{ matrix.platform.target }}
</code></pre>
<p>🐇 We can just convert the current configuration to this and see if it works.</p>
<p>🐢 Let’s do that. I created <code>.github/workflows/release.yml</code> and will update it.</p>
<p>🐇 Let’s check the results with <code>gh</code>:</p>
<pre><code class="language-bash">$ gh -R iesahin/xvc run list
completed	failure	v0.6.13	Release	v0.6.13	pull_request	12525070531	34s	2024-12-28T08:17:28Z
</code></pre>
<p>🐢 It says the release has completed.</p>
<pre><code class="language-bash">$ gh -R iesahin/xvc run view 12525070531

X v0.6.13 Release iesahin/xvc#263 · 12525070531
Triggered via pull_request about 3 minutes ago

JOBS
X Release - FreeBSD-x86_64 in 12s (ID 34936257653)
  ✓ Set up job
  ✓ Checkout
  X Build binary
  - Publish artifacts and release
  ✓ Post Build binary
  ✓ Post Checkout
  ✓ Complete job
...
</code></pre>
<p>Now let’s look at the failure:</p>
<pre><code class="language-bash">$ gh -R iesahin/xvc run view 12525070531 --log-failed
</code></pre>
<p>🐇 Change the order and see what happens. Maybe we should first succeed in a non-cross-compilation build.</p>
<pre><code class="language-bash">$ gh -R iesahin/xvc run list
completed	failure	v0.6.13	Release	v0.6.13	pull_request	12525178664	35s	2024-12-28T08:34:18Z
...
$ gh -R iesahin/xvc run view   12525178664 --log-failed
...
</code></pre>
<p>🐢 It’s the same error. The <a href="https://github.com/houseabsolute/ubi/blob/master/.github/workflows/ci.yml">usage example</a> is actually much more sophisticated.</p>
<p>🐇 The issue is that we’re asking for <code>command</code> from the matrix, but the matrix doesn’t define it. That’s the second time today an example from the documentation has failed. I’ve added <code>command</code> for each platform. We can also add separate features this way to make platform-specific functionality work.</p>
<p>🐢 Agreed. Let’s look at it once more.</p>
<pre><code class="language-bash">$ gh -R iesahin/xvc run list | rg Release | head -n 1
completed	failure	v0.6.13	Release	v0.6.13	pull_request	12525253982	41s	2024-12-28T08:46:23Z

$ gh -R iesahin/xvc run view 12525253982 --log-failed
...
Release - macOS-x86_64	Build binary	2024-12-28T08:46:44.8464010Z  [1m [31merror [0m [1m: [0m the lock file /Users/runner/work/xvc/xvc/Cargo.lock needs to be updated but --locked was passed to prevent this
</code></pre>
<p>🐇 Ah, that’s a different error. Let’s remove the <code>--locked</code> flag and retry.</p>
<pre><code class="language-bash">$ gh -R iesahin/xvc run list | rg Release | head -n 1
completed	failure	v0.6.13	Release	v0.6.13	pull_request	12525289176	1m8s	2024-12-28T08:53:23Z
</code></pre>
<p>🐢 We’re finally starting to get some good news. Now we’re getting OpenSSL errors. We need feature flags for these platforms or to specify where OpenSSL is. Let’s add <code>bundled-openssl</code> to the failed ones. We also need <code>bundled-sqlite</code> for Windows binaries.</p>
<pre><code class="language-bash">gh -R iesahin/xvc run list | rg Release | head -n 1
completed	failure	v0.6.13	Release	v0.6.13	pull_request	12525354179	3m9s	2024-12-28T09:04:55Z
</code></pre>
<p>🐇 It looks like the <code>Changes.md</code> file is missing, and it can’t upload the binaries as a release because of this. I’ll set the changes file to <code>CHANGELOG.md</code>.</p>
<p>🐢 It’s weird to fail because of that. I think we should report these errors and possibly send a PR to make that file optional.</p>
<pre><code class="language-bash">gh -R iesahin/xvc run list | rg Release | head -n 1
in_progress		v0.6.13	Release	v0.6.13	pull_request	12525449368	1m38s	2024-12-28T09:18:40Z
...
✓ Release - macOS-x86_64 in 2m39s (ID 34937017695)
✓ Release - macOS-aarch64 in 2m42s (ID 34937017787)
..
ARTIFACTS
xvc-macOS-x86_64.tar.gz
xvc-macOS-arm64.tar.gz
</code></pre>
<p>🐢 It looks like <code>Linux-riscv64</code> has OpenSSL compilation errors. I think we can skip this platform for now. The goal was to add <code>aarch64</code> for macOS, and that seems to have succeeded.</p>
<pre><code class="language-bash">gh -R iesahin/xvc run list | rg Release | head -n 1
completed	failure	v0.6.13	Release	v0.6.13	pull_request	12525504079	3m24s	2024-12-28T09:26:43Z
...
X Release - Linux-x86_64	Build binary	2024-12-28T09:29:16.2262518Z  [0m [1m [38;5;9merror[E0308] [0m [0m [1m: mismatched types [0m
...
</code></pre>
<p>🐢 It looks like <code>Linux-x86_64</code> doesn’t support reflinks. Maybe we can make reflinks an optional feature and add it specifically to macOS and Windows targets.</p>
<p>🐇 I’ve removed <code>reflink</code> from the default features. This is a breaking change, but <em>fortunately</em> we don’t have many users who will be affected by it.</p>
<p>Let’s check the results once more.</p>
<pre><code class="language-bash">gh -R iesahin/xvc run list | rg Release | head -n 1
completed	failure	v0.6.13	Release	v0.6.13	pull_request	12525597866	3m39s	2024-12-28T09:42:44Z
</code></pre>
<p>🐢 Now turn off the <code>NetBSD</code> target as well.</p>
<p>🐇 <code>FreeBSD</code> and <code>Linux-x86_64</code> targets are building. Let’s see why the ARM Linux targets are failing.</p>
<pre><code class="language-bash">gh -R iesahin/xvc run list | rg Release | head -n 1
completed	failure	v0.6.13	Release	v0.6.13	pull_request	12525643284	4m46s	2024-12-28T09:52:04Z
</code></pre>
<p>🐇 It looks like those targets don’t have <code>libsqlite3</code> installed. Let’s add <code>bundled-sqlite</code> to these targets too.</p>
<pre><code class="language-bash">gh -R iesahin/xvc run list | rg Release | head -n 1
completed	failure	v0.6.13	Release	v0.6.13	pull_request	12525734325	4m11s	2024-12-28T10:07:36Z
</code></pre>
<p>🐢 Adding Android as a target didn’t work. Let’s remove it for now; it seems to require more work.</p>
<p>🐇 I think we’ll eventually move to using Xvc to distribute binaries. We could build the Android binary on Termux and link it on the releases page or push it as a release artifact.</p>
<p>🐢 I need to learn more about GitHub releases. If we can add artifacts to the release, maybe we can do some of this work locally.</p>
<p>🐇 We can start by looking at the capabilities of <code>gh</code> commands.</p>
<p>🐢 It looks like it’s possible to upload assets to releases. Let’s check how that works.</p>
<pre><code class="language-bash">gh release upload --help
</code></pre>
<p>🐇 So basically, we can just upload files to tags. We can list releases and work with them like anything else.</p>
<pre><code class="language-bash">gh -R iesahin/xvc release list
</code></pre>
<p>And we can delete releases:</p>
<pre><code class="language-bash">for r in v0.4.2-alpha.8 v0.4.2-alpha.7 v0.4.2-alpha.6 v0.4.2-alpha.5 v0.4.2-alpha.0 v0.4.1-alpha.0; do
  gh -R iesahin/xvc release delete "${r}"
done
</code></pre>
<p>🐢 Now let’s list them again.</p>
<pre><code class="language-bash">gh -R iesahin/xvc release list
</code></pre>
<p>🐇 The latest one has failed again.</p>
<pre><code class="language-bash">X Failed to CreateArtifact: Received non-retryable error: Failed request: (409) Conflict: an artifact with this name already exists on the workflow run
</code></pre>
<p>🐢 It looks like we’re coming to the end of this session. Under what conditions will we create a release?</p>
<p>🐇 I think it’s better to release only non-alpha tags.</p>
<p>🐢 Then the rule in the workflow will be something like:</p>
<pre><code class="language-yaml">on:
  workflow_dispatch:
  push:
    tags:
      - "v*.*.*"
      - "!v.*.*-alpha.*"
</code></pre>
<p>🐇 And now we have this result:</p>
<pre><code class="language-bash">✓ v0.6.13 Release iesahin/xvc#263 · 12533498361
Triggered via pull_request about 21 minutes ago

JOBS
✓ Release - FreeBSD-x86_64 in 4m14s
✓ Release - Linux-x86_64 in 4m3s
✓ Release - Linux-aarch64 in 4m50s
✓ Release - Windows-x86_64 in 8m47s
✓ Release - Windows-aarch64 in 6m42s
✓ Release - macOS-x86_64 in 2m37s
✓ Release - macOS-aarch64 in 2m36s
</code></pre>
<p>🐢 Nice! Let’s check the release list.</p>
<pre><code class="language-bash">gh -R iesahin/xvc release list
</code></pre>
<p>🐇 Why is there no “latest” release?</p>
<p>🐢 We have to tag it first.</p>
<p>🐇 Ah, right. Let’s tag it then.</p>
<p>🐢 I’ve pushed the changes and tagged them with <code>v0.6.13-alpha.5</code>.</p>
<p>🐇 I think it’s possible to make a release today.</p>
<p>🐢 It looks like it, yes.</p>
<p>🐇 We can merge the PR and tag it. Then everything should work.</p>
<p>🐢 We need some cleanup in the YAML files, though.</p>
<p>🐇 We can leave that for the next release.</p>
<pre><code class="language-bash">gh -R iesahin/xvc run list
in_progress		v0.6.13	Rust-CI	v0.6.13	pull_request	12533690474	9m18s	2024-12-29T08:02:29Z
completed	success	Release	Release	v0.6.13-alpha.5	push	12533689099	9m11s	2024-12-29T08:02:19Z
</code></pre>
<p>🐢 Now we have another failure in the regular CI. Let’s look at it.</p>
<p>🐇 The issue seems to be in the doc tests:</p>
<pre><code class="language-diff">- Total #: 8 Workspace Size:         276 Cached Size:          19
+ Total #: 8 Workspace Size:         278 Cached Size:          19
</code></pre>
<p>🐢 These tests are brittle, but they provide valuable information. Let’s fix it and push again.</p>
<p>🐇 Done. We can also remove some of the watches that produce so many logs.</p>
<p>🐢 I’m a bit ambivalent about them. I thought we could use these watches when debugging, but experience has shown that we need more granular watches during debugging and almost never use these otherwise. Let’s remove some of them.</p>
<p>🐇 We can increase the output for certain commands, but the tracing output doesn’t help much in regular runs. If Xvc gets popular enough that we can’t cope with bug reports, we can always add more watches.</p>
<p>🐢 Another option is to exclude the watch code from the release build, but that won’t change anything for our debug cycles.</p>
<p>🐇 I think removing them is a fair trial. We can always put them back when debugging.</p>
<p>🐢 Watches could also produce regular output instead of tracing. That way we won’t forget to remove them.</p>
<p>🐇 Ah, yep, that’s a good option too.</p>
<p>🐢 We can have a <code>trace!</code> macro similar to the current one for user consumption, and a <code>watch!</code> macro that sends output to <code>stderr</code>.</p>
<p>🐇 Good idea. Let’s do that in the next release.</p>
<p>🐢 Let’s check the tests before pushing this cleanup.</p>
<pre><code class="language-bash">gh -R iesahin/xvc run list
completed	success	v0.6.13	Rust-CI	v0.6.13	pull_request	12533993607	13m49s	2024-12-29T08:47:50Z
</code></pre>
<p>🐇 CI succeeded, and the release didn’t run. Let’s do some more cleanup.</p>
<pre><code class="language-bash">gh -R iesahin/xvc run list
in_progress		v0.6.13	Rust-CI	v0.6.13	pull_request	12534416152	43s	2024-12-29T09:57:02Z
</code></pre>]]></content:encoded>
    </item>
    <item>
      <title>devlog 8</title>
      <published>2025-01-04T12:41:16+00:00</published>
      <updated>2025-01-04T12:41:16+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Sat, 04 Jan 2025 12:41:16 +0000</pubDate>
      <link>https://emresahin.net/devlog-8/</link>
      <guid isPermaLink="true">https://emresahin.net/devlog-8/</guid>
      <description>🐢 The only failure was the patch coverage in Xvc. Let’s see what needs to be done. 🐇 It looks, from the coverage page , that our additions to HStore don’t have any tests. Maybe we can add some unit tests to the new joins. 🐢 I don’t find unit tests particularly useful, but let’s use GitHub Copilot...</description>
      <category>devlog</category>
      <category>Software Development</category>
      <category>xvc</category>
      <category>rust</category>
      <category>coverage</category>
      <category>cargo-publish</category>
      <category>github-copilot</category>
      <category>claude</category>
      <category>dufs</category>
      <category>neovim</category>
      <category>testing</category>
      <content:encoded><![CDATA[<p>🐢 The only failure was the patch coverage in Xvc. Let’s see what needs to be done.</p>
<p>🐇 It looks, from the <a href="https://app.codecov.io/gh/iesahin/xvc/pull/263?src=pr&amp;el=tree&amp;utm_medium=referral&amp;utm_source=github&amp;utm_content=comment&amp;utm_campaign=pr+comments&amp;utm_term=Emre+Sahin">coverage page</a>, that our additions to HStore don’t have any tests. Maybe we can add some unit tests to the new joins.</p>
<p>🐢 I don’t find unit tests particularly useful, but let’s use GitHub Copilot to add some for us.</p>
<p>🐇 I added a unit test and a doc test for <code>full_join</code>. I think doc tests have more value; they provide documentation, and we can readily see how to use a function from its docs. It’s better to increase coverage with doc tests.</p>
<p>🐢 There are things I should keep in mind while writing doc tests. Imports must use the full path, not <code>crate</code>. Also, the tested struct doesn’t have implicit imports.</p>
<p>🐇 The ceremony for adding keys and values is a bit too much. It might be worthwhile to add an <code>insert</code> method for anything that implements <code>Into&lt;XvcEntity&gt;</code>.</p>
<p>🐢 That would certainly save time—no more typing <code>.into()</code> for each key! :)</p>
<p>🐇 Pushed to test again. Should we have some way to test coverage locally?</p>
<p>🐢 I don’t think we need to check coverage locally. It’s not worth our time right now.</p>
<p>🐇 Now, while waiting for the tests to complete, what else can we do?</p>
<pre><code class="language-bash">gh -R iesahin/xvc run list
completed failure v0.6.13 Rust-CI v0.6.13 pull_request 12543831360 3m54s 2024-12-30T08:15:06Z
...
</code></pre>
<p>🐢 That didn’t take too long. Let’s view the results:</p>
<pre><code class="language-bash">gh -R iesahin/xvc run view 12543831360
...
  X Run Current Dev Tests
...
To see what failed, try: 
View this run on GitHub: https://github.com/iesahin/xvc/actions/runs/12543831360
</code></pre>
<p>🐢 The current dev tests are failing for some reason. Let’s run them locally.</p>
<p>🐇 We’re missing <code>llvm-tools-preview</code> locally. How do I install this?</p>
<p>🐢 The command is <code>rustup component add llvm-tools-preview</code>:</p>
<pre><code class="language-text">info: component 'llvm-tools' for target 'aarch64-apple-darwin' is up to date
</code></pre>
<p>🐇 It’s already installed. We just need to set the environment variables.</p>
<p>🐢 Instead, we can just turn off dev tests for the time being. We don’t really need them; our local tests pass.</p>
<p>🐇 Yeah, okay. We don’t need to solve every single bit of these issues right now.</p>
<pre><code class="language-bash">gh -R iesahin/xvc run list
...
</code></pre>
<p>🐇 Okay, let’s take a look at the run again.</p>
<pre><code class="language-bash">gh -R iesahin/xvc run list
completed failure v0.6.13 Rust-CI v0.6.13 pull_request 12544020064 3m39s 2024-12-30T08:33:25Z

gh -R iesahin/xvc run view 12544020064
...
  X Test and Coverage
...

gh run view 12544020064 --log-failed
...
Test and Coverage (stable) Test and Coverage 2024-12-30T08:36:58.8911690Z Error: ProcessError { stdout: "", stderr: "  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current\n                                 Dload  Upload   Total   Spent    Left  Speed\n\r  0     0    0     0    0     0      0      0 --:--:-- --:--:-- --:--:--     0\r  0     0    0     0    0     0      0      0 --:--:-- --:--:-- --:--:--     0\ncurl: (7) Failed to connect to e1.xvc.dev port 80 after 160 ms: Couldn't connect to server\n" }
...
</code></pre>
<p>🐢 We need to start Nginx on the server. We forgot to do that yesterday.</p>
<p>🐇 Ah, right. After that <code>dufs</code> installation. Okay.</p>
<p>🐢 We also need to add a reverse proxy for <code>dufs</code> somehow, but that’s for later.</p>
<p>🐇 For this use case, I don’t think it’s necessary. We can just adjust the port to a non-standard one if we need 443 for something else. Let’s look at the tests again.</p>
<p>🐢 Let’s add another doc test, this time to <code>XvcStore</code>.</p>
<pre><code class="language-bash">cargo test -p xvc-ecs --doc
...
test result: ok. 8 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 2.87s
</code></pre>
<p>🐇 Sent the files again. Waiting for the tests to finish.</p>
<p>🐢 Let’s check the keymaps file in the meantime.</p>
<p>🐇 I tried reading the documentation but didn’t see an error. Maybe we should just set it to non-lazy and disallow remaps.</p>
<p>🐢 We’ve already spent too much time on this.</p>
<p>🐇 Yes, let’s check the tests again.</p>
<pre><code class="language-bash">gh -R iesahin/xvc run list
completed success v0.6.13 Rust-CI v0.6.13 pull_request 12544291068 8m10s 2024-12-30T09:00:42Z
...
</code></pre>
<p>🐢 Oh, nice, the merge is ready.</p>
<p>🐇 Patch coverage is still behind the target, though.</p>
<p>🐢 Yeah, but let’s release this one and ensure the next one is better covered. Also, I’m not sure if the doc tests actually affected the coverage report.</p>
<p>🐇 If we look at the <a href="https://app.codecov.io/gh/iesahin/xvc/pull/263?src=pr&amp;el=tree&amp;utm_medium=referral&amp;utm_source=github&amp;utm_content=comment&amp;utm_campaign=pr+comments&amp;utm_term=Emre+Sahin">coverage page</a> again, we can see.</p>
<p>🐢 It seems codecov.io doesn’t consider coverage for doc tests. That’s a bit weird, but let’s not spend more time on it.</p>
<p>🐇 Sure, let’s merge.</p>
<p>🐢 I think we forgot to bump the versions in the <code>Cargo.toml</code> files. We’ll have to do that in <code>main</code>.</p>
<p>🐇 Oh, yeah. Let’s bump them and tag the release as well.</p>
<p>🐢 Now we can wait for all the files to be produced. What’s next?</p>
<p>🐇 We can release the Python version too. It shouldn’t need any changes.</p>
<p>🐢 Right. Maybe we can add a few tests there as well.</p>
<p>🐇 Let’s bump the version first and see.</p>
<p>🐢 I’ve bumped the versions in <code>Cargo.toml</code> and run <code>maturin develop</code>.</p>
<p>🐇 It seems ready now.</p>
<p>🐢 The main branch is failing, though. The “Publish crates” action is looking for <code>libsqlite3</code>. Let’s investigate.</p>
<pre><code class="language-bash">gh -R iesahin/xvc run list
completed failure Release v0.6.13 Publish Crates v0.6.13 push 12544753359 6m1s 2024-12-30T09:41:56Z
</code></pre>
<p>🐢 The issue is that the VM doesn’t have <code>libsqlite3-dev</code>. Let’s add it.</p>
<p>🐇 We need to restart the job manually. Let’s skip tagging this time.</p>
<p>🐢 Some of the packages were already published, and now they’re causing failures because crates.io says they already exist. Maybe we can check if a package is already published before trying.</p>
<p>🐇 Let’s see if we can make <code>cargo publish</code> more forgiving.</p>
<p>🐢 There doesn’t seem to be an easy option. Let’s search for “how to skip published packages in workspace to avoid errors with cargo publish.”</p>
<p>🐇 Claude is hallucinating again. Let’s try a manual approach: how to skip already published packages?</p>
<p>🐢 It might be easier to just add a check. How do we get that info?</p>
<p>🐇 Or we can just move on to the next package if one is already available.</p>
<p>🐢 Let’s push the missing packages manually this time.</p>
<p>🐢 We should have a key to open garden files quickly. What does <code>Fzf-Lua files</code> receive as arguments?</p>
<p>🐇 Let’s check the help page: <code>fzf-lua</code>.</p>
<p>🐢 Before that, maybe we can try to fix why selected lines are not searched in Visual-Line mode?</p>
<p>🐇 Yep, let’s fix our search first.</p>
<p>🐢 How do you set a key in visual line mode in Neovim Lua?</p>
<p>🐇 The abbreviation is <code>V</code>. Let’s try that.</p>
<p>🐢 Looks like we need to restart the session.</p>
<p>🐇 It’s still not working. Let’s check <code>:map</code>.</p>
<p>🐢 It seems it needs more care; I’ll check it later.</p>
<p>👨🏾‍🦲
🐢 Bence kendini biraz daha anlamlı bir işle uğraştırmalısın.
🐇 Ne gibi?
🐢 Belki biraz daha yayın yapmalısın, biraz daha işe başvurmalısın.
🐇 “Meli”, “malı” ile geçiyor ömrümüz.</p>
<p>⌚
🐢 It looks like we can move most of the daily template to links or commands. We can call it <code>ref/daily</code>. My idea is to fill the page intentionally, without templates.</p>
<p>🐇 It might be better to start with a blank page, yeah. The template makes me a bit nervous. There are too many things to fill in, and most of them aren’t things I like being pushed into doing.</p>
<p>🐢 Let’s start by moving the daily templates to <code>ref/daily</code>. No more daily template.</p>
<p>🐇 Now we can delete the rest of this page. We’ll use the daily links page and maybe have reminders at the end of these sessions.</p>
<p>👨🏽‍⚕️</p>]]></content:encoded>
    </item>
    <item>
      <title>More features for my Telegram bot</title>
      <published>2024-12-13T16:52:45+00:00</published>
      <updated>2024-12-13T16:52:45+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Fri, 13 Dec 2024 16:52:45 +0000</pubDate>
      <link>https://emresahin.net/more-features-for-my-telegram-bot/</link>
      <guid isPermaLink="true">https://emresahin.net/more-features-for-my-telegram-bot/</guid>
      <description>I have a Telegram bot to save files to my zettelkasten. It’s always there, waiting for my muses to visit. Currently, it only supports text. When I send a message, the bot adds it to a new file or appends it to a file modified within the last 30 minutes, which simplifies multi-paragraph inputs. I ...</description>
      <category>Automation</category>
      <category>Development</category>
      <category>Inboxbot</category>
      <category>Xvc</category>
      <category>Telegram</category>
      <category>Mime-types</category>
      <category>Zettelkasten</category>
      <category>Bots</category>
      <content:encoded><![CDATA[<p>I have a <a href="https://github.com/iesahin/inboxbot">Telegram bot</a> to save files to my zettelkasten. It’s always there, waiting for my muses to visit.</p>
<p>Currently, it only supports text. When I send a message, the bot adds it to a new file or appends it to a file modified within the last 30 minutes, which simplifies multi-paragraph inputs.</p>
<p>I need some more features to make it more useful.</p>
<p>I want it to save audio and image files to my inbox as well. I plan to run <a href="https://github.com/iesahin/xvc">Xvc</a> pipelines on these files to convert audio notes to text and make images searchable via OCR or tagging.</p>
<p>The first step to achieve this is to understand what kind of files Telegram can upload and download. I want to track these files with Xvc to avoid bloating the Git repository. I have a script that processes these files, but currently, it only commits changes to Git. I want to determine whether a file should be tracked by Xvc or Git.</p>
<p>I skimmed the <a href="https://core.telegram.org/bots/api#sending-files">Telegram API documentation</a> and found that it supports a wide range of file types. This led me to change my approach: I will track only <code>.txt</code> and <code>.md</code> files with Git and use Xvc for everything else.</p>
<p>In the future, I plan to add a <code>--binary-only</code> option to <code>xvc file track</code> to track binary files only. This will help me to track only non-text files when using globs as targets. I might also extend this to a general predicate to decide which files to track based on their size, modification time, or other properties.</p>]]></content:encoded>
    </item>
    <item>
      <title>devlog 7</title>
      <published>2024-08-06T06:43:01+00:00</published>
      <updated>2024-08-06T06:43:01+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Tue, 06 Aug 2024 06:43:01 +0000</pubDate>
      <link>https://emresahin.net/devlog-7/</link>
      <guid isPermaLink="true">https://emresahin.net/devlog-7/</guid>
      <description>I noticed that I often forget to update the Xvc CHANGELOG. To fix this, I added a pre-push hook that checks the files I’m pushing. If the CHANGELOG is not among them and there are changes to Rust files, it prevents the push. I hope this will remind me to update the logs more frequently. #!/bin/ba...</description>
      <category>devlog</category>
      <category>Software Development</category>
      <category>Git</category>
      <category>git-hooks</category>
      <category>automation</category>
      <category>changelog</category>
      <category>xvc</category>
      <category>bash</category>
      <content:encoded><![CDATA[<p>I noticed that I often forget to update the Xvc CHANGELOG. To fix this, I added
a pre-push hook that checks the files I’m pushing. If the CHANGELOG is not among
them and there are changes to Rust files, it prevents the push. I hope this will
remind me to update the logs more frequently.</p>
<pre><code class="language-bash">#!/bin/bash

# Git pre-push hook to check if CHANGELOG.md is included in the push and if the branch is develop

# Get the current branch name

current_branch=$(git rev-parse --abbrev-ref HEAD)

# Check if the branch is develop

if [ "$current_branch" == "main" ]; then
	echo "You are on the main branch. Skipping CHANGELOG.md check."
	exit 0
fi

# remote="$1"
# url="$2"

# Get the list of commits to be pushed
commits=$(git rev-list '@{u}..HEAD')

has_rust_files=$(false)

# TODO: We can iterate to get file list only once
# Get the list of files that are going to be pushed
for commit in $commits; do
	if git diff-tree --no-commit-id --name-only -r "${commit}" | grep -q "\\.rs$"; then
		has_rust_files=$(true)
	fi
done

if [[ ! $has_rust_files ]]; then
	echo "No .rs files in the push, no need to check CHANGELOG"
	exit 0
fi

# Check if CHANGELOG.md is among the files in the commits
for commit in $commits; do
	if git diff-tree --no-commit-id --name-only -r "${commit}" | grep -q "CHANGELOG.md"; then
		echo "CHANGELOG.md is included in the push."
		exit 0
	fi
done

echo "ERROR: CHANGELOG.md is not included in the push."
exit 1
</code></pre>
<p><strong>Edit (2025-02-08):</strong> Added a check for when only code files are changed.</p>]]></content:encoded>
    </item>
    <item>
      <title>Differences between DVC and Xvc</title>
      <published>2024-07-17T09:57:44+00:00</published>
      <updated>2024-07-17T09:57:44+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Wed, 17 Jul 2024 09:57:44 +0000</pubDate>
      <link>https://emresahin.net/differences-between-dvc-and-xvc/</link>
      <guid isPermaLink="true">https://emresahin.net/differences-between-dvc-and-xvc/</guid>
      <description>I wrote this on Reddit; let’s put it here too. Full list of similarities and differences is rather long. Let me summarize it. Xvc has different commands; xvc file track is used instead of dvc add . Xvc doesn’t add files (like .dvc files) to your repository and keeps all metadata tracking under th...</description>
      <category>xvc</category>
      <category>MLOps</category>
      <category>xvc</category>
      <category>dvc</category>
      <category>mlops</category>
      <category>data-versioning</category>
      <category>rclone</category>
      <category>s5cmd</category>
      <category>data-science</category>
      <content:encoded><![CDATA[<p>I wrote <a href="https://www.reddit.com/r/mlops/comments/1e4vvv3/xvc_a_free_as_in_freedom_cli_and_python_tool_to/">this</a> on Reddit; let’s put it here too.</p>
<p><a href="https://docs.xvc.dev/start/from-dvc">Full list of similarities and differences</a> is rather long. Let me summarize it.</p>
<p>Xvc has different commands; <code>xvc file track</code> is used instead of <code>dvc add</code>. Xvc doesn’t add files (like <code>.dvc</code> files) to your repository and keeps all metadata tracking under the <code>.xvc</code> directory. <a href="https://docs.xvc.dev/ref/xvc-file-recheck"><em>Checkout method</em></a> is per-file, not configured globally, so you can keep track of your data directory with symlinks and your model directory as copies. Xvc uses BLAKE3 as the default hashing algorithm, and you can configure this to be BLAKE2, SHA-2, or SHA-3.</p>
<p>Pipelines are not defined using YAML files. You can write a shell script with <code>xvc pipeline step ...</code> or use Python <code>xvc.pipeline().step().dependency(step_name="preprocess", param="hyperparams.yaml::batch_size")</code> to define pipelines first. Then you can use <code>xvc pipeline export</code> and <code>xvc pipeline import</code> to modify the pipeline in YAML, JSON, or TOML.</p>
<p>There are <a href="https://docs.xvc.dev/ref/xvc-pipeline-step-dependency">more dependency options</a>; e.g., a pipeline step may depend on a text file partially, by <a href="https://docs.xvc.dev/ref/xvc-pipeline-step-dependency#regex-item-dependencies">regex</a> or by the <a href="https://docs.xvc.dev/ref/xvc-pipeline-step-dependency#line-item-dependencies">line</a> options. There is a <a href="https://docs.xvc.dev/ref/xvc-pipeline-step-dependency#generic-command-dependencies">generic dependency</a> option; the output of a shell command can be used as a dependency to a step.</p>
<p><a href="https://docs.xvc.dev/ref/xvc-storage-new">Remote storage options</a> are rather limited for Xvc; local, ssh+rsync, and S3-compatible storages are supported for now. There is also a <code>generic</code> storage option where you can define upload and download commands for the tool you’re using, e.g., <code>rclone</code> or <code>s5cmd</code>, and Xvc can use it. I’ll add Azure and rclone as natively supported storage options eventually, but I don’t like the idea of keeping credentials, so there won’t be any OAuth-required storage options, e.g., Google Drive. (You’ll be able to use these through rclone, though.) All Xvc storages use environment variables for authentication.</p>
<p>Xvc doesn’t have experiment tracking yet. <a href="https://docs.xvc.dev/how-to/git-branches">You can use <code>--from-ref</code> and <code>--to-branch</code> options</a> to store artifacts from the pipeline to different branches. I’ll add features to run pipelines and commands quickly and compare these eventually (I need one too), but it may take some time.</p>
<p>Xvc doesn’t track anything about the user. It shouldn’t make any network connections if you don’t specifically ask it to do so. I’m planning to add binaries that only do file operations or pipeline operations, so if someone doesn’t need pipeline features, they will simply use <code>xvc-file</code>.</p>]]></content:encoded>
    </item>
    <item>
      <title>devlog 6</title>
      <published>2024-07-12T08:53:14+00:00</published>
      <updated>2024-07-12T08:53:14+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Fri, 12 Jul 2024 08:53:14 +0000</pubDate>
      <link>https://emresahin.net/devlog-6/</link>
      <guid isPermaLink="true">https://emresahin.net/devlog-6/</guid>
      <description>I have a habit of testing against the CLI’s help string output. It allows me to keep the documentation up to date and makes me aware of any undocumented options. When new features are added, the help text changes and the test fails, which prompts me to add those options to the documentation. I tr...</description>
      <category>devlog</category>
      <category>Software Development</category>
      <category>xvc</category>
      <category>Python</category>
      <category>pytest</category>
      <category>clap</category>
      <category>CLI</category>
      <category>Testing</category>
      <content:encoded><![CDATA[<p>I have a habit of testing against the CLI’s help string output. It allows me to
keep the documentation up to date and makes me aware of any undocumented options.
When new features are added, the help text changes and the test fails, which
prompts me to add those options to the documentation.</p>
<p>I tried the same approach when testing Python bindings with Pytest:</p>
<pre><code class="language-python">def test_pipeline_step_dependency(empty_xvc_repo):
    dep_help = empty_xvc_repo.pipeline().step().dependency(help=True)
    expected = """
Usage: xvc pipeline step dependency [OPTIONS] --step-name &lt;STEP_NAME&gt;

Options:
  -s, --step-name &lt;STEP_NAME&gt;
          Name of the step to add the dependency to
"""
    assert dep_help == expected
</code></pre>
<p>This doesn’t work because the help text is generated by <a href="https://docs.rs/clap/latest/clap/">clap</a> and skips the usual
thread-based output handler. All command output and errors in Xvc are returned
as strings from the command, except for the help text that’s generated by <a href="https://docs.rs/clap/latest/clap/">clap</a>
automatically.</p>
<p>There are probably workarounds for this, but I won’t pursue them further as I don’t
test the actual functionality in the Python bindings anyway.</p>]]></content:encoded>
    </item>
    <item>
      <title>devlog 5</title>
      <published>2024-07-08T20:35:24+00:00</published>
      <updated>2024-07-08T20:35:24+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Mon, 08 Jul 2024 20:35:24 +0000</pubDate>
      <link>https://emresahin.net/devlog-5/</link>
      <guid isPermaLink="true">https://emresahin.net/devlog-5/</guid>
      <description>While writing Xvc tests for Python, I hit an error caused by the ECS single-load protection. The single-loader allows only one instance of Xvc to be run in a single process. This is no problem for the shell, but it looks like it won’t be possible to use multiple Xvc instances in a single Python p...</description>
      <category>devlog</category>
      <category>Xvc</category>
      <category>devlog</category>
      <category>multiprocessing</category>
      <category>Jupyter</category>
      <category>python</category>
      <category>ecs</category>
      <category>testing</category>
      <content:encoded><![CDATA[<p>While writing Xvc tests for Python, I hit an error caused by the ECS single-load protection.</p>
<p>The single-loader allows only one instance of Xvc to be run in a single process. This is no problem for the shell, but it looks like it won’t be possible to use multiple Xvc instances in a single Python process.</p>
<p>It’s possible to overcome this with an elaborate multiprocessing setup in the wrapper, but I won’t bother with it for now.</p>]]></content:encoded>
    </item>
    <item>
      <title>devlog 4</title>
      <published>2024-06-05T10:04:54+00:00</published>
      <updated>2024-06-05T10:04:54+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Wed, 05 Jun 2024 10:04:54 +0000</pubDate>
      <link>https://emresahin.net/devlog-4/</link>
      <guid isPermaLink="true">https://emresahin.net/devlog-4/</guid>
      <description>Let’s start by looking at the debug output issue. We can start by replacing the eprintln! macros with println! , perhaps. I replaced the eprintln! s with println! , but it didn’t make any difference. Maybe we should remove those statements completely. I can’t really find the place that kills the ...</description>
      <category>devlog</category>
      <category>xvc</category>
      <category>xvc storage</category>
      <category>python</category>
      <category>clap</category>
      <category>debug</category>
      <category>rust</category>
      <category>s3</category>
      <category>cli</category>
      <content:encoded><![CDATA[<p>Let’s start by looking at the debug output issue. We can start by replacing the <code>eprintln!</code> macros with <code>println!</code>, perhaps.</p>
<p>I replaced the <code>eprintln!</code>s with <code>println!</code>, but it didn’t make any difference. Maybe we should remove those statements completely.</p>
<p>I can’t really find the place that kills the kernel. The last command to run is <code>xvc file list</code>:</p>
<pre><code class="language-python">print(xvc_test_data.file().list("test-data/dir-0002"))
</code></pre>
<p>and the output it produces is:</p>
<pre><code>[src/output.rs:144:13] &amp;output_str = "SS         131 2024-06-05 08:57:11 41e16be7          test-data/dir-0002/file-0003.bin\nSS         131 2024-06-05 08:57:11 27f0
efd0          test-data/dir-0002/file-0002.bin\nSS         131 2024-06-05 08:57:11 66de5084          test-data/dir-0002/file-0001.bin\nTotal #: 3 Workspace Size:
      393 Cached Size:        6006\n"
</code></pre>
<p><code>print</code> may be causing the crash, but the more likely cause is the command that comes after this:</p>
<pre><code>!ls -l test-data/dir-0001/
</code></pre>
<p>I replaced this with <code>lsd</code>, which also failed. Maybe it’s actually a Python crash or bug.</p>
<p>The way to understand is to create a notebook file with only that cell and try to run it.</p>
<p>The <code>ls</code> line runs fine with a new notebook. It even runs on the <code>README</code> file when run at the beginning. The line that makes the kernel crash is:</p>
<pre><code class="language-python">xvc_test_data.storage().new_s3(name="backup", bucket_name="xvc-test", region="eu-central-1", storage_prefix="xvc-storage")
</code></pre>
<p>We can start by removing the <code>new_s3</code> part.</p>
<p>The <code>storage()</code> method runs fine. It returns an <code>XvcStorage()</code> object, as it should.</p>
<p>When I run <code>storage().list()</code>, it takes a very long time. The bug is likely related to <code>storage()</code>.</p>
<p>It looks like the <code>storage</code> object was adding <code>file</code> instead of <code>storage</code> as a subcommand. I’ve fixed it now.</p>
<p>That was the bug. The <code>README</code> notebook now creates the S3 storage.</p>
<p>What was the reason behind this?</p>
<p>Parsing the CLI to the <code>XvcCLI</code> object was perhaps the culprit. Let’s look at it more clearly.</p>
<p>Let’s try <code>xvc file new s3</code> as a command to see how it behaves.</p>
<p>It says <em>unrecognized subcommand</em> for <code>new</code>.</p>
<p>This is how it should be, but I wonder why it doesn’t work for the <code>XvcCLI</code> parser.</p>
<p>Anyway, it’s already 13:00, so let’s stop here for today.</p>]]></content:encoded>
    </item>
    <item>
      <title>devlog 2</title>
      <published>2024-05-26T17:29:09+00:00</published>
      <updated>2024-05-26T17:29:09+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Sun, 26 May 2024 17:29:09 +0000</pubDate>
      <link>https://emresahin.net/devlog-2/</link>
      <guid isPermaLink="true">https://emresahin.net/devlog-2/</guid>
      <description>We were debugging Python bindings for Xvc. xvc.file.track enters an infinite loop when given a non-existent path. To make debugging easier, we can add file deletions in the ./start-readme script to the notebook itself. Instead, I’ll run my watcher run-after-commit.sh to ensure that the files are ...</description>
      <category>devlog</category>
      <category>xvc</category>
      <category>xvc</category>
      <category>jupyter</category>
      <category>xvc-config</category>
      <category>xvc-file</category>
      <category>repr</category>
      <category>python-bindings</category>
      <category>debugging</category>
      <content:encoded><![CDATA[<p>We were debugging Python bindings for Xvc. <code>xvc.file.track</code> enters an infinite
loop when given a non-existent path.</p>
<p>To make debugging easier, we can add file deletions in the <code>./start-readme</code> script
to the notebook itself.</p>
<p>Instead, I’ll run my watcher <code>run-after-commit.sh</code> to ensure that the files
are deleted after commits.</p>
<p>That may also work; you can use both as well.</p>
<p>I noticed cli-opts pass <code>--no-system-config</code> etc. by default. Let’s deal with
this first.</p>
<p>I fixed that.</p>
<p>I’m testing whether we are in the directory that we should be in with
<code>xvc.root("--absolute")</code>, but it looks like it’s not possible to pass a string
argument to the command.</p>
<p>Let me take a look at this.</p>
<p>It looks like there have been some changes in optional parameter handling in PyO3.
You may need to deal with keyword arguments with decorators.</p>
<p>Now, let’s try <code>xvc.file().track()</code> once more with <code>dir-0001/</code>.</p>
<p>It seems to be working now. With an existing directory, it doesn’t show an
error.</p>
<p>What does <code>xvc.file().list()</code> show?</p>
<p>It shows a single string with <code>\n</code> in it. It looks like we need to handle this in
the output thread.</p>
<p>I added a <code>replace("\\n", "\n")</code> to <code>output_str</code> at the end, but it didn’t make a
difference. I tested in the notebook with <code>list_result.split("\n")</code>, and these
<code>\n</code> characters are indeed CRLF. So Jupyter shows CRLF in strings with <code>\n</code>, and
this is not something we should try to fix, I think.</p>
<p>You can search how to show <code>\n</code> characters in a Jupyter notebook with CRLF.</p>
<p>The output string should be fed into <code>repr</code>, as in:</p>
<pre><code class="language-python"># Define a string with CRLF characters
text_crlf = "Hello\r\nWorld\r\nThis is a test string."

# Use repr to show the \r\n characters explicitly
print(repr(text_crlf))
</code></pre>
<p>I tested this, and GPT misleads. <code>repr</code> is when you <em>want</em> to show <code>\n</code>, not
vice versa. When I <code>print(list_result)</code>, it prints the results properly.</p>
<p>It looks like we don’t need to make this a priority now. We can tell the user in the
notebook that the commands are intentionally returning strings, and they can
process or print them however they want.</p>]]></content:encoded>
    </item>
    <item>
      <title>Devlog 1: XVC Root and Python Bindings Debugging</title>
      <published>2024-05-24T10:12:50+00:00</published>
      <updated>2024-05-24T10:12:50+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Fri, 24 May 2024 10:12:50 +0000</pubDate>
      <link>https://emresahin.net/devlog-1/</link>
      <guid isPermaLink="true">https://emresahin.net/devlog-1/</guid>
      <description>I should have a dispatch method that receives an XvcRootOpt and runs a command with it. The dispatcher can also update the XvcRoot from None to Some(XvcRoot) in some cases, so it should receive a mutable XvcRootOpt or return one after receiving ownership. The issue with having a mutable element i...</description>
      <category>Devlog</category>
      <category>XVC</category>
      <category>xvc</category>
      <category>debugging</category>
      <category>shell</category>
      <category>jupyter-lab</category>
      <category>rust</category>
      <category>python-bindings</category>
      <content:encoded><![CDATA[<p>I should have a dispatch method that receives an <code>XvcRootOpt</code> and runs a command
with it. The dispatcher can also update the <code>XvcRoot</code> from <code>None</code> to <code>Some(XvcRoot)</code>
in some cases, so it should receive a mutable <code>XvcRootOpt</code> or return one after
receiving ownership.</p>
<p>The issue with having a mutable element is that we need to update <code>XvcRoot</code> from
<code>Arc&lt;XvcRootInner&gt;</code> to <code>Arc&lt;RwLock&lt;XvcRootInner&gt;&gt;</code>, which will cause almost all
machinery around <code>XvcRoot</code> to require locking the object first. This is too
large a refactoring.</p>
<p>However, I can write an <code>xvc_root!</code> macro to replace <code>xvc_root.read()</code> or
whatever is required to minimize the code changes. <code>XvcRoot</code> can also be a wrapper
object, but that’s too much fuss, and the responsibility may be misplaced.</p>
<p>Let’s update the type and let the dispatcher update <code>XvcRootInner</code> only
when necessary. Let’s see what will require updating.</p>
<hr>
<p>We have a bug in the Python bindings where <code>xvc.file.track()</code> enters an infinite loop.</p>
<p>There can be multiple reasons, but are you sure that the <code>xvc</code> CLI works with the
same command?</p>
<p>Let’s start testing by creating another notebook server.</p>
<p>The notebook creates an Xvc repo successfully and returns the root with
<code>xvc.root()</code>, but <code>xvc.file().track("test-data/dir-0001/")</code> never finishes.</p>
<p>It’s likely that this is caused by something in the background threads.</p>
<p>The CLI command completes successfully, but we may be passing the command
incorrectly. Let’s try <code>dir-0001</code> only.</p>
<p>Even if an incorrect path is provided, it shouldn’t enter an infinite loop.</p>
<p>When I provide the correct path, <code>dir-0001</code>, it returns. There may be something going on with finding the root of the Xvc repo.</p>
<p>It returns, but it’s also giving an error that it cannot find a repository. I should take a closer look.</p>
<p>The <code>Xvc</code> struct wasn’t implementing <code>Debug</code>, so I added it.</p>
<p>Let’s add it to others, <code>XvcFile</code>, etc.</p>
<p>Added it. Recompiling to get more info when we run <code>track</code> with an incorrect path.</p>
<p>You can also make it run after a commit to avoid starting a new server if one is already running.</p>
<p>Added conditionals, and now it runs the server if none is running in the background.</p>
<pre><code class="language-zsh">PORT=7979
if [[ ! -s "$(ps ax | rg -v ps | rg jupyter-lab | rg $PORT)" ]] ; then
  jupyter lab --port=$PORT --notebook-dir=Readme/ &amp;
  open http://localhost:${PORT}/lab/workspaces/auto-H/tree/Readme.ipynb
fi
</code></pre>]]></content:encoded>
    </item>
    <item>
      <title>A Brief History of Xvc</title>
      <published>2024-01-22T09:17:15+00:00</published>
      <updated>2024-01-22T09:17:15+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Mon, 22 Jan 2024 09:17:15 +0000</pubDate>
      <link>https://emresahin.net/a-brief-history-of-xvc/</link>
      <guid isPermaLink="true">https://emresahin.net/a-brief-history-of-xvc/</guid>
      <description>In the first months of 2021, I decided to return to life after a long legal battle for divorce. Covid was still raging. I wasn’t keen to start a company or work in my country due to my half-deaf ears. I decided to find some open source projects and contribute, maybe get recognition, maybe hired. ...</description>
      <category>xvc</category>
      <category>development</category>
      <category>free software</category>
      <category>Software Engineering</category>
      <category>xvc</category>
      <category>history</category>
      <category>dvc</category>
      <category>Rust</category>
      <category>MLOps</category>
      <category>Open Source</category>
      <category>Iterative.ai</category>
      <category>Git Annex</category>
      <category>serde</category>
      <category>Blake3</category>
      <category>PyO3</category>
      <content:encoded><![CDATA[<p>In the first months of 2021, I decided to return to life after a long legal battle for divorce. Covid was still raging. I wasn’t keen to start a company or work in my country due to my half-deaf ears. I decided to find some open source projects and contribute, maybe get recognition, maybe hired.</p>
<p>I saw an ad on Stack Overflow Jobs those days about <em>employment by contributing to open source projects.</em> I applied to that. A few weeks later, the CTO of [iterative.ai] got in touch and I started working on DVC documentation. Initially on a per-hour basis, and after May 2021, as a full-time employee.</p>
<p>Initially, I liked the tool we were building very much. The team was awesome. (Still, they are.) It was one of the best periods of my life, especially in my turbulent still-ongoing-divorce-period pressures. I know I will always miss them.</p>
<p>My job was learning DVC, documenting it, and helping newcomers grasp it easily. It was a fun job. Until then, I didn’t see myself as a technical writer. English is not my native tongue and I never have lived in an English-speaking country. Nevertheless, I think I wasn’t <em>too bad</em> at it.</p>
<p>When I was first learning the tool, I began to use it everywhere. I was an avid user of Git Annex once. DVC looked better. I don’t remember why I lost interest in Git Annex after many years, but it was probably related to symbolic links not working on Windows (or on Termux). DVC had multiple ways of connecting the cache and the files in the workspace, including hardlinks and copy, so it was a breath of fresh air for me.</p>
<p>I began to use it for my large collections. Keeping track of my binary files in Git was something I always desired. Git is the <em>least sucking</em> version control system among the ones I used previously (SVN, hg, darcs…) and I’d rather keep using it everywhere rather than learning new tools for binary files.</p>
<p>After some time I began to use the tool for my personal file collections. I noticed its performance became a burden. I was tracking maybe a few gigabytes of files with it and basic file operations became slower as I added more. I noticed I was becoming distracted after I wrote a <code>dvc</code> command. It took some time to confess that the tool I liked once and was earning my salary with was not a tool that I liked to use.</p>
<p>I don’t know what <em>real professionals</em> would do at this point. I never had a good LinkedIn profile. When I met a similar problem with the example repository that’s supposed to contain 70,000 small files, I brought the issue forward. I wrote a shell script that was basically doing the same thing as <code>dvc add</code> and it worked much faster than the actual command. The shell script was naïve and I thought DVC must have <em>at least</em> that level of speed. It didn’t. Simply calling <code>md5sum</code> on files and copying them to appropriate location in <code>.dvc/cache</code> was way faster. How could this be?</p>
<p>I had cursory observations on the codebase. I know some decisions (like a large central class that connects everything, separate <code>.dvc</code> files for each tracked file) that may lead to degradation. Although I don’t see it as <em>the problem</em>, Python was also not helpful. These are rough observations.</p>
<p>It was September 2021. I was also teaching myself Rust. I wrote an email to the CTO and CEO of the company to request a sabbatical to work on DVC. My plan was to rewrite certain portions (or commands) in Rust and wrap them with PyO3. It could fail. So to have <em>skin in the game</em>, I said I’ll work for free during this time and if I fail to make DVC faster for some reason, I’ll return to my writing position.</p>
<p>They didn’t accept. I didn’t try to persuade them. The decision was rational and although I’d say <em>go ahead and see what happens</em> if I were in their shoes just to make my employee happy, they aren’t <em>crazy-managers</em> as I once was. Probably there are many factors that I’m not aware of. I returned to my post and continued to write documentation for another 9 months. In the meantime I studied Rust and thought about how I could architect a similar tool. Where does DVC go wrong?</p>
<p>In April 2022, I informed the CTO that I’d like to take a sabbatical for my book. I have a political-SF book and after the Ukrainian war started with a (albeit minor) probability of nuclear attack <em>on the other shore of Black Sea</em>, I thought it’s not a time to work on something I stopped liking. My performance in the last quarter was also not something I was proud of. I didn’t feel good.</p>
<p>When I retired to sabbatical in July though, while writing the book, I thought writing the software that I wanted to see was also <em>something on my mind before nuclear war.</em> I had notes about the architecture I was planning. I wanted to see if I could apply an Entity-Component System to this basic problem, without any Object-Oriented conceptions. I believe it looks cool. I’m still simplifying and testing the idea, and it looks better to my mind than mixing data and functions for no reason.</p>
<p>After I made the repository public, I resigned from Iterative.</p>
<p>In a sense, Xvc owes its existence to DVC, and the name is a tribute to this. I hope they squash their bugs, and improve their user experience, and be a long-term player in the crowded market they are in. I don’t intend to be a “competitor”, because I prefer being developer/architect rather than a “technical steward to VC money”, and the license of Xvc is GPL-3 to signal this.</p>
<p>The Xvc command line interface, however, is as different as it can be from DVC. The command names are different; DVC has commands similar to Git (<code>push</code>, <code>fetch</code>, <code>pull</code>, <code>commit</code>), while Xvc tries to be different from Git to reduce the user’s mental load. For example, as a writer, I noticed that “Git remotes” and “DVC remotes” was confusing, so I called them “Xvc storages”. DVC calls the units of a pipeline <em>stages</em>; the same concept is called <em>steps</em> in Xvc, because <em>stage</em> in Git is something completely different.</p>
<p>Internally, the architecture is also very different. Xvc uses serialization (with serde) instead of YAML. It can export/import pipelines from YAML (or JSON), but YAML is not as central as in DVC. (I believe YAML is overused in our industry, and it’s an employment guarantee for another generation of developers but there is better work than keeping up a half-baked configuration format.) Xvc doesn’t keep its artifacts in the user’s workspace (except <code>.xvcignore</code> files). They are all stored in the <code>.xvc/</code> directory. The DVC way of doing things makes merging <code>.dvc</code> files easier. To overcome the problems caused by merging large metadata files, Xvc keeps track of events and replays them to get the final state of the repository. All metadata storage and retrieval operations revolve around the <code>XvcStore&lt;T&gt;</code> struct in Xvc. Typically, if the user runs an <code>xvc</code> command, only the updated store events (added files, changed pipelines, etc.) are stored. There are optimizations in this front, but I profile first and optimize later.</p>
<p>Algorithms for data digests are configurable; by default Xvc uses Blake3, but it is configurable to use SHA2-256, SHA3-256, or Blake2s. It can be modified to use any 256-bit digest quickly. There are some features that are not found in DVC, and more will come. So, although I’m solving a similar problem, Xvc is not “DVC rewritten in Rust,” it’s a different tool completely.</p>
<p>Currently, it doesn’t have as much eye candy as DVC. In time, I plan to add Python, Julia, and R APIs, notebook integration, experiment tracking (without relying on Git internals), data labeling and filtering, and other MLOps features. I’m building with a goal to make these features available without making the rest of the software slower.</p>
<p>I’ve found the tool I was looking for to track my kids’ photos and Ottoman OCR datasets in a Git repository. I’m tracking more than 1TB of files in a single repository with Xvc and adding another 10TB looks feasible now.</p>]]></content:encoded>
    </item>
    <item>
      <title>Perl Error Building rust-openssl with Maturin, Ubuntu, and Manylinux</title>
      <published>2023-11-02T09:19:26+00:00</published>
      <updated>2023-11-02T09:19:26+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Thu, 02 Nov 2023 09:19:26 +0000</pubDate>
      <link>https://emresahin.net/perl-error-building-rust-openssl-with-maturin--ubuntu-and-manylinux/</link>
      <guid isPermaLink="true">https://emresahin.net/perl-error-building-rust-openssl-with-maturin--ubuntu-and-manylinux/</guid>
      <description>While building Python packages for Xvc with Maturin, I was receiving an error in GitHub Actions CI for Linux packages. Can't locate IPC/Cmd.pm in @INC (@INC contains: /home/runner/work/xvc.py/xvc.py/target/x86_64-unknown-linux-gnu/release/build/openssl-sys-844a96d66ae533b1/out/openssl-build/build...</description>
      <category>Xvc</category>
      <category>CI/CD</category>
      <category>Rust</category>
      <category>Xvc</category>
      <category>Maturin</category>
      <category>GitHub Actions</category>
      <category>Perl</category>
      <category>Python</category>
      <category>OpenSSL</category>
      <category>CentOS</category>
      <category>Debian</category>
      <content:encoded><![CDATA[<p>While building <a href="https://github.com/iesahin/xvc.py">Python packages for Xvc</a> with Maturin, I was receiving an error in GitHub Actions CI for Linux packages.</p>
<pre><code class="language-text">Can't locate IPC/Cmd.pm in @INC (@INC contains: /home/runner/work/xvc.py/xvc.py/target/x86_64-unknown-linux-gnu/release/build/openssl-sys-844a96d66ae533b1/out/openssl-build/build/src/util/perl /usr/local/lib64/perl5 /usr/local/share/perl5 /usr/lib64/perl5/vendor_perl /usr/share/perl5/vendor_perl /usr/lib64/perl5 /usr/share/perl5 . /home/runner/work/xvc.py/xvc.py/target/x86_64-unknown-linux-gnu/release/build/openssl-sys-844a96d66ae533b1/out/openssl-build/build/src/external/perl/Text-Template-1.56/lib)
</code></pre>
<p>The issue seemed to be a missing Perl package. Building <code>rust-openssl</code> now appears to require the <code>perl-core</code> package.</p>
<p>I added <code>apt-get update &amp;&amp; apt-get install perl-core</code> to the CI configuration, but that didn’t work.</p>
<p>However, as <a href="https://github.com/sfackler/rust-openssl/issues/2036#issuecomment-1724324145">this GitHub issue comment suggests</a>, Maturin with its <code>manylinux</code> support uses different kinds of Docker containers to build the packages. We must detect whether it is a CentOS or Debian-based container and install the missing packages accordingly.</p>
<p>The <em>Building wheels</em> step in the configuration should look similar to:</p>
<pre><code class="language-yaml">- name: Build wheels
  uses: PyO3/maturin-action@v1
  with:
    target: ${{ matrix.target }}
    manylinux: auto
    args: --release --out dist
    before-script-linux: |
      # If we're running on RHEL/CentOS, install needed packages.
      if command -v yum &amp;&gt; /dev/null; then
          yum update -y &amp;&amp; yum install -y perl-core openssl openssl-devel pkgconfig libatomic

          # If we're running on i686, we need to symlink libatomic
          # in order to build openssl with the -latomic flag.
          if [[ ! -d "/usr/lib64" ]]; then
              ln -s /usr/lib/libatomic.so.1 /usr/lib/libatomic.so
          fi
      else
          # If we're running on a Debian-based system.
          apt update -y &amp;&amp; apt-get install -y libssl-dev openssl pkg-config
      fi
</code></pre>
<p>This will run the specified script before executing the Maturin build command in the container, ensuring all missing packages are installed.</p>]]></content:encoded>
    </item>
    <item>
      <title>bits 4</title>
      <published>2023-03-14T11:19:00+00:00</published>
      <updated>2023-03-14T11:19:00+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Tue, 14 Mar 2023 11:19:00 +0000</pubDate>
      <link>https://emresahin.net/bits-4/</link>
      <guid isPermaLink="true">https://emresahin.net/bits-4/</guid>
      <description>I want to use channels for Xvc pipelines. A thread will be created for each step and will receive dependency state changes via channels. Note that there may be thousands of dependencies (e.g., files for a step), and creating these channels beforehand is not feasible. The Crossbeam library has Sel...</description>
      <category>bits</category>
      <category>xvc</category>
      <category>Rust</category>
      <category>crossbeam</category>
      <category>select</category>
      <category>channels</category>
      <category>concurrency</category>
      <content:encoded><![CDATA[<p>I want to use channels for Xvc pipelines. A thread will be created for each
step and will receive dependency state changes via channels.</p>
<p>Note that there may be thousands of dependencies (e.g., files for a step), and
creating these channels beforehand is not feasible.</p>
<p>The Crossbeam library has
<a href="https://docs.rs/crossbeam/latest/crossbeam/channel/struct.Select.html#">Select</a>
which allows waiting for an arbitrary number of channels. Now, I need to figure out
how to structure the threads and update their states. A minor task! 😆</p>]]></content:encoded>
    </item>
    <item>
      <title>XVC State Machine</title>
      <published>2022-12-10T20:49:21+00:00</published>
      <updated>2022-12-10T20:49:21+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Sat, 10 Dec 2022 20:49:21 +0000</pubDate>
      <link>https://emresahin.net/xvc-state-machine/</link>
      <guid isPermaLink="true">https://emresahin.net/xvc-state-machine/</guid>
      <description>I began writing Xvc’s pipeline and dependency handling. The best way to handle dependency states seems to be through a state machine. A state machine is a simple abstraction that shows state changes with respect to inputs. It can also have outputs for these state changes. There are some varieties...</description>
      <category>development</category>
      <category>xvc</category>
      <category>xvc</category>
      <category>rust</category>
      <category>finite state machines</category>
      <category>state machines</category>
      <category>pipeline</category>
      <category>programming</category>
      <content:encoded><![CDATA[<p>I began writing Xvc’s pipeline and dependency handling.
The best way to handle dependency states seems to be through a state machine.
A state machine is a simple abstraction that shows state changes with respect to inputs.
It can also have outputs for these state changes.
There are some varieties of this, but Xvc’s state machine (SM) is a simple one.</p>
<p>I first tried to use the <a href="https://crates.io/crates/rust-fsm"><code>rust-fsm</code></a> library, but it became apparent that Xvc pipeline steps’ states are tied to <code>XvcRoot</code>. That is, if we are to check the presence of a file or the value of a parameter, we have to do it relative to the repository root.
The root directory should be taken into consideration in every transition.</p>
<p>I checked the code and noticed that the FSM is actually very simple.
I copied it, added <code>&amp;XvcRoot</code> to the <code>transition</code> and <code>output</code> functions in the trait definition, and implemented it for <code>XvcOutput</code>, <code>XvcDependency</code>, and <code>XvcStep</code>.</p>
<p>These are the constituents of a pipeline.
A pipeline is composed of an <code>XvcStep</code> that defines a command, and each step can have multiple <code>XvcDependency</code> and <code>XvcOutput</code> definitions.
For each of these structs, I’ve added fields that represent their current state.</p>
<p>For example, an <code>XvcOutput</code> can be <code>Missing</code>, <code>Found</code>, <code>Old</code>, or <code>Ok</code>.
An <code>XvcStep</code> that produces this <code>XvcOutput</code> doesn’t check the dependency content hash if an output is missing.
However, if an <code>XvcOutput</code> is <code>Found</code>, the <code>XvcStepStateMachine</code> checks the <code>XvcDependency</code> states and their modification times to see if they have changed since the output was generated.
An <code>XvcStep</code> is invalidated when an <code>XvcOutput</code> is <code>Missing</code> or an <code>XvcDependency</code> has changed after the last command run.</p>
<p>Unlike DVC, I added the ability for an <code>XvcStep</code> to depend on other <code>XvcStep</code>s.
They communicate through outputs. I’ve added <code>XvcDependency::Step(XvcStep)</code> to the <code>XvcDependency</code> definition.
I’m also planning <code>XvcDependency::Pipeline(XvcPipeline)</code> to allow steps to depend on other pipelines, so that pipelines can be run in order.</p>
<p>Currently, the following are included as <code>XvcDependency</code>:</p>
<ul>
<li><code>File</code>: A (binary or text) file in the repository. If the metadata (size and modification time) or the content changes, the dependent step becomes invalidated.</li>
<li><code>Directory</code>: A directory that contains files. If a file is added to or removed from the directory, or any of the files are changed, the associated step becomes invalidated.</li>
<li><code>Glob</code>: A glob such as <code>my-data/*.png</code>. If the list of files changes or their content has changed, the associated step becomes invalidated.</li>
<li><code>Parameter</code>: Xvc can parse YAML, TOML, and JSON files and get the values of variables. It’s possible to define these (hyper)parameters as dependencies.</li>
<li><code>URL</code>: An HTTPS URL, which is checked first by metadata and then by content to see whether it has changed.</li>
<li><code>Step</code>: A previously defined step; if it’s invalidated, the depending step also becomes invalidated.</li>
</ul>
<p>Additionally, I’m planning to add the following items to Xvc as dependencies:</p>
<ul>
<li><code>Lines {path, begin, end}</code>: Lines in a text file. This can be used for general-purpose input tracking. If the given lines in a file are changed, the dependent step becomes invalidated.</li>
<li><code>Regex {path, regex}</code>: If the regular expression result on the file changes, the dependent stage becomes invalidated.</li>
<li><code>Pipeline { name }</code>: If any of the steps in a pipeline is invalidated, the pipeline is also invalidated, or the step that depends on this pipeline becomes invalidated.</li>
</ul>
<p>Each of these dependencies is checked minimally; that is, when their size on disk is detected to have changed, they are considered changed without checking the content hash. It needs a very detailed state machine to track the changes without bugs.</p>
<p>I’ve noticed that if I can write such a state machine, most of the I/O operations can be done in parallel. If two steps do not depend on each other in the dependency graph, they can be run in parallel. The state machine’s granularity allows this.</p>]]></content:encoded>
    </item>
    <item>
      <title>Xvc Devlog - 221109</title>
      <published>2022-11-10T09:26:00+00:00</published>
      <updated>2022-11-10T09:26:00+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Thu, 10 Nov 2022 09:26:00 +0000</pubDate>
      <link>https://emresahin.net/xvc-devlog---221109/</link>
      <guid isPermaLink="true">https://emresahin.net/xvc-devlog---221109/</guid>
      <description>🐇 How do you want to proceed from here, Mr. 🐢? 🐢 I think I can implement Rsync today. Looking at ssh2-rs , though, I think we can implement file transfer without relying on rsync . It might be easier to implement everything within the code. 🐇 Then you should rename the issue to new ssh . 🐢 Fair. ...</description>
      <category>devlog</category>
      <category>development</category>
      <category>xvc</category>
      <category>ssh</category>
      <category>rsync</category>
      <category>storage</category>
      <category>feature-flags</category>
      <category>rust</category>
      <category>libssh2</category>
      <content:encoded><![CDATA[<p>🐇 How do you want to proceed from here, Mr. 🐢?</p>
<p>🐢 I think I can implement Rsync today. Looking at <a href="https://docs.rs/ssh2/latest/ssh2/">ssh2-rs</a>, though, I think we can implement file transfer without relying on <code>rsync</code>. It might be easier to implement everything within the code.</p>
<p>🐇 Then you should rename the issue to <code>new ssh</code>.</p>
<p>🐢 Fair. There is also the <a href="https://docs.rs/ssh-rs/0.2.2/ssh_rs/">ssh_rs</a> crate, but it doesn’t have full support for the protocol. Instead, we can have another command, like <code>xvc storage new ssh</code>, that uses <code>libssh2</code> via the crate mentioned above. It has some limitations with OpenSSH on macOS.</p>
<p>🐇 From the <a href="https://github.com/alexcrichton/ssh2-rs">crate’s README</a>, it looks like you can enable the <code>vendored-openssl</code> feature to compile it statically.</p>
<p>🐢 Let’s go ahead then. It’s better to compile it behind a feature flag, though.</p>
<p>🐇 Yup. Rsync can be separate. I think for now you can implement rsync via <code>Exec::cmd</code> and make <code>new ssh</code> a new issue.</p>
<p>🐢 I’ll copy this conversation there.</p>
<hr>
<p>🐇 Now, let’s start implementing <code>rsync</code>.</p>
<p>🐢 Do we really want to hide it behind a feature flag? It doesn’t bring any extra complexity to <code>generic</code>, for example—just using the commands and returning the errors.</p>
<p>🐇 I think so. If the user doesn’t have <code>rsync</code> on their system, they’ll just get errors. We don’t need to make the implementation optional, but the tests might be.</p>
<p>🐢 OK.</p>]]></content:encoded>
    </item>
  </channel>
</rss>
