<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>emre şahin's digital garden 🍃 - transcription</title>
    <link>https://emresahin.net/tags/transcription/</link>
    <description>Posts in the transcription tag</description>
    <language>en</language>
    <managingEditor>contact@emresahin.net (Emre Şahin)</managingEditor>
    <lastBuildDate>Tue, 15 Sep 2026 19:46:32 +0000</lastBuildDate>
    <atom:link href="https://emresahin.net/tags/transcription/rss.xml" rel="self" type="application/rss+xml"/>
    <item>
      <title>A Need for Yet Another Transliteration Alphabet for Ottoman</title>
      <published>2012-09-25T14:00:00+00:00</published>
      <updated>2012-09-25T14:00:00+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Tue, 25 Sep 2012 14:00:00 +0000</pubDate>
      <link>https://emresahin.net/12126-3-1424/</link>
      <guid isPermaLink="true">https://emresahin.net/12126-3-1424/</guid>
      <description>The Ottoman Text Archival Project has its own reversible transcription system. However, for word labels, this is an overkill and requires too much work from experts. I’m looking for a one-to-one mapping between the different visual elements of a word and its representation in UTF-8. The labels sh...</description>
      <category>Dervaze</category>
      <category>Linguistics</category>
      <category>Typography</category>
      <category>transliteration</category>
      <category>Ottoman</category>
      <category>visual encoding</category>
      <category>UTF-8</category>
      <category>alphabet</category>
      <category>transcription</category>
      <content:encoded><![CDATA[<p>The Ottoman Text Archival Project has its own reversible transcription
system. However, for word labels, this is an overkill and requires
too much work from experts.</p>
<p>I’m looking for a one-to-one mapping between the different <em>visual</em> elements
of a word and its representation in UTF-8. The labels should be simple
to remember, yet distinctive enough to represent visual variations of
words.</p>
<p>I’m considering creating letter+digit codes. The letter part will reflect the
most similar sound, and the digit will reflect the visual variation. In this
case, I need a Latin letter for each letter in the Ottoman alphabet, but
there are not enough letters in the standard Turkish alphabet to denote all
the letters of Ottoman. So, how can I solve this?</p>]]></content:encoded>
    </item>
    <item>
      <title>Converting Latin-based Turkish spelling to Ottoman</title>
      <published>2012-09-23T14:00:00+00:00</published>
      <updated>2012-09-23T14:00:00+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Sun, 23 Sep 2012 14:00:00 +0000</pubDate>
      <link>https://emresahin.net/12124-4-2782/</link>
      <guid isPermaLink="true">https://emresahin.net/12124-4-2782/</guid>
      <description>I’m working on a system to search Ottoman document collections. In order to query a large collection in Ottoman, the user needs to write the query in Ottoman, which uses an Arabic-based alphabet with a completely different set of spelling rules. This limits usability, since most users will not be...</description>
      <category>Dervaze</category>
      <category>search</category>
      <category>conversion</category>
      <category>query</category>
      <category>transcription</category>
      <category>Ottoman</category>
      <category>Turkish</category>
      <category>NLP</category>
      <content:encoded><![CDATA[<p>I’m working on a system to search Ottoman document collections.</p>
<p>In order to query a large collection in Ottoman, the user needs to write
the query in Ottoman, which uses an Arabic-based alphabet with a completely
different set of spelling rules. This limits usability, since most
users will not be familiar with the spelling. Experts are, but we
can’t assume all users will be able to use it.</p>
<p>There are various methods of transcribing Ottoman to modern Turkish.
Many of these use diacritics to denote different long vowels. When it
comes to consonants that are spelled identically in Turkish but correspond
to different letters in Ottoman, most of these transcription systems are
silent. They don’t represent the difference between the letters <em>tha</em> and
<em>sin</em> for example, although the first is used in Arabic words
considerably. Turks pronounce these two letters identically since
Ottoman times, so they are written identically as <em>s</em> in Turkish.</p>
<p>The information content of these two writing systems is different. One
has three <em>s</em> letters, the other has a rich set of vowels etc. The Latin-based
Turkish script looks more verbose, so I decided to reduce this verbosity
to get a set of <em>probable</em> Ottoman spellings.</p>]]></content:encoded>
    </item>
  </channel>
</rss>
