<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>emre şahin's digital garden 🍃 - historical documents</title>
    <link>https://emresahin.net/tags/historical-documents/</link>
    <description>Posts in the historical documents tag</description>
    <language>en</language>
    <managingEditor>contact@emresahin.net (Emre Şahin)</managingEditor>
    <lastBuildDate>Tue, 15 Sep 2026 19:46:32 +0000</lastBuildDate>
    <atom:link href="https://emresahin.net/tags/historical-documents/rss.xml" rel="self" type="application/rss+xml"/>
    <item>
      <title>Paper Review: Text Line Segmentation of Historical Documents: A Survey</title>
      <published>2012-07-27T14:00:00+00:00</published>
      <updated>2012-07-27T14:00:00+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Fri, 27 Jul 2012 14:00:00 +0000</pubDate>
      <link>https://emresahin.net/text-line-segmentation-of-historical-documents/</link>
      <guid isPermaLink="true">https://emresahin.net/text-line-segmentation-of-historical-documents/</guid>
      <description>Authors: Laurance Likforman-Sulem, Abderrezak Sahour, Bruno Taconet URL: http://arxiv.org/pdf/0704.1267.pdf Keywords: page segmentation overlapping components image quality document complexity preprocessing projection based smearing based grouping based hough transform based repulsive attractive ...</description>
      <category>Paper Review</category>
      <category>text line segmentation</category>
      <category>historical documents</category>
      <category>Hough transform</category>
      <category>document analysis</category>
      <category>image processing</category>
      <content:encoded><![CDATA[<h1 id="authors-laurance-likforman-sulem-abderrezak-sahour-bruno-taconet">Authors: Laurance Likforman-Sulem, Abderrezak Sahour, Bruno Taconet</h1>
<h1 id="url-httparxivorgpdf07041267pdf">URL: <a href="http://arxiv.org/pdf/0704.1267.pdf">http://arxiv.org/pdf/0704.1267.pdf</a></h1>
<h1 id="keywords">Keywords:</h1>
<ul>
<li>page segmentation</li>
<li>overlapping components</li>
<li>image quality</li>
<li>document complexity</li>
<li>preprocessing</li>
<li>projection based</li>
<li>smearing based</li>
<li>grouping based</li>
<li>hough transform based</li>
<li>repulsive attractive</li>
<li>stochastic</li>
<li>touching components</li>
</ul>
<h1 id="q1-what-are-the-most-usable-techniques-for-ottoman-divans">Q1: What are the most usable techniques for Ottoman divans?</h1>
<p>Likforman-Sulem and Faure’s technique, which uses Gestalt criteria to associate text elements, might be of use. Feldbach and Tennies’ work, which was tested on Church Registers, may also be helpful. The Hough transform can be used. The Repulsive-Attractive method of Öztop et al. is also applicable. Stochastic methods by Tseng and Lee, which use a probabilistic Viterbi algorithm, can also be utilized.</p>
<h1 id="q2-how-are-touching-components-successfully-delimited">Q2: How are touching components successfully delimited?</h1>
<p>A touching component can be detected by its size. Subsequently, it should either be assigned to a lower or upper line, or be separated. Successful separation requires letter images or skeletons (which we lack).</p>
<h1 id="q3-how-is-the-hough-transform-used">Q3: How is the Hough transform used?</h1>
<p>Centroids of the connected components (CCs) are used as units of the Hough transform. Line hypotheses are developed in the Hough domain and verified in the image domain.</p>
<h1 id="q4-what-are-the-problems-specific-to-non-latin-texts">Q4: What are the problems specific to non-Latin texts?</h1>
<p>The baseline of Hebrew is at the upper part of the letters because of their box shape. Devanagari and similar scripts also have a headline on top of them. Diacritics and inter-letter shapes pose problems for Arabic.</p>]]></content:encoded>
    </item>
    <item>
      <title>Paper Review: Computerized Paleography: Tools for Historical Manuscripts</title>
      <published>2012-07-23T07:08:00+00:00</published>
      <updated>2012-07-23T07:08:00+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Mon, 23 Jul 2012 07:08:00 +0000</pubDate>
      <link>https://emresahin.net/12061-15-2433/</link>
      <guid isPermaLink="true">https://emresahin.net/12061-15-2433/</guid>
      <description>Authors: Lior Wolf, Liza Potikha, Nachum Dershowitz, Roni Shweka, Yaacov Choueka Keywords: handwritten paleography fragments SIFT sparse coding dictionaries Q1: What is the ultimate goal of the authors? The two main goals are providing tools to bring together fragments of the same page (specifica...</description>
      <category>Computer Vision</category>
      <category>Paper Review</category>
      <category>Paleography</category>
      <category>handwriting</category>
      <category>historical documents</category>
      <category>SIFT</category>
      <category>classic CV</category>
      <category>sparse coding</category>
      <category>Cairo Genizah</category>
      <category>paleography</category>
      <content:encoded><![CDATA[<p>Authors: Lior Wolf, Liza Potikha, Nachum Dershowitz, Roni Shweka, Yaacov Choueka</p>
<h1 id="keywords">Keywords:</h1>
<ul>
<li>handwritten</li>
<li>paleography</li>
<li>fragments</li>
<li>SIFT</li>
<li>sparse coding</li>
<li>dictionaries</li>
</ul>
<h1 id="q1-what-is-the-ultimate-goal-of-the-authors">Q1: What is the ultimate goal of the authors?</h1>
<p>The two main goals are providing tools to bring together fragments of the same page (specifically from the Cairo Genizah) and trying to classify handwriting and dates.</p>
<h1 id="q2-how-is-sift-used">Q2: How is SIFT used?</h1>
<p>SIFT is used at various points of a letter to generate descriptors. There are 100,000 descriptors overall before inputting them into k-means. SIFT serves as the main classification technique.</p>
<h1 id="q3-how-did-they-produce-the-letter-dictionaries">Q3: How did they produce the letter dictionaries?</h1>
<p>They produced letter dictionaries using the generated SIFT descriptors and k-means clustering to find representative visual words.</p>
<h1 id="q4-what-is-sparse-coding-and-its-importance">Q4: What is sparse coding and its importance?</h1>
<p>Sparse coding is used to code documents (or any other codable thing) with a separating/descriptive code which also shows the similarity between items. An example might be the bag of visual words approach.</p>
<h1 id="q5-are-there-any-relevant-techniques-for-our-research">Q5: Are there any relevant techniques for our research?</h1>
<p>This is relevant to the Divan matching problem. This work might be cited in historical document matching, although it covers techniques that are generally well-known in the field.</p>]]></content:encoded>
    </item>
  </channel>
</rss>
