<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>emre şahin's digital garden 🍃 - deep learning</title>
    <link>https://emresahin.net/tags/deep-learning/</link>
    <description>Posts in the deep learning tag</description>
    <language>en</language>
    <managingEditor>contact@emresahin.net (Emre Şahin)</managingEditor>
    <lastBuildDate>Tue, 29 Sep 2026 14:57:43 +0000</lastBuildDate>
    <atom:link href="https://emresahin.net/tags/deep-learning/rss.xml" rel="self" type="application/rss+xml"/>
    <item>
      <title>Types of regularization in ML</title>
      <published>2020-12-26T01:07:49+00:00</published>
      <updated>2020-12-26T01:07:49+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Sat, 26 Dec 2020 01:07:49 +0000</pubDate>
      <link>https://emresahin.net/regularization-types/</link>
      <guid isPermaLink="true">https://emresahin.net/regularization-types/</guid>
      <description>What is regularization? Regularization is a technique used to reduce the complexity of a model, thereby preventing overfitting. There are three common types of regularization used in Deep Neural Networks (DNN): L2 Regularization: We define the complexity of a model by the sum of the squares of it...</description>
      <category>AI</category>
      <category>Machine Learning</category>
      <category>deep learning</category>
      <category>regularization</category>
      <category>ml</category>
      <category>neural networks</category>
      <category>l1 regularization</category>
      <category>l2 regularization</category>
      <category>dropout</category>
      <content:encoded><![CDATA[<p>What is regularization?</p>
<p>Regularization is a technique used to reduce the complexity of a model, thereby preventing overfitting. There are three common types of regularization used in Deep Neural Networks (DNN):</p>
<p><strong>L2 Regularization:</strong> We define the complexity of a model by the sum of the squares of its weights: $W = w_0^2 + w_1^2 + … + w_n^2$. We add this term to the loss function to obtain:</p>
<p>$L(\text{data}, \text{model}) = \text{loss}(\text{data}, \text{model}) + \lambda \sum w_i^2$</p>
<p>We then aim to minimize this total loss. As the derivative of $W$ with respect to each weight $w_i$ is $2w_i$, backpropagation reduces the weights by penalizing larger values, effectively “decaying” them.</p>
<p><strong>L1 Regularization:</strong> This is similar to L2 regularization, but $W$ is defined as the sum of the absolute values of the weights:</p>
<p>$W = \sum |w_i|$</p>
<p>The derivative of $W$ with respect to $w_i$ is a constant ($\pm 1$) this time, so weights can be reduced exactly to zero, unlike in L2 regularization. This often leads to sparse models.</p>
<p><strong>Dropout:</strong> Unlike the previous two methods, dropout is implemented as a layer within the neural network rather than a modification to the loss function.</p>
<p>A dropout layer randomly sets a subset of activations to zero during training. For example, a dropout layer with a rate of 0.3 will randomly deactivate 30% of the neurons in that layer for each training step.</p>]]></content:encoded>
    </item>
    <item>
      <title>Recurrent Neural Networks</title>
      <published>2014-10-09T21:00:00+00:00</published>
      <updated>2014-10-09T21:00:00+00:00</updated>
      <author>Emre Şahin</author>
      <pubDate>Thu, 09 Oct 2014 21:00:00 +0000</pubDate>
      <link>https://emresahin.net/rnn-notes/</link>
      <guid isPermaLink="true">https://emresahin.net/rnn-notes/</guid>
      <description>These notes have been gathered from various sources. I provide credits and links whenever possible, but even where omitted, these are certainly not original ideas. Sequence Learning in RNNs An example of a sequence is a set of words in a sentence. Sequence learning and transformation allow comput...</description>
      <category>AI</category>
      <category>Machine Learning</category>
      <category>rnn</category>
      <category>recurrent neural networks</category>
      <category>deep learning</category>
      <category>sequence learning</category>
      <category>hmm</category>
      <category>linear dynamical systems</category>
      <content:encoded><![CDATA[<p>These notes have been gathered from various sources. I provide credits and links
whenever possible, but even where omitted, these are certainly not original
ideas.</p>
<h1 id="sequence-learning-in-rnns">Sequence Learning in RNNs</h1>
<p>An example of a sequence is a set of words in a sentence. Sequence learning and
transformation allow computers to translate one sequence into another language.</p>
<p>Alternatively, if no explicit target exists, RNNs can predict the next element
in a sequence. This type of prediction often blurs the line between supervised
and unsupervised learning.</p>
<h2 id="models-with-state">Models with State</h2>
<p>Autoregressive models calculate the current value based on previous ones:</p>
<p>$$x_t = f(x_{t-1}, x_{t-2}, \ldots)$$</p>
<p>By incorporating hidden states, it becomes much easier to perform complex
tasks:</p>
<p>$$x_t = f(h_t, x_{t-1}, x_{t-2}, \ldots)$$</p>
<p>These hidden states are typically nonlinear.</p>
<h2 id="similarity-to-quantum-mechanics">Similarity to Quantum Mechanics</h2>
<p>In Feed-Forward Neural Networks (FFNNs), the hidden state is not directly
observable. Is this similar to a quantum state?</p>
<h2 id="two-earlier-models">Two Earlier Models</h2>
<p>There are two general types of models worth mentioning.</p>
<h3 id="linear-dynamical-systems">Linear Dynamical Systems</h3>
<p>Used extensively in engineering. The system state is always linear; therefore,
Kalman filtering is utilized.</p>
<h3 id="hidden-markov-models">Hidden Markov Models</h3>
<p>Stochastic models with discrete states that store $\log(N)$ bits for $N$ states.
HMMs have efficient learning and prediction algorithms.</p>
<p>An important <strong>limitation</strong> of HMMs is their <strong>memory.</strong> They can only keep
$\log(N)$ bits of information. For a full-fledged linguistic application, we
might need at least 100 bits of state, which would require $2^{100}$ states—an
infeasible number.</p>
<h3 id="differences-between-rnns-hmms-and-ldss">Differences Between RNNs, HMMs, and LDSs</h3>
<p>Unlike HMMs, RNNs feature a distributed hidden state and complex, nonlinear
hidden units. Furthermore, they are typically deterministic.</p>
<p>RNN state behaviors can include:</p>
<ul>
<li><strong>Oscillation:</strong> Potentially useful for motor control.</li>
<li><strong>Settling to point attractors:</strong> Potentially useful for retrieving memories.</li>
<li><strong>Chaos:</strong> Generally undesirable for information processing.</li>
</ul>
<p>RNNs can learn to implement many small programs that run in parallel.</p>
<p>A significant disadvantage of RNNs: <strong>RNNs are hard to train.</strong></p>]]></content:encoded>
    </item>
  </channel>
</rss>
