<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.9.3">Jekyll</generator><link href="https://lorenlugosch.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://lorenlugosch.github.io/" rel="alternate" type="text/html" /><updated>2023-04-06T09:19:59-07:00</updated><id>https://lorenlugosch.github.io/feed.xml</id><title type="html">Loren Lugosch</title><subtitle>personal description</subtitle><author><name>Loren Lugosch</name></author><entry><title type="html">What does Hegel mean by “Reason”?</title><link href="https://lorenlugosch.github.io/posts/2021/05/hegel-reason/" rel="alternate" type="text/html" title="What does Hegel mean by “Reason”?" /><published>2021-05-15T00:00:00-07:00</published><updated>2021-05-15T00:00:00-07:00</updated><id>https://lorenlugosch.github.io/posts/2021/05/hegel-reason</id><content type="html" xml:base="https://lorenlugosch.github.io/posts/2021/05/hegel-reason/">&lt;p&gt;I finally finished reading the big and baffling &lt;a href=&quot;https://www.google.ca/books/edition/Phenomenology_of_Spirit/xOnhG9tidGsC?hl=en&amp;amp;gbpv=0&quot;&gt;&lt;em&gt;Phenomenology of Spirit&lt;/em&gt;&lt;/a&gt; by Georg Wilhelm Friedrich Hegel.&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/hegel/knives_out.jpg&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;The gist of the book is that something called “Spirit” (“Geist”) develops from simple consciousness of the “Here and Now”, to abstract concepts, to knowledge of self and the social/ethical world, to more sophisticated forms of art, science, religion, and finally &lt;em&gt;Absolute Knowledge&lt;/em&gt;. Along the way, Spirit invents big world-historical things like Stoicism, Skepticism, Christianity, and Kant. (See the German audiobook cover below, which incidentally cracks me up because it looks like something from a cult.)&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/hegel/stages.jpeg&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;For me, as for other people, the most fruitful way to get through the book turned out to be &lt;em&gt;not&lt;/em&gt; to try following the logic (if it exists!) of the text from section to section, but rather to sit back and enjoy the stream-of-consciousness of an extremely erudite man, who occasionally drops an interesting phrase or formulation that people like Marx and Sartre later picked up.&lt;sup id=&quot;fnref:sadler&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:sadler&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;Still, there was one aspect of the &lt;em&gt;Phenomenology of Spirit&lt;/em&gt; that defied my stream-of-consciousness-style reading and gave me pause: namely, the way in which Hegel defines the word “Reason” (“Vernunft”).&lt;/p&gt;

&lt;p&gt;Here are a few selections from the beginning of the “Reason” section of the book:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;[Self-consciousness as Reason] is certain that it is itself reality, or that everything actual is none other than itself… [Sec. 232, p. 139]&lt;/p&gt;
&lt;/blockquote&gt;

&lt;blockquote&gt;
  &lt;p&gt;Reason is the certainty of consciousness that it is all reality; thus does idealism express its Notion. [Sec. 233, p. 140]&lt;/p&gt;
&lt;/blockquote&gt;

&lt;blockquote&gt;
  &lt;p&gt;Reason is the certainty of being all &lt;em&gt;reality&lt;/em&gt;. [Sec. 235, p. 142]&lt;/p&gt;
&lt;/blockquote&gt;

&lt;blockquote&gt;
  &lt;p&gt;Reason, as it &lt;em&gt;immediately&lt;/em&gt; comes before us as the certainty of consciousness that is is all reality, … [Sec. 242, p. 146]&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is an odd definition for “Reason”. It sounds more like a definition of “idealism”. I think most philosophers, or AI people like me, would instead define “Reason” as “producing (or the faculty of producing) valid new assertions given other assertions”, or something like that. But Hegel is definitely very keen on his own unusual definition, as he makes a point of saying it at least four times.&lt;/p&gt;

&lt;p&gt;Yet Hegel’s definition doesn’t seem that useful or fitting even in his own book. Take, for example, this description of what Reason does:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Reason sets to work to &lt;em&gt;know&lt;/em&gt; the truth, to find in the form of a Notion that which, for ‘meaning’ and ‘perceiving’, is a Thing; i.e. it seeks to possess in thinghood the consciousness only of itself. [Sec. 240, p. 145]&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Plugging in “the certainty of consciousness that it is all reality” for “Reason” does not really seem to work here:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;[The certainty of consciousness that it is all reality] sets to work to &lt;em&gt;know&lt;/em&gt; the truth…&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Or here, in which Hegel starts to take one of his many potshots at “sound common sense”:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Since self-consciousness knows itself to be a moment of the &lt;em&gt;being-for-self&lt;/em&gt; of this substance, it expresses the existence of the law within itself as follows: sound Reason knows immediately what is right and good. [Sec. 422, p. 253]&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;$\rightarrow$&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;sound [certainty of consciousness that it is all reality] knows immediately what is right and good (?)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;One possibility is that Hegel is not actually using “the certainty of consciousness that it is all reality” as a &lt;em&gt;definition&lt;/em&gt; of “Reason”. Instead, maybe he just wants to firmly establish the notion that “consciousness is all reality”, and this advanced stage of the development of Spirit in his book seemed like a good place to do it. We &lt;em&gt;know&lt;/em&gt; what Reason is; so instead of wasting the reader’s time giving a “definition” they already have, why not use the gap where a definition would normally go as a free space to stick an idea he wants to make sure we share with him? Hegel seems to pull &lt;a href=&quot;https://www.youtube.com/watch?v=N5iKYhnPpV4&quot;&gt;many such tricks&lt;/a&gt; in this book.&lt;/p&gt;

&lt;p&gt;Another possibility is that when Hegel says “Reason &lt;em&gt;is&lt;/em&gt; the certainty of being all reality”, the “&lt;em&gt;is&lt;/em&gt;” is not indicating &lt;em&gt;identity&lt;/em&gt; but rather just marking a &lt;em&gt;predicate&lt;/em&gt;: that is, “the certainty of being all reality” is &lt;em&gt;one&lt;/em&gt; aspect of the thing Hegel calls “Reason”, but not the complete definition (which he never provides). But the predicate of “the certainty of being all reality” still does not really jive with our commonsensical notion of “Reason”. So maybe Hegel is just casually asserting that “&lt;em&gt;if&lt;/em&gt; you’re a thinking person, a &lt;em&gt;reasonable&lt;/em&gt; German intellectual in the year 1807, and you’ve read your Kant, your Fichte, and your Schelling, then &lt;em&gt;of course&lt;/em&gt; you’re an idealist and you already &lt;em&gt;know&lt;/em&gt; that consciousness is all reality”.&lt;/p&gt;

&lt;p&gt;A third possibility is that “Vernunft” has overtones in German that “Reason” does not have in English and that could make Hegel’s odd usage more sensible. This is at least true of “Geist” (cognate with English “ghost”), which we could translate in English as “Spirit” (as is done for Hegel), “Mind”, or maybe “intellectual activity” (“humanities” in German is “Geisteswissenschaften”, in contrast to “Naturwissenschaften”, “natural science”). My German isn’t good enough for me to know whether something similar can be said of “Vernunft”.&lt;/p&gt;

&lt;p&gt;(A fourth possibility is that Hegel is full of shit and we should just put the book down. In math, this is called “&lt;a href=&quot;https://en.wikipedia.org/wiki/Triviality_(mathematics)#Trivial_and_nontrivial_solutions&quot;&gt;the trivial solution&lt;/a&gt;”.)&lt;/p&gt;

&lt;p&gt;It’s also interesting to look at how Hegel uses “Reason” (“Vernunft”) in some of his later writings. In the preface&lt;sup id=&quot;fnref:owl&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:owl&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;2&lt;/a&gt;&lt;/sup&gt; to his 1820 &lt;em&gt;Philosophy of Right&lt;/em&gt; Hegel attributes an idea about Reality and Reason to Plato:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Yet [Plato] has proved himself to be a great mind because the very principle and central distinguishing feature of his idea is the pivot upon which the world-wide revolution then in process turned: What is rational is real; And what is real is rational. (Was vernünftig ist, das ist wirklich; und was wirklich ist, das ist vernünftig.)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Another famous instance is his phrase “the Cunning of Reason” (“die List der Vernunft”) in his &lt;em&gt;Lectures on the Philosophy of History&lt;/em&gt;. The idea is that individual humans who think they are simply acting in their own self-interest, with no regard for whatever shenanigans Spirit is trying to pull with world history at the moment, are, by being &lt;em&gt;rational&lt;/em&gt;, unwittingly carrying out what &lt;em&gt;Reason&lt;/em&gt; wants to have happen in the world:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;The particular interest linked to passion is thus inseparable from the actualization of the universal principle; for the universal is the outcome of the particular and determinate, and from its negation. … This may be called the &lt;em&gt;Cunning of Reason&lt;/em&gt;, that it allows the passions to work for it, while what it brings into existence suffers loss and injury.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Reason as logic; Reason as idealism; Reason as Spirit moving history forward by tricking us rational peasants into doing what it wants. Is there some simple notion underlying these conceptions for Hegel? You’d probably have to get a philosophy degree to prove it. As a programmer, polysemy frustrates me. As a poet, it fascinates me.&lt;/p&gt;

&lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
  &lt;ol&gt;
    &lt;li id=&quot;fn:sadler&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;For a more systematic reading by someone who is &lt;em&gt;not&lt;/em&gt; an amateur like me, I recommend &lt;a href=&quot;https://www.youtube.com/playlist?list=PL4gvlOxpKKIgR4OyOt31isknkVH2Kweq2&quot;&gt;Gregory Sadler’s “Half Hour Hegel”&lt;/a&gt;, a years-long lecture series in which Sadler attempts to explain or interpret &lt;em&gt;every section&lt;/em&gt; of the book. I’m not that patient! &lt;a href=&quot;#fnref:sadler&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:owl&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;At the end of that preface, Hegel gives a beautiful metaphor for his conception of philosophy: &lt;em&gt;“When philosophy paints its grey in grey, one form of life has become old, and by means of grey it cannot be rejuvenated, but only known. The owl of Minerva takes its flight only when the shades of night are gathering.” (“Wenn die Philosophie ihr Grau in Grau malt, dann ist eine Gestalt des Lebens alt geworden, und mit Grau in Grau läßt sie sich nicht verjüngen, sondern nur erkennen; die Eule der Minerva beginnt erst mit der einbrechenden Dämmerung ihren Flug.”)&lt;/em&gt; &lt;a href=&quot;#fnref:owl&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
&lt;/div&gt;</content><author><name>Loren Lugosch</name></author><summary type="html">I finally finished reading the big and baffling Phenomenology of Spirit by Georg Wilhelm Friedrich Hegel.</summary></entry><entry><title type="html">First PC build</title><link href="https://lorenlugosch.github.io/posts/2021/03/pc/" rel="alternate" type="text/html" title="First PC build" /><published>2021-03-03T00:00:00-08:00</published><updated>2021-03-03T00:00:00-08:00</updated><id>https://lorenlugosch.github.io/posts/2021/03/pc</id><content type="html" xml:base="https://lorenlugosch.github.io/posts/2021/03/pc/">&lt;p&gt;My everyday work computer is a MacBook Pro that I’ve had since 2013. It’s a great machine and continues to serve me well, but I was moved in a moment of pandemic malaise to treat myself to a little upgrade.&lt;/p&gt;

&lt;p&gt;I decided to build (for the first time!) a desktop PC instead of getting another laptop, for a few reasons:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;I’m working from home now, and I’ll probably have to for the foreseeable future.&lt;/li&gt;
  &lt;li&gt;As a holder of electrical and computer engineering degrees, I felt a little embarrassed that I had never built my own computer. (Cf. the &lt;a href=&quot;https://www.imdb.com/title/tt0582462/&quot;&gt;engine repair episode of &lt;em&gt;Frasier&lt;/em&gt;&lt;/a&gt;.)&lt;/li&gt;
  &lt;li&gt;I wanted to try modern PC games that require a little more horsepower than &lt;em&gt;Undertale&lt;/em&gt; and that I can’t play on my Nintendo Switch.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So here’s how my first PC build went.&lt;/p&gt;

&lt;h2 id=&quot;the-gpu&quot;&gt;The GPU&lt;/h2&gt;

&lt;p&gt;As you might guess from my website’s favicon, and my all-around Stack More Layers demeanor, I started with the GPU.&lt;/p&gt;

&lt;p&gt;I wanted a not-too-expensive GPU that I could use both for gaming and for training reasonably large neural nets. (I do have access to lots of powerful GPUs in the Mila cluster, but it’s nice not to have to share with other people sometimes, and to be able to leave a model training for, like, a month.)&lt;/p&gt;

&lt;p&gt;The best option looked to be the Nvidia (&lt;a href=&quot;https://en.wikipedia.org/wiki/Talk%3ANvidia#Naming_Conventions&quot;&gt;nVidia? NVIDIA?&lt;/a&gt;) &lt;a href=&quot;https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3080/&quot;&gt;RTX 3080&lt;/a&gt;. But whether due to pandemic disruptions, or Bitcoin bros, or BERT bros, or something else, it’s off shelves everywhere.&lt;/p&gt;

&lt;p&gt;I waited a month to see if the RTX 3080 would come back in stock anywhere, but no luck. Instead, I consulted Tim Dettmers’s magisterial &lt;a href=&quot;https://timdettmers.com/2020/09/07/which-gpu-for-deep-learning/&quot;&gt;deep learning GPU guide&lt;/a&gt; and found that, among not-unavailable GPUs, the &lt;a href=&quot;https://www.nvidia.com/en-us/geforce/products/10series/ultimate-4k/&quot;&gt;GTX 1080 Ti&lt;/a&gt; had the best performance per dollar, and good performance for training transformer models. So I picked one up from a guy nearby in Montreal off Kijiji.&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/pc/gpu.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h2 id=&quot;the-rest&quot;&gt;The rest&lt;/h2&gt;

&lt;p&gt;Using &lt;a href=&quot;https://pcpartpicker.com/builds/&quot;&gt;PCPartPicker’s Completed Builds&lt;/a&gt; search feature, I found a &lt;a href=&quot;https://pcpartpicker.com/b/BPgJ7P&quot;&gt;build&lt;/a&gt; for a PC based around the RTX 3080 and followed it to the letter. I probably could have optimized the components more for what I want, but having not ever built a PC before, I didn’t want to accidentally select components that were incompatible with each other. (PCPartPicker does have a compatibility checker, but I don’t know how idiot-proof it is.)&lt;/p&gt;

&lt;p&gt;Then I bought all the components, put it together, and breathed a sigh of relief when I pushed the power button for the first time and the fans turned on.&lt;/p&gt;

&lt;p&gt;The build process felt like assembling a LEGO (Lego?) set, but with the instructions distributed across multiple boxes and the additional stress of unfamiliar sensitive electronic components that would cost a lot of money to replace.&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/pc/build.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;I started by reading the motherboard manual, which told me to put in the CPU. That was a little nerve-wracking because I really had to pull the little whammy bar (or so I’m inclined to call it) &lt;em&gt;tight&lt;/em&gt; to get the chip in, which felt wrong.&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/pc/cpu.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;Then I installed the fan, with the help of a &lt;a href=&quot;https://www.youtube.com/watch?v=Ascz5P0-jyU&quot;&gt;video guide&lt;/a&gt;. I wasn’t even sure what the orientation of the fan within the case was supposed to be, so I tried to infer it from a photo taken by the creator of the build. Because I hadn’t put the motherboard into the case yet and had to operate solely on other landmarks, this was mystifying, until I realized that the photo showed a machine with glowy RAM sticks instead of the non-glowy RAM I had. Thus I learned that people use RAM that &lt;em&gt;glows&lt;/em&gt;.&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/pc/vidicus_build.jpg&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;The PC case itself was a complex enough beast that I had to follow yet another (thorough, excellent) &lt;a href=&quot;https://www.youtube.com/watch?v=2mMBA4lzDvo&quot;&gt;video guide&lt;/a&gt; on how to get it open and wire things up to the motherboard.&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/pc/complete.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;It’s alive! And it &lt;em&gt;glows&lt;/em&gt;. But the RGB effects are a little distracting, so I’ll probably deactivate them. Besides, I don’t need them to feel like a real &lt;em&gt;Gamer&lt;/em&gt;. As it is written: &lt;em&gt;“When you Game, do not be like the hypocrites, for they love to be seen as Gamers with their showy RGB effects. Instead, when you Game, go into your room, close the door, and Game in secret.”&lt;/em&gt;&lt;/p&gt;

&lt;h2 id=&quot;some-extras&quot;&gt;Some extras&lt;/h2&gt;

&lt;p&gt;The build I followed did not include any peripheral devices, so I picked some up, again using PCPartPicker. I’ve never been picky about keyboards, so I just filtered for low prices, sorted by rating, and picked the top one. Same for the monitor. For the mouse, I bought this nice glowy one with some sort of fangly beast on it. More RGB!&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/pc/mouse.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;Also, apparently, motherboards cannot connect to the Internet without help! As I discovered when I turned on the PC. (Stop laughing!) I don’t have an easily-reachable Ethernet cable, so I had to wait a couple days more for a Wi-Fi PCIe card to arrive before I could actually use the computer.&lt;/p&gt;

&lt;p&gt;Incidentally, the package for the Wi-Fi card featured maybe my favorite acronym ever:&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/pc/fast.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h2 id=&quot;testing-er-out&quot;&gt;Testing ‘er out&lt;/h2&gt;

&lt;p&gt;For my first game, I picked up &lt;em&gt;Star Wars: Squadrons&lt;/em&gt;, the recently released spiritual successor to the old &lt;em&gt;Rogue Squadron&lt;/em&gt; series of dogfighting games. I’m bad at it. But it’s fun! And it runs great on my new machine. It’s very pretty, but I haven’t looked up how to do a screenshot yet, so here’s a picture swiped from the Wikipedia article on the game.&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/pc/squadrons.jpg&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;I also downloaded the original &lt;em&gt;Half-Life&lt;/em&gt;, a game I bought on Steam a while back that my MacBook rudely made unplayable with the Catalina macOS update.&lt;/p&gt;

&lt;p&gt;Yay! Almost done. The last step for me is getting Linux dual-booted so that I can use all my machine learning tools.&lt;/p&gt;</content><author><name>Loren Lugosch</name></author><summary type="html">My everyday work computer is a MacBook Pro that I’ve had since 2013. It’s a great machine and continues to serve me well, but I was moved in a moment of pandemic malaise to treat myself to a little upgrade.</summary></entry><entry><title type="html">End-to-end models falling short on SLURP</title><link href="https://lorenlugosch.github.io/posts/2020/12/slurp/" rel="alternate" type="text/html" title="End-to-end models falling short on SLURP" /><published>2020-12-19T00:00:00-08:00</published><updated>2020-12-19T00:00:00-08:00</updated><id>https://lorenlugosch.github.io/posts/2020/12/slurp</id><content type="html" xml:base="https://lorenlugosch.github.io/posts/2020/12/slurp/">&lt;p&gt;There’s a hot new dataset for spoken language understanding: &lt;a href=&quot;https://www.aclweb.org/anthology/2020.emnlp-main.588.pdf&quot;&gt;SLURP&lt;/a&gt;. I’m excited about SLURP for a couple reasons:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;It’s way bigger and more challenging than existing open-source SLU datasets:&lt;/li&gt;
&lt;/ul&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/slurp/dataset-info.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;The authors found that an end-to-end model did not work:&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
  &lt;p&gt;We have tested several SOTA E2E-SLU systems on SLURP, including (Lugosch et al., 2019b) which produces SOTA results on the FSC corpus. However, re-training these models on this more complex domain did not converge or result in meaningful outputs. Note that these models were developed to solve much easier tasks (e.g. a single domain).&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Having &lt;a href=&quot;https://lorenlugosch.github.io/posts/2020/12/slu/&quot;&gt;just written&lt;/a&gt; a SpeechBrain recipe for end-to-end SLU with my much simpler Timers and Such dataset, I thought I’d give that recipe a whirl on SLURP. The model used in the recipe looks like this:&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/slu/direct.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;The output dictionaries for SLURP look a bit different from the ones in Timers and Such or Fluent Speech Commands:&lt;/p&gt;
&lt;pre&gt;&lt;code style=&quot;font-size:14px&quot;&gt;
{
  &quot;scenario&quot;: &quot;alarm&quot;,
  &quot;action&quot;: &quot;query&quot;,
  &quot;entities&quot;: [
    {&quot;type&quot;: &quot;event_name&quot;, &quot;filler&quot;: &quot;dance class&quot;}
  ]
}
&lt;/code&gt;
&lt;/pre&gt;

&lt;p&gt;But—much like a honey badger—our autoregressive sequence-to-sequence model doesn’t care. It just generates the dictionary character-by-character, no matter what the format looks like.&lt;/p&gt;

&lt;p&gt;The authors of the SLURP paper provide baseline results with a more task-specific model, &lt;a href=&quot;https://www.aclweb.org/anthology/W19-5931.pdf&quot;&gt;HerMiT&lt;/a&gt;, which uses an ASR model to predict a transcript and a conditional random field for each of (scenario, action, entities). The outputs are generated using the Viterbi algorithm.&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/slurp/hermit.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;The authors kindly &lt;a href=&quot;https://github.com/pswietojanski/slurp&quot;&gt;released&lt;/a&gt; their tool for computing performance metrics, including a new “SLU-F1” metric they propose. I used this tool and got the following results &lt;em&gt;(EDIT: I left the model training a little longer and updated the numbers EDIT EDIT: I got rid of a coverage penalty term in the beam search and updated the numbers again)&lt;/em&gt;:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Model&lt;/th&gt;
      &lt;th&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;scenario&lt;/code&gt; (accuracy)&lt;/th&gt;
      &lt;th&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;action&lt;/code&gt; (accuracy)&lt;/th&gt;
      &lt;th&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;intent&lt;/code&gt; (accuracy)&lt;/th&gt;
      &lt;th&gt;Word-F1&lt;/th&gt;
      &lt;th&gt;Char-F1&lt;/th&gt;
      &lt;th&gt;SLU-F1&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;End-to-end&lt;/td&gt;
      &lt;td&gt;&lt;s&gt;71.17&lt;/s&gt;&lt;br /&gt;81.73&lt;/td&gt;
      &lt;td&gt;&lt;s&gt;65.43&lt;/s&gt;&lt;br /&gt;77.11&lt;/td&gt;
      &lt;td&gt;&lt;s&gt;61.85&lt;/s&gt;&lt;br /&gt;75.05&lt;/td&gt;
      &lt;td&gt;&lt;s&gt;45.57&lt;/s&gt;&lt;br /&gt;61.24&lt;/td&gt;
      &lt;td&gt;&lt;s&gt;49.23&lt;/s&gt;&lt;br /&gt;65.42&lt;/td&gt;
      &lt;td&gt;&lt;s&gt;47.33&lt;/s&gt;&lt;br /&gt;63.26&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;HerMiT&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;85.69&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;81.42&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;78.33&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;69.34&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;72.39&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;70.83&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Turns out HerMiT does way better than our end-to-end model. This isn’t totally surprising because HerMiT has a lot of structure built in that our autoregressive model has to learn from scratch. For instance, I don’t think it’s possible for their model to output &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;&quot;brussia&quot;&lt;/code&gt; (one of the more charming output mistakes I noticed early in training).&lt;/p&gt;
&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/slurp/brussia.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;
Another difference is that their ASR model is trained on Multi-ASR, a massive &lt;strong&gt;24,000 hour&lt;/strong&gt; dataset formed by Captain-Planeting the LibriSpeech, Switchboard, Fisher, CommonVoice, AMI, and ICSI datasets—whereas my encoder is only pre-trained using the 1,000 hours of LibriSpeech.&lt;/p&gt;

&lt;p&gt;So: how much of the gap is due to more/better audio? The authors also report some results when applying HerMiT to the gold transcripts instead of the ASR output; similarly, we can feed the gold transcripts into our sequence-to-sequence model instead of audio and compare the results &lt;em&gt;(EDIT again, I let the model train a bit longer/got rid of the pesky coverage penalty and updated the numbers)&lt;/em&gt;:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Model&lt;/th&gt;
      &lt;th&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;scenario&lt;/code&gt; (accuracy)&lt;/th&gt;
      &lt;th&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;action&lt;/code&gt; (accuracy)&lt;/th&gt;
      &lt;th&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;intent&lt;/code&gt; (accuracy)&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;End-to-end (text input, gold transcripts)&lt;/td&gt;
      &lt;td&gt;&lt;s&gt;89.91&lt;/s&gt;&lt;br /&gt; &lt;strong&gt;90.81&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;s&gt;86.54&lt;/s&gt;&lt;br /&gt; &lt;strong&gt;88.29&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;s&gt;85.43&lt;/s&gt;&lt;br /&gt; &lt;strong&gt;87.28&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;HerMiT (gold transcripts)&lt;/td&gt;
      &lt;td&gt;90.15&lt;/td&gt;
      &lt;td&gt;86.99&lt;/td&gt;
      &lt;td&gt;84.84&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Here the simple sequence-to-sequence model actually does about as well as HerMiT &lt;em&gt;(EDIT actually a bit better!)&lt;/em&gt;. This suggests that the audio side of things is more where our problems lie.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;In summary,&lt;/em&gt; I have attempted to defend the honor of end-to-end models: we can indeed train one on SLURP and get semi-reasonable outputs. Note, though, that I’ve done no hyperparameter tuning on the model (except to increase the number of training epochs and getting rid of the coverage penalty term), so it’s possible we could do better with a little elbow grease—maybe starting by swapping out the now-unfashionable RNNs I used in the encoder and decoder with ✨Transformers✨.&lt;/p&gt;</content><author><name>Loren Lugosch</name></author><category term="sequence modeling" /><summary type="html">There’s a hot new dataset for spoken language understanding: SLURP. I’m excited about SLURP for a couple reasons:</summary></entry><entry><title type="html">Siri from scratch! (Not really.)</title><link href="https://lorenlugosch.github.io/posts/2020/12/slu/" rel="alternate" type="text/html" title="Siri from scratch! (Not really.)" /><published>2020-12-11T00:00:00-08:00</published><updated>2020-12-11T00:00:00-08:00</updated><id>https://lorenlugosch.github.io/posts/2020/12/slu</id><content type="html" xml:base="https://lorenlugosch.github.io/posts/2020/12/slu/">&lt;p&gt;I make fairly heavy use of the voice assistant on my phone for things like setting timers while cooking. As a result, when I spent some time this summer at my in-laws’ place—where there was no cell signal and not-very-good Wi-Fi—I often tried using Siri only to get a sad little “sorry, no Internet :(“ response. (#FirstWorldProblems.)&lt;/p&gt;

&lt;p&gt;This reminded me of a tweet I saw a while back:&lt;/p&gt;

&lt;blockquote class=&quot;twitter-tweet&quot;&gt;&lt;p lang=&quot;en&quot; dir=&quot;ltr&quot;&gt;like 90% of my voice assistant usage is setting timers, math, and unit conversions&lt;br /&gt;&lt;br /&gt;can we ship an offline on-device model that does this well and nothing else?&lt;/p&gt;&amp;mdash; Arkadiy Kukarkin (@parkan) &lt;a href=&quot;https://twitter.com/parkan/status/1119334813960429569?ref_src=twsrc%5Etfw&quot;&gt;April 19, 2019&lt;/a&gt;&lt;/blockquote&gt;
&lt;script async=&quot;&quot; src=&quot;https://platform.twitter.com/widgets.js&quot; charset=&quot;utf-8&quot;&gt;&lt;/script&gt;

&lt;p&gt;Now, it just so happens that I 1) am a big fan of doing things offline rather than in the cloud, 2) have some experience training speech models, and 3) wanted to procrastinate. (&lt;em&gt;The perfect storm.&lt;/em&gt;) So here’s the record of me taking a crack at it.&lt;/p&gt;

&lt;h2 id=&quot;the-goal&quot;&gt;The goal&lt;/h2&gt;
&lt;p&gt;The goal is to make a box that can take as input the speech signal and output what the speaker wants. The output takes the form of a dictionary containing the semantics of the utterance&lt;sup id=&quot;fnref:RL&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:RL&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;—i.e. the intent, slots, and slot values, and maybe some other fancy things like Named Entities. I shall casually refer to this output simply as the “intent”, a common synecdoche. This is called “spoken language understanding” (SLU).&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/slu/SLU.png&quot; style=&quot;max-width:100%&quot; /&gt;&lt;/center&gt;

&lt;h2 id=&quot;the-decoupled-approach&quot;&gt;The “decoupled” approach&lt;/h2&gt;

&lt;p&gt;The first thing I tried was a straightforward “decoupled” approach: I trained a general-purpose automatic speech recognition (ASR) part and a separate domain-specific natural language understanding (NLU) part.&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/slu/decoupled.png&quot; style=&quot;max-width:100%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;
To train the ASR part, I used LibriSpeech, a big dataset of American English speakers reading audiobooks. (Or rather, my SpeechBrain colleagues did, and I just loaded their best model checkpoint. Thanks, Ju-Chieh + Mirco + Abdel + Peter!)&lt;/p&gt;

&lt;p&gt;To train the NLU part, we need text labeled with intents. For this, I wrote a script to generate a bunch of labeled random phrases for the four types of commands I wanted: setting timers, converting units (length, volume, temperature), setting alarms, and simple math. Here’s a few examples:&lt;/p&gt;

&lt;pre&gt;&lt;code style=&quot;font-size:14px&quot;&gt;
(&quot;how many inches are there in 256 centimeters&quot;,
{
  'intent': 'UnitConversion', 
  'slots': {
    'unit1': 'centimeter', 
    'unit2': 'inch', 
    'amount': 256
  }
})

(&quot;set my alarm for 8:03AM&quot;,
{
  'intent': 'SetAlarm', 
  'slots': {
    'am_or_pm': 'AM', 
    'alarm_hour': 8, 
    'alarm_minute': 3
  }
})

(&quot;what's 37.67 minus 75.7&quot;,
{
  'intent': 'SimpleMath', 
  'slots': {
    'number1': 37.67, 
    'number2': 75.7, 
    'op': ' minus '
  }
})
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;I trained an attention model to ingest the transcript and autoregressively predict these dictionaries as strings, one character at a time.&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/slu/nlu.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;
This model gets perfect accuracy on the test transcripts, which makes sense, since this is a pretty simple domain.&lt;/p&gt;

&lt;p&gt;Does that mean our system as a whole will get perfect accuracy on test audio? No, because there is the chance that the ASR part will incorrectly transcribe the input and that the NLU part will fail as a result. To measure how well our system works, we need some actual in-domain audio data.&lt;/p&gt;

&lt;h3 id=&quot;some-end-to-end-test-and-training-data&quot;&gt;Some end-to-end test (and training) data&lt;/h3&gt;

&lt;p&gt;So I recorded a few friends and colleagues speaking the generated prompts. My recordees kindly gave me their consent to release their anonymized recordings, which you can find &lt;a href=&quot;https://zenodo.org/record/4110812&quot;&gt;here&lt;/a&gt;. I manually segmented and cleaned (= fixed the label, when someone misspoke) their recordings, yielding me a modest 271 audios. (This exercise builds character. I highly recommend it! &lt;a href=&quot;https://karpathy.github.io/2019/04/25/recipe/&quot;&gt;So does Andrej Karpathy&lt;/a&gt;.)&lt;/p&gt;

&lt;p&gt;I split the recordings into train/dev/test sets so that each speaker was in only one of the sets, with:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;144 audios (4 speakers) for the train set (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;train-real&lt;/code&gt;),&lt;/li&gt;
  &lt;li&gt;72 audios (2 speakers) for the dev set (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;dev-real&lt;/code&gt;), and&lt;/li&gt;
  &lt;li&gt;55 audios (5 speakers) for the test set (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;test-real&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;(For the decoupled approach, we don’t need audio during training, but we will shortly for an alternative approach.)&lt;/p&gt;

&lt;p&gt;It’s hard to get meaningful accuracy estimates with only 55 test examples, so I also generated a bunch of &lt;em&gt;synthetic&lt;/em&gt; audio (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;train-synth&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;dev-synth&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;test-synth&lt;/code&gt;) by synthesizing all my NLU training text data with Facebook’s &lt;a href=&quot;https://github.com/facebookarchive/loop/&quot;&gt;VoiceLoop&lt;/a&gt; text-to-speech model.&lt;sup id=&quot;fnref:TTS&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:TTS&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;2&lt;/a&gt;&lt;/sup&gt; Like I did for the real speakers, I split the 22 synthetic speakers into train/dev/test sets with no speaker overlap.&lt;/p&gt;

&lt;p&gt;In total, I generated around 200,000 audios. I couldn’t figure out how to parallelize the generation process—with simple multithreading I ran into some concurrency problems with the vocoder tool called by VoiceLoop—so I just did it sequentially, which took more than a week to finish.&lt;/p&gt;

&lt;p&gt;Lucky for you, I’ve uploaded both the real speech and synthesized speech &lt;a href=&quot;https://zenodo.org/record/4110812&quot;&gt;here&lt;/a&gt;! I call the complete dataset &lt;strong&gt;Timers and Such v0.1&lt;/strong&gt;. It’s v0.1 because the test set of real speakers is probably too small for it to be used to meaningfully compare results from different approaches. It would be nice to scale this up to more real speakers and make a v1.0 that people can actually use for R&amp;amp;D.&lt;/p&gt;

&lt;h3 id=&quot;first-results&quot;&gt;First results&lt;/h3&gt;

&lt;p&gt;I tested the system out on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;test-real&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;test-synth&lt;/code&gt; and measured the overall accuracy (= if any slot is wrong, the whole utterance is considered wrong). Here’s the results, averaged over 5 seeds:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Model&lt;/th&gt;
      &lt;th&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;test-real&lt;/code&gt;&lt;/th&gt;
      &lt;th&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;test-synth&lt;/code&gt;&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Decoupled&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;23.6%&lt;/strong&gt; $\pm$ 7.3%&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;18.7%&lt;/strong&gt; $\pm$ 5.1%&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Wow, it’s only getting the answer right a fifth of the time. That sucks! What could the problem be?&lt;/p&gt;

&lt;p&gt;Let’s look at some of the ASR outputs for &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;test-real&lt;/code&gt;:&lt;/p&gt;

&lt;hr /&gt;
&lt;pre&gt;
&lt;code style=&quot;font-size:14px&quot;&gt;
True transcript: &quot;SET A TIMER FOR EIGHT MINUTES&quot;
ASR transcript: &quot;SAID A TIMID FOR EIGHT MINUTES&quot;

True transcript: &quot;HOW MANY TEASPOONS ARE THERE IN SIXTY SEVEN TABLESPOONS&quot;
ASR transcript: &quot;AWMAN IN TEASPOILS ADHERED IN SEVEN ESSAYS SEVERN TABLESPOONS&quot;
&lt;/code&gt;
&lt;/pre&gt;
&lt;hr /&gt;

&lt;p&gt;Now, this ASR model gets a WER of 3% (&lt;a href=&quot;https://paperswithcode.com/sota/speech-recognition-on-librispeech-test-clean&quot;&gt;not SOTA, but good&lt;/a&gt;) on the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;test-clean&lt;/code&gt; subset of LibriSpeech. So why do the outputs look so bad here?&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;
    &lt;p&gt;One issue is &lt;strong&gt;accent mismatch&lt;/strong&gt;. LibriSpeech has only American English speakers, whereas only 3 of 11 real speakers in Timers and Such have American accents. There’s not much we could do about this, save re-train the ASR model on a more diverse set of accents.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;The other issue is &lt;strong&gt;language model (LM) mismatch&lt;/strong&gt;. The ASR model has an LM trained on LibriSpeech’s LM text data. That text data comes from Project Gutenberg books, which is a very different domain: for instance, “SAID” is more likely to appear at the beginning of a transcript from a book than “SET”.&lt;/p&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So I trained an LM on the Timers and Such transcripts, and used that as the ASR model’s LM instead of the LibriSpeech LM. This ends up fixing a lot of ASR mistakes, but not all of them.&lt;sup id=&quot;fnref:FST&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:FST&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Model&lt;/th&gt;
      &lt;th&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;test-real&lt;/code&gt;&lt;/th&gt;
      &lt;th&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;test-synth&lt;/code&gt;&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Decoupled (LibriSpeech LM)&lt;/td&gt;
      &lt;td&gt;23.6% $\pm$ 7.3%&lt;/td&gt;
      &lt;td&gt;18.7% $\pm$ 5.1%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Decoupled (Timers and Such LM)&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;44.4%&lt;/strong&gt; $\pm$ 6.9%&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;31.9%&lt;/strong&gt; $\pm$ 3.9%&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;h2 id=&quot;sort-of-end-to-end-training-the-multistage-approach&quot;&gt;Sort-of end-to-end training: the “multistage” approach&lt;/h2&gt;

&lt;p&gt;Instead of training on the true transcript, we might want to train the NLU model using the ASR transcript, to make the NLU part more resilient to ASR errors. Google &lt;a href=&quot;https://arxiv.org/abs/1809.09190&quot;&gt;calls this&lt;/a&gt; a “multistage” end-to-end SLU model: it still uses distinct ASR and NLU parts, but the complete system is trained on audio data.&lt;/p&gt;

&lt;!-- [^decoupled]: There are some [software engineering benefits](https://papers.nips.cc/paper/2015/file/86df7dcfd896fcaf2674f757a2463eba-Paper.pdf) to _not_ doing things end-to-end. By separating the problem into two stages, the ASR people do not need to care about the intent structure and can focus on optimizing word error rate, and the NLU people do not need to care about sampling rates and FFTs and can focus on optimizing semantic accuracy. --&gt;

&lt;p&gt;Running this experiment, multistage training works a lot better than the decoupled approach, with or without an appropriate LM.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Model&lt;/th&gt;
      &lt;th&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;test-real&lt;/code&gt;&lt;/th&gt;
      &lt;th&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;test-synth&lt;/code&gt;&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Decoupled (LibriSpeech LM)&lt;/td&gt;
      &lt;td&gt;23.6% $\pm$ 7.3%&lt;/td&gt;
      &lt;td&gt;18.7% $\pm$ 5.1%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Decoupled (Timers and Such LM)&lt;/td&gt;
      &lt;td&gt;44.4% $\pm$ 6.9%&lt;/td&gt;
      &lt;td&gt;31.9% $\pm$ 3.9%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Multistage (LibriSpeech LM)&lt;/td&gt;
      &lt;td&gt;69.8% $\pm$ 3.5%&lt;/td&gt;
      &lt;td&gt;69.9% $\pm$ 2.5%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Multistage (Timers and Such LM)&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;75.3%&lt;/strong&gt; $\pm$ 4.2%&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;73.1%&lt;/strong&gt; $\pm$ 8.7%&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;h2 id=&quot;for-real-end-to-end-training-the-direct-approach&quot;&gt;For-real end-to-end training: the “direct” approach&lt;/h2&gt;

&lt;p&gt;Still, there’s a few disadvantages to the multistage ASR-NLU approach.&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;We need an intermediate search step to predict a transcript during training. The search is inherently sequential and ends up being the slowest part of training.&lt;/li&gt;
  &lt;li&gt;It’s difficult&lt;sup id=&quot;fnref:backprop&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:backprop&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;4&lt;/a&gt;&lt;/sup&gt; to backpropagate through the discrete search into the encoder, so the model can’t learn to give more priority to recognizing words that are more relevant to the SLU task, as opposed to less informative words like “the” and “please”.&lt;/li&gt;
  &lt;li&gt;Ultimately, we don’t actually care about the transcript for this application: we just want the intent. By predicting the transcript, we’re wasting FLOPs.&lt;sup id=&quot;fnref:non-transcript&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:non-transcript&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So: why not train a model to just map directly from speech to intent?&lt;sup id=&quot;fnref:direct&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:direct&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;6&lt;/a&gt;&lt;/sup&gt; To quote Vapnik: &lt;em&gt;“When solving a problem of interest, do not solve a more general problem as an intermediate step.”&lt;/em&gt;&lt;sup id=&quot;fnref:vapnik&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:vapnik&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;7&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/slu/direct.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;
The direct approach faces an additional difficulty: it has to learn what speech sounds like from scratch. To make the comparison with the ASR-based models more fair, we can use transfer learning. This can be done simply by popping the encoder out of the pre-trained LibriSpeech ASR model and using it as a feature extractor in the SLU model.&lt;sup id=&quot;fnref:pretrain&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:pretrain&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;8&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;We get pretty good results with this approach: around the same performance as the multistage model with an appropriate language model (slightly worse, within a standard deviation). The direct model also trained a lot faster: 1h 42m for one epoch (running on a beastly Quadro RTX 8000), compared with 2h 53m for the multistage model on the same machine.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Model&lt;/th&gt;
      &lt;th&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;test-real&lt;/code&gt;&lt;/th&gt;
      &lt;th&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;test-synth&lt;/code&gt;&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Decoupled (LibriSpeech LM)&lt;/td&gt;
      &lt;td&gt;23.6% $\pm$ 7.3%&lt;/td&gt;
      &lt;td&gt;18.7% $\pm$ 5.1%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Decoupled (Timers and Such LM)&lt;/td&gt;
      &lt;td&gt;44.4% $\pm$ 6.9%&lt;/td&gt;
      &lt;td&gt;31.9% $\pm$ 3.9%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Multistage (LibriSpeech LM)&lt;/td&gt;
      &lt;td&gt;69.8% $\pm$ 3.5%&lt;/td&gt;
      &lt;td&gt;69.9% $\pm$ 2.5%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Multistage (Timers and Such LM)&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;75.3%&lt;/strong&gt; $\pm$ 4.2%&lt;/td&gt;
      &lt;td&gt;73.1% $\pm$ 8.7%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Direct&lt;/td&gt;
      &lt;td&gt;74.5% $\pm$ 6.9%&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;96.1%&lt;/strong&gt; $\pm$ 0.2%&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Notice anything interesting? The direct model performs a lot better on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;test-synth&lt;/code&gt; than the other models. This makes sense: the direct model has access to the raw speech features, so it can learn the idiosyncrasies of the speech synthesizer and recognize the synthetic test speech more easily. (Of course, we don’t care about the performance on synthetic speech; we only care about how well this works for human speakers.)&lt;/p&gt;

&lt;h2 id=&quot;are-we-done&quot;&gt;Are we done?&lt;/h2&gt;
&lt;p&gt;Did we achieve our goal from the outset of “an offline on-device model that [recognizes numeric commands] well and nothing else”?&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;em&gt;As for “on-device”:&lt;/em&gt; The model has 170 million parameters, most of which live in the pre-trained ASR part. This requires 680MB of storage using single-precision floats; for reference, right now the maximum app download size for Android is 100MB, so we would probably have a hard time putting this on a phone. We would need to fiddle with the ASR hyperparameters a bit to shrink the encoder, but this is definitely doable. In fact, the direct SLU models I trained in my previous papers had a little over 1 million parameters—this was one of the main selling points for end-to-end SLU models in the &lt;a href=&quot;https://research.fb.com/wp-content/uploads/2018/02/towards-end-to-end-spoken-language-understanding.pdf&quot;&gt;original paper&lt;/a&gt;.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;em&gt;As for “well”:&lt;/em&gt; Does 75% accuracy count as good? Probably not, unless you’re OK with your cooking timer being set for 10 hours instead of 10 minutes now and then. For starters, we saw that the training data—LibriSpeech and the synthetic Timers and Such speech—is almost entirely American English, so we would need to collect some more accents for the training data. But I’m American, so it works well for me! (The linguistic equivalent of “it runs on my machine”.)&lt;/p&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;code-and-data&quot;&gt;Code and data&lt;/h2&gt;
&lt;p&gt;I wrote the code for all these experiments as a set of recipes to be included in the SpeechBrain toolkit. It’s not available to the public yet, but it will be soon. In the meantime, you can train a model on &lt;a href=&quot;https://zenodo.org/record/4110812&quot;&gt;Timers and Such v0.1&lt;/a&gt; using my older end-to-end SLU code &lt;a href=&quot;https://github.com/lorenlugosch/end-to-end-SLU&quot;&gt;here&lt;/a&gt;—though I would recommend waiting until SpeechBrain comes out, since my SpeechBrain SLU recipes are a lot cleaner and easier to use.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
  &lt;ol&gt;
    &lt;li id=&quot;fn:RL&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;The disadvantage of formulating the problem this way is that we need someone to design this output format and write a program to map the intent to a sequence of actions. A better way to formulate the problem might be to use reinforcement learning: let the agent act in response to requests and learn to act to maximize some reward signal (“did the agent do what I want?”), without the need for any hard-coded semantics, as was originally suggested in &lt;a href=&quot;https://pdfs.semanticscholar.org/9f62/db97e65e042657d43b5739e9bbdba14ed159.pdf&quot;&gt;this paper&lt;/a&gt;. The question then becomes: what should our action space look like? High-level actions, like pushing buttons in an app (inflexible, but easier to learn)? Or low-level actions, like reading and writing to locations in memory (flexible, but more difficult to learn)? And can we train the model with some sort of imitation learning or simulation, so that we don’t have to wait forever for it to learn from human feedback? Interesting and challenging questions that I won’t linger on here, but which I’d like to think more about in the future. &lt;a href=&quot;#fnref:RL&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:TTS&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;I wrote a &lt;a href=&quot;https://ieeexplore.ieee.org/abstract/document/9053063&quot;&gt;whole paper&lt;/a&gt; about this idea! &lt;a href=&quot;#fnref:TTS&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:FST&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;We could probably do better using an &lt;a href=&quot;https://arxiv.org/abs/2010.01003&quot;&gt;FST&lt;/a&gt;-based speech recognizer, which would allow us to perfectly constrain the model to only outputting sentences that fit a certain grammar (that is, the “G” part of “&lt;a href=&quot;https://cs.nyu.edu/~mohri/pub/csl01.pdf&quot;&gt;HCLG&lt;/a&gt;”). &lt;a href=&quot;#fnref:FST&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:backprop&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Though not impossible, using tricks like Gumbel-Softmax and the straight-through estimator. &lt;a href=&quot;#fnref:backprop&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:non-transcript&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;A fourth (more subtle) disadvantage of ASR-based SLU models is that the speech signal may contain information that is not present in the transcript. For example, sarcasm is not always apparent from just looking at a transcript. This is not really relevant for the simple numeric commands we’re dealing with here, but for more general-purpose robust language understanding in robots of the future, non-transcript information might be crucial, in addition to other multimodal information like visual cues. &lt;a href=&quot;#fnref:non-transcript&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:direct&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;This was proposed by my friend Dima Serdyuk in his &lt;a href=&quot;https://research.fb.com/wp-content/uploads/2018/02/towards-end-to-end-spoken-language-understanding.pdf&quot;&gt;2018 ICASSP paper&lt;/a&gt;. A few other groups—including &lt;a href=&quot;fluent.ai&quot;&gt;Fluent.ai&lt;/a&gt;, where I worked before my PhD—had had similar ideas earlier, but as far as I’m aware, Dima was the first to get a truly end-to-end SLU model to work, without any sort of ASR-based inductive bias or transfer learning. &lt;a href=&quot;#fnref:direct&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:vapnik&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Actually, there are tons of examples of solving a more general problem yielding better results for the problem of interest, like language model pre-training for text classification. But as the amount of data we have for the problem of interest goes to infinity, Vapnik is right. &lt;a href=&quot;#fnref:vapnik&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:pretrain&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;In &lt;a href=&quot;https://arxiv.org/abs/1904.03670&quot;&gt;this paper&lt;/a&gt;, I proposed a somewhat more complicated way of pre-training the encoder using phoneme and word targets from a forced aligner. The idea was that using word targets would be ideal (and more amenable to an idea I had for using pre-trained word embeddings to help the model understand the meaning of synonyms not present in the SLU training set), but using too many word targets would be expensive, which is why we used phoneme targets as well. Using the pre-trained ASR model’s encoder was a lot simpler to implement, though I haven’t done a fair comparison with the forced alignment approach yet. &lt;a href=&quot;#fnref:pretrain&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
&lt;/div&gt;</content><author><name>Loren Lugosch</name></author><summary type="html">I make fairly heavy use of the voice assistant on my phone for things like setting timers while cooking. As a result, when I spent some time this summer at my in-laws’ place—where there was no cell signal and not-very-good Wi-Fi—I often tried using Siri only to get a sad little “sorry, no Internet :(“ response. (#FirstWorldProblems.)</summary></entry><entry><title type="html">Sequence-to-sequence learning with Transducers</title><link href="https://lorenlugosch.github.io/posts/2020/11/transducer/" rel="alternate" type="text/html" title="Sequence-to-sequence learning with Transducers" /><published>2020-11-16T00:00:00-08:00</published><updated>2020-11-16T00:00:00-08:00</updated><id>https://lorenlugosch.github.io/posts/2020/11/transducer</id><content type="html" xml:base="https://lorenlugosch.github.io/posts/2020/11/transducer/">&lt;p&gt;The &lt;strong&gt;Transducer&lt;/strong&gt; (sometimes called the “RNN Transducer” or “RNN-T”, though it need not use RNNs) is a sequence-to-sequence model proposed by Alex Graves in “&lt;a href=&quot;https://arxiv.org/abs/1211.3711&quot;&gt;Sequence Transduction with Recurrent Neural Networks&lt;/a&gt;”. The paper was published at the &lt;a href=&quot;https://sites.google.com/site/representationworkshopicml2012/&quot;&gt;ICML 2012 Workshop on Representation Learning&lt;/a&gt;. Graves showed that the Transducer was a sensible model to use for speech recognition, achieving good results on a small dataset (TIMIT).&lt;/p&gt;

&lt;p&gt;Since then, the Transducer hasn’t been used as much compared to &lt;a href=&quot;https://www.cs.toronto.edu/~graves/icml_2006.pdf&quot;&gt;CTC&lt;/a&gt; models (like &lt;a href=&quot;https://arxiv.org/abs/1512.02595&quot;&gt;Deep Speech 2&lt;/a&gt;) or &lt;a href=&quot;https://arxiv.org/abs/1409.0473&quot;&gt;attention&lt;/a&gt; models (like &lt;a href=&quot;https://arxiv.org/abs/1508.01211&quot;&gt;Listen, Attend, and Spell&lt;/a&gt;). Last year, however, the Transducer got some serious attention when Google researchers showed that it could enable &lt;a href=&quot;https://ai.googleblog.com/2019/03/an-all-neural-on-device-speech.html&quot;&gt;entirely on-device low-latency speech recognition&lt;/a&gt; for Pixel phones. And more recently, the Transducer was used to achieve a &lt;a href=&quot;https://arxiv.org/pdf/2010.10504.pdf&quot;&gt;new state-of-the-art&lt;/a&gt; word error rate for the LibriSpeech benchmark.&lt;sup id=&quot;fnref:foot&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:foot&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;So what is the Transducer, and when might you want to use it? In this post, we will see where Transducer models fit in with other sequence-to-sequence models and a detailed explanation of how they work.&lt;/p&gt;

&lt;p&gt;This post also includes a Colab notebook with a PyTorch implementation of the Transducer for a toy problem—which you can skip straight to &lt;a href=&quot;https://github.com/lorenlugosch/transducer-tutorial/blob/main/transducer_tutorial_example.ipynb&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;

&lt;h2 id=&quot;attention-models&quot;&gt;Attention models&lt;/h2&gt;
&lt;p&gt;The problems we’re interested in here are &lt;strong&gt;sequence transduction&lt;/strong&gt; problems, where the goal is to map an input sequence $\mathbf{x} = \{x_1, x_2, \dots x_T\}$ to an output sequence $\mathbf{y} = \{y_1, y_2, \dots, y_U\}$.&lt;/p&gt;

&lt;p&gt;The go-to models for sequence transduction problems are attention-based sequence-to-sequence models, like &lt;a href=&quot;https://arxiv.org/abs/1409.0473&quot;&gt;RNN encoder-decoder models&lt;/a&gt; or &lt;a href=&quot;https://arxiv.org/abs/1706.03762&quot;&gt;Transformers&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Here’s a diagram of an attention model. (In the diagrams below, I’ll use &lt;span style=&quot;color:red&quot;&gt;&lt;strong&gt;red&lt;/strong&gt;&lt;/span&gt; to indicate that a module has access to $\mathbf{x}$, &lt;span style=&quot;color:blue&quot;&gt;&lt;strong&gt;blue&lt;/strong&gt;&lt;/span&gt; to indicate access to $\mathbf{y}$, and &lt;span style=&quot;color:purple&quot;&gt;&lt;strong&gt;purple&lt;/strong&gt;&lt;/span&gt; to indicate access to both $\mathbf{x}$ and $\mathbf{y}$.)&lt;/p&gt;

&lt;!-- &lt;describe attention model&gt; --&gt;
&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/transducer/attention-model.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;
The model encodes the input $\mathbf{x}$ into a sequence of feature vectors, then computes the probability of the next output $y_u$ as a function of the encoded input and previous outputs. The attention mechanism allows the decoder to look at different parts of the input sequence when predicting each output. Here, for example, is a heatmap of where the decoder is looking during a translation task (from &lt;a href=&quot;https://arxiv.org/abs/1409.0473&quot;&gt;Bahdanau et al.&lt;/a&gt;):&lt;/p&gt;

&lt;!-- &lt;show attention heatmap for speech recognition&gt;  --&gt;
&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/transducer/attention-heatmap.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;
Attention models can be applied to any problem, but they are not always the best choice for certain problems, like speech recognition, for a few reasons&lt;sup id=&quot;fnref:mocha&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:mocha&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;The attention operation is expensive for long input sequences. The complexity of attending to the entire input for every output is $O(TU)$—and for audio, $T$ and $U$ are big.&lt;/li&gt;
  &lt;li&gt;Attention models cannot be run &lt;em&gt;online&lt;/em&gt; (in real time), since the entire input sequence needs to be available before the decoder can attend to it.&lt;/li&gt;
  &lt;li&gt;Attention models also don’t take advantage of the fact that, for speech recognition, the alignment between inputs and outputs is &lt;strong&gt;monotonic&lt;/strong&gt;: that is, if word A comes after word B in the transcript, word A must come after word B in the audio signal (see image below, from &lt;a href=&quot;https://arxiv.org/pdf/1508.01211.pdf&quot;&gt;Chan et al.&lt;/a&gt;, for an example of a monotonic alignment). The fact that attention models lack this inductive bias seems to make them &lt;a href=&quot;https://awni.github.io/train-sequence-models/&quot;&gt;harder to train&lt;/a&gt; for speech recognition; it’s common to add &lt;a href=&quot;https://arxiv.org/abs/1609.06773&quot;&gt;auxiliary loss terms&lt;/a&gt; to stabilize training.&lt;/li&gt;
&lt;/ul&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/transducer/attention-ASR.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;This leads us to Connectionist Temporal Classification (CTC) models, which are more suitable for some problems than attention models.&lt;/p&gt;

&lt;h2 id=&quot;ctc-models&quot;&gt;CTC models&lt;/h2&gt;

&lt;p&gt;CTC models assume that there is a monotonic input-output alignment&lt;sup id=&quot;fnref:CTC&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:CTC&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;. This ends up making the model a lot simpler.&lt;/p&gt;

&lt;!-- &lt;show lattice&gt; --&gt;
&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/transducer/ctc-model.png&quot; style=&quot;max-width:30%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;
So simple! We only need a single neural net to implement a CTC model, and no expensive global attention mechanism.&lt;/p&gt;

&lt;p&gt;But CTC models have a couple problems of their own:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Problem 1:&lt;/strong&gt; &lt;em&gt;The output sequence length $U$ has to be smaller than the input sequence length $T$.&lt;/em&gt; This might not seem like a problem for speech recognition, where $T$ is much larger than $U$—but it prevents us from using a model architecture that does a lot of pooling, which can make the model a lot faster.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Problem 2:&lt;/strong&gt; &lt;em&gt;The outputs are assumed to be independent of each other.&lt;/em&gt; The result is that CTC models often produce outputs that are obviously wrong, like “I eight food” instead of “I ate food”. Getting good results with CTC usually requires a search algorithm that incorporates a secondary language model.&lt;sup id=&quot;fnref:Jasper&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:Jasper&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Can we do better than CTC? Yes: using Transducer models.&lt;/p&gt;

&lt;h2 id=&quot;transducer-models&quot;&gt;Transducer models&lt;/h2&gt;

&lt;p&gt;The Transducer elegantly solves both problems associated with CTC, while retaining some of its advantages over attention models.&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;It solves &lt;strong&gt;Problem 1&lt;/strong&gt; by allowing multiple outputs for each input.&lt;/li&gt;
  &lt;li&gt;It solves &lt;strong&gt;Problem 2&lt;/strong&gt; by adding a predictor network and joiner&lt;sup id=&quot;fnref:joiner&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:joiner&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;5&lt;/a&gt;&lt;/sup&gt; network.&lt;/li&gt;
&lt;/ul&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/transducer/transducer-model.png&quot; style=&quot;max-width:60%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;
The predictor is autoregressive: it takes as input the previous outputs and produces features that can be used for predicting the next output, like a standard language model.&lt;/p&gt;

&lt;p&gt;The joiner is a simple feedforward network that combines the encoder vector $f_t$ and predictor vector $g_u$ and outputs a softmax $h_{t,u}$ over all the labels, as well as a “null” output $\varnothing$.&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/transducer/joiner-output.png&quot; style=&quot;max-width:33%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;
Given an input sequence $\mathbf{x}$, generating an output sequence $\mathbf{y}$ can be done using a simple greedy search algorithm:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;
    &lt;p&gt;Start by setting $t := 1$, $u := 0$, and $\mathbf{y} :=$ an empty list.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;Compute $f_t$ using $\mathbf{x}$ and $g_u$ using $\mathbf{y}$.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;Compute $h_{t,u}$ using $f_t$ and $g_u$.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;If the argmax of $h_{t,u}$ is a &lt;em&gt;label&lt;/em&gt;, set $u := u + 1$, and output the label (append it to $\mathbf{y}$ and feed it back into the predictor).&lt;br /&gt;&lt;br /&gt;If the argmax of $h_{t,u}$ is $\varnothing$, set $t := t + 1$ (in other words, just move to the next input timestep and output nothing).&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;If $t=T+1$, we’re done. Else, go back to step 2.&lt;/p&gt;
  &lt;/li&gt;
&lt;/ol&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/transducer/greedy-search.png&quot; style=&quot;max-width:90%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;
A couple cool things about Transducers to note here:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;
    &lt;p&gt;If the encoder is causal (i.e., we’re not using something like a bidirectional RNN), then the search can run in an online/&lt;a href=&quot;https://twitter.com/lorenlugosch/status/1327330577104695297?s=20&quot;&gt;streaming&lt;/a&gt; fashion, where we process each $x_t$ as soon as it arrives.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;The predictor only has access to $\mathbf{y}$, and not $\mathbf{x}$—unlike the decoder in an attention model, which sees both $\mathbf{x}$ and $\mathbf{y}$. That means we can easily pre-train the predictor on text-only data, which there’s a lot more of than paired (speech, text) data.&lt;/p&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;alignment&quot;&gt;Alignment&lt;/h2&gt;

&lt;p&gt;Given an $(\mathbf{x}, \mathbf{y})$ pair, the Transducer defines a set of possible monotonic alignments between $\mathbf{x}$ and $\mathbf{y}$. For example, consider an input sequence of length $T = 4$ and an output sequence (“CAT”) of length $U = 3$. We can illustrate the set of alignments using a graph&lt;sup id=&quot;fnref:FST&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:FST&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;6&lt;/a&gt;&lt;/sup&gt; like this:&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/transducer/transducer-graph.png&quot; style=&quot;max-width:60%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;
Here’s one alignment: $\mathbf{z} = \varnothing, C, A, \varnothing, T, \varnothing, \varnothing$&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/transducer/cat-align-1.png&quot; style=&quot;max-width:60%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;
Here’s another alignment: $\mathbf{z} = C, \varnothing, A, \varnothing, T, \varnothing, \varnothing$&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/transducer/cat-align-2.png&quot; style=&quot;max-width:60%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;
&lt;!-- *(Why the dangling $\varnothing$ at the end, you ask? I'm actually not sure!*  --&gt;
&lt;!-- Training a model without it seems to work just fine. But that's how the original model is specified, so I've included it here.)* --&gt;&lt;/p&gt;

&lt;p&gt;We can calculate the probability of one of these alignments by multiplying together the values of each edge along the path:&lt;/p&gt;

&lt;p&gt;$\mathbf{z} = \varnothing, C, A, \varnothing, T, \varnothing, \varnothing$&lt;br /&gt;
↓
$p(\mathbf{z} | \mathbf{x}) = h_{1,0}[\varnothing] \cdot h_{2,0}[C] \cdot h_{2,1}[A] \cdot h_{2,2}[\varnothing] \cdot h_{3,2}[T] \cdot h_{3,3}[\varnothing] \cdot h_{4,3}[\varnothing],$&lt;/p&gt;

&lt;p&gt;where the value of an edge is the corresponding entry of $h_{t,u}$.&lt;/p&gt;
&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/transducer/edge-weights.png&quot; style=&quot;max-width:50%&quot; /&gt;&lt;/center&gt;

&lt;!-- If you're wondering how I got the $(t,u)$ indices, we start from $(t=1, u=0)$, increment $t$ when the edge is $\varnothing$, and increment $u$ when the edge is a label. --&gt;

&lt;h2 id=&quot;training&quot;&gt;Training&lt;/h2&gt;

&lt;p&gt;How do we train the model? If we knew the true alignment&lt;sup id=&quot;fnref:alignment&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:alignment&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;7&lt;/a&gt;&lt;/sup&gt; $\mathbf{z}$, we could minimize the cross-entropy between $\mathbf{h}$ and $\mathbf{z}$, like a normal classifier. However, we usually don’t know the true alignment (and for some tasks, a “true” alignment might not even exist).&lt;/p&gt;

&lt;p&gt;Instead, the Transducer defines $p(\mathbf{y}|\mathbf{x})$ as the sum of the probabilities of &lt;em&gt;all&lt;/em&gt; possible alignments between $\mathbf{x}$ and $\mathbf{y}$. We train the model by minimizing the loss function $-\log p(\mathbf{y}|\mathbf{x})$.&lt;/p&gt;

&lt;p&gt;There are usually too many possible alignments to compute the loss function by just adding them all up directly. To compute the sum efficiently, we compute the “forward variable” $\alpha_{t,u}$, for $1 \leq t \leq T$ and $0 \leq u \leq U$:&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*} \alpha_{t,u} = \alpha_{t-1,u} \cdot h_{t-1,u}[\varnothing] \\+ \alpha_{t,u-1} \cdot h_{t,u-1}[y_{u-1}] \end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;We can visualize this computation as passing values along the edges of the alignment graph:&lt;/p&gt;
&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/transducer/forward-messages.png&quot; style=&quot;max-width:50%&quot; /&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;
After we’ve computed $\alpha_{t,u}$ for every node in the alignment graph, we get $p(\mathbf{y}|\mathbf{x})$ using the forward variable at the last node of the graph:&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*} p(\mathbf{y}|\mathbf{x}) = \alpha_{T,U} \cdot h_{T,U}[\varnothing]\end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;We need to do everything in the log domain, for the &lt;a href=&quot;https://lorenlugosch.github.io/posts/2020/06/logsumexp/&quot;&gt;usual reasons&lt;/a&gt;. In the log domain, the computation becomes:&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*} \log \alpha_{t,u} = \text{logsumexp}([\log \alpha_{t-1,u} + \log h_{t-1,u}[\varnothing], \\ \log \alpha_{t,u-1} + \log h_{t,u-1}[y_{u-1}] ]) \end{eqnarray*}$$&lt;/center&gt;

&lt;center&gt;$$\begin{eqnarray*} \log p(\mathbf{y}|\mathbf{x}) = \log \alpha_{T,U} + \log h_{T,U}[\varnothing]\end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;Finally, to compute the gradient of the loss $-\log p(\mathbf{y}|\mathbf{x})$, there is a second algorithm that computes a backward variable $\beta_{t,u}$, using the same computation as $\alpha_{t,u}$, but in reverse, starting from the last node.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;In the notebook, I provide a simple PyTorch implementation of the loss function that only writes out the forward computation and uses automatic differentiation to compute the gradient. This is a lot slower than a lower-level implementation, but easier to program and to read.&lt;/em&gt;&lt;/p&gt;

&lt;h2 id=&quot;memory-usage&quot;&gt;Memory usage&lt;/h2&gt;

&lt;p&gt;In general, Transducer models seem like a good idea. But here’s the catch (and possibly the unspoken reason that the Transducer never caught on until recently):&lt;/p&gt;

&lt;p&gt;Suppose we have $T=1000$, $U=100$, $L=1000$ labels, and batch size $B=32$. Then to store $h_{t,u}$ for all $(t,u)$ to run the forward-backward algorithm, we need a tensor of size $B \times T \times U \times L = $ 3,200,000,000, or 12.8 GB if we’re using single-precision floats. And that’s just the output tensor: there’s also the hidden unit activations of the joiner network, which are of size $B \times T \times U \times d_{\text{joiner}}$.&lt;/p&gt;

&lt;p&gt;So unless you are, &lt;em&gt;ahem&lt;/em&gt;, a certain tech company in possession of TPUs with plentiful RAM (guess who’s been publishing the most Transducer papers!), you may need to find some way to reduce memory consumption during training—e.g., by pooling in the encoder to reduce $T$, or by using a small batch size $B$.&lt;/p&gt;

&lt;p&gt;Ironically, this is only a problem during training; during inference, we only need a small amount of memory to store the current activations and hypotheses for $\mathbf{y}$.&lt;/p&gt;

&lt;h2 id=&quot;search&quot;&gt;Search&lt;/h2&gt;

&lt;p&gt;We saw earlier that you can predict $\mathbf{y}$ using a greedy search, always picking the top output of $h_{t,u}$. Better results can be obtained using a beam search instead, maintaining a list of multiple hypotheses for $\mathbf{y}$ and updating them at each input timestep.&lt;/p&gt;

&lt;p&gt;The Transducer beam search algorithm can be found in the original paper—though it is somewhat gnarlier than the simple attention model beam search, and I confess I haven’t implemented it myself yet. (Check out the soon-to-be-released &lt;a href=&quot;https://speechbrain.github.io/&quot;&gt;SpeechBrain&lt;/a&gt; toolkit for my colleagues’ implementation.)&lt;/p&gt;

&lt;h2 id=&quot;code&quot;&gt;Code&lt;/h2&gt;

&lt;p&gt;Finally, the Colab notebook for the Transducer can be found &lt;a href=&quot;https://github.com/lorenlugosch/transducer-tutorial/blob/main/transducer_tutorial_example.ipynb&quot;&gt;here&lt;/a&gt;. The notebook implements a Transducer model in PyTorch for a toy sequence transduction problem (filling in missing vowels in a sentence: “hll wrld” –&amp;gt; “hello world”), including the loss function, the greedy search, and a function for computing the probability of a single alignment. Enjoy!&lt;/p&gt;

&lt;h2 id=&quot;citation&quot;&gt;Citation&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;If you found this tutorial helpful and would like to cite it, you can use the following BibTeX entry:&lt;/em&gt;&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;@misc{
	lugosch_2020, 
	title={Sequence-to-sequence learning with Transducers}, 
	url={https://lorenlugosch.github.io/posts/2020/11/transducer/}, 
	author={Lugosch, Loren}, 
	year={2020}, 
	month={Nov}
}
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;
&lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
  &lt;ol&gt;
    &lt;li id=&quot;fn:foot&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;It always seems to take a few years between Alex Graves publishing a good idea and the research community fully recognizing it. There was an 8 year gap between CTC (2006) and Baidu’s Deep Speech (2014), and an 8 year gap between the Transducer (2012) and Google’s latest result (2020). This suggests a simple algorithm for achieving state-of-the-art results: select a paper written by Alex Graves from 8 years ago, and reimplement it using whatever advances in deep learning have been made since then. Maybe 2022 will be the year &lt;a href=&quot;https://arxiv.org/abs/1410.5401&quot;&gt;Neural Turing Machines&lt;/a&gt; really shine! &lt;a href=&quot;#fnref:foot&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:mocha&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;There’s been some interesting work developing attention models that do not have these three issues, like &lt;a href=&quot;https://arxiv.org/abs/1712.05382&quot;&gt;monotonic chunkwise attention (MoChA)&lt;/a&gt;. &lt;a href=&quot;#fnref:mocha&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:CTC&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;As do their older cousins, &lt;a href=&quot;https://lorenlugosch.github.io/posts/2020/01/hmm/&quot;&gt;Hidden Markov Models&lt;/a&gt;, &lt;a href=&quot;https://pdfs.semanticscholar.org/62d7/9ced441a6c78dfd161fb472c5769791192f6.pdf&quot;&gt;Graph Transformer Networks&lt;/a&gt;, and the more recent &lt;a href=&quot;https://arxiv.org/pdf/1609.03193.pdf&quot;&gt;AutoSegCriterion&lt;/a&gt;. See Awni Hannun’s excellent &lt;a href=&quot;https://distill.pub/2017/ctc/&quot;&gt;introduction&lt;/a&gt; if you want to learn more about CTC. &lt;a href=&quot;#fnref:CTC&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:Jasper&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Alternately, you can use a very big and deep network like &lt;a href=&quot;https://arxiv.org/pdf/1904.03288.pdf&quot;&gt;Jasper&lt;/a&gt;. Jasper was a CTC model proposed by NVIDIA researchers that, astonishingly, achieved nearly state-of-the-art performance using only a greedy search. If the model is big and deep, it can intelligently coordinate its outputs so as to not produce dumb predictions like “I eight food” instead of “I ate food”. Still, it seems to be more parameter-efficient to use a model that explicitly assumes that outputs are not independent, like attention models and Transducer models. &lt;a href=&quot;#fnref:Jasper&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:joiner&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;In the original paper, there was no joiner; the encoder vector and predictor vector were simply added together. Graves and his co-authors added the joiner in a &lt;a href=&quot;https://www.cs.toronto.edu/~fritz/absps/RNN13.pdf&quot;&gt;subsequent paper&lt;/a&gt;, finding that it reduced the number of deletion errors. &lt;a href=&quot;#fnref:joiner&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:FST&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;If you’re familiar with the powerful gadgets known as &lt;a href=&quot;http://www.opengrm.org/twiki/bin/view/GRM/PyniniDocs&quot;&gt;finite state transducers (FSTs)&lt;/a&gt;, you may recognize that the Transducer graph is a weighted FST, where an alignment forms the input labels, $\mathbf{y}$ forms the output labels, and the weight for each edge is dynamically generated by the joiner network. &lt;a href=&quot;#fnref:FST&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:alignment&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;The true alignment of neural networks is known to be Chaotic Good. &lt;a href=&quot;#fnref:alignment&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
&lt;/div&gt;</content><author><name>Loren Lugosch</name></author><category term="sequence modeling" /><summary type="html">The Transducer (sometimes called the “RNN Transducer” or “RNN-T”, though it need not use RNNs) is a sequence-to-sequence model proposed by Alex Graves in “Sequence Transduction with Recurrent Neural Networks”. The paper was published at the ICML 2012 Workshop on Representation Learning. Graves showed that the Transducer was a sensible model to use for speech recognition, achieving good results on a small dataset (TIMIT).</summary></entry><entry><title type="html">My research goals</title><link href="https://lorenlugosch.github.io/posts/2020/08/goals/" rel="alternate" type="text/html" title="My research goals" /><published>2020-08-14T00:00:00-07:00</published><updated>2020-08-14T00:00:00-07:00</updated><id>https://lorenlugosch.github.io/posts/2020/08/goals</id><content type="html" xml:base="https://lorenlugosch.github.io/posts/2020/08/goals/">&lt;p&gt;I wanted to clarify to myself and others what some of my research goals are, and why I’m working on certain problems. The hope is that putting this online for the world to see will help challenge me to keep focused and working towards those goals—sort of like telling your friends that you’re going to quit smoking, or something like that.&lt;/p&gt;

&lt;p&gt;My broad long-term goal is to build reliable, competent domestic robots: in other words, robots that can help you around the house with tidying up, folding laundry, and the kind of things you might now ask Alexa/Siri/Google Home for, like playing music and setting timers.&lt;/p&gt;

&lt;p&gt;In addition to the interesting technical challenge of it, domestic robots are just something I personally would love to have. They could also make life a lot better for elderly people and people who need long-term care—if you can’t pay for a caregiver, and if your loved ones aren’t able to take on the role of caregiver, a robot might be an affordable alternative.&lt;/p&gt;

&lt;p&gt;An important aspect of my goal is: I don’t want these robots to rely on the Internet—I want their AI to live on the robot, offline. There’s a few reasons for this.&lt;/p&gt;

&lt;p&gt;1) &lt;em&gt;Privacy.&lt;/em&gt; Yes, yes, I know there’s a lot of people working on things like computation on encrypted data (“some of my best friends work on data privacy!”), but barring some big breakthroughs, I’d prefer to just cut the Gordian Knot and avoid sending my data to the cloud altogether.&lt;/p&gt;

&lt;p&gt;2) &lt;em&gt;Latency/reliability.&lt;/em&gt; Suppose a robot lives with an elderly person, and the person is about to slip and fall. The robot needs to detect this and quickly move to keep the person from getting hurt. We might use a neural network to map the robot’s camera feed to a sequence of physical actions to take. If we store that neural network in the cloud, and the Internet cuts out, the robot won’t be able to act.&lt;/p&gt;

&lt;p&gt;3) &lt;em&gt;The Internet shouldn’t even be necessary.&lt;/em&gt; If it’s not something that inherently requires the Internet, like checking the weather forecast or buying groceries, I shouldn’t have to use the Internet to do it. My brain takes up the space of about a hard drive and runs on 20 Watts—I don’t have to store it in the cloud and run it on a supercomputer. It’s annoying that I need the Internet just to tell my phone to set a reminder.&lt;/p&gt;

&lt;p&gt;And this last one admittedly isn’t a very good reason, but I want to be honest with myself:&lt;/p&gt;

&lt;p&gt;4) &lt;em&gt;I want to own all my stuff.&lt;/em&gt; I’m annoyed that less and less of my stuff consists of physical things that I definitely own, and instead is more virtual things in the cloud that I maybe kind of own or maybe am just renting. I want to own my data, for instance; I don’t want to have to worry about where it lives and who can see it. (But maybe this is just a bit of nostalgia, or a bit of the tin-foil-hat-wearing libertarian’s instinct to convert his digital money into a pile of gold and hide it under the bed.)&lt;/p&gt;

&lt;blockquote class=&quot;twitter-tweet&quot;&gt;&lt;p lang=&quot;en&quot; dir=&quot;ltr&quot;&gt;industry: would you like to move all your code and data to a rich man&amp;#39;s personal computer?&lt;br /&gt;programmer: what no that sounds like a bad idea&lt;br /&gt;industry: ok would you like to run everything in &amp;quot;the cloud&amp;quot;?&lt;br /&gt;programmer: oh that sounds fluffy and nice yes please&lt;/p&gt;&amp;mdash; Computer Facts (@computerfact) &lt;a href=&quot;https://twitter.com/computerfact/status/1192938091201335296?ref_src=twsrc%5Etfw&quot;&gt;November 8, 2019&lt;/a&gt;&lt;/blockquote&gt;
&lt;script async=&quot;&quot; src=&quot;https://platform.twitter.com/widgets.js&quot; charset=&quot;utf-8&quot;&gt;&lt;/script&gt;

&lt;p&gt;I suspect that, to make really good and reliable domestic robots happen, solving the AI problems—true spoken language understanding, fine-grained motor control, long-term planning—will require &lt;em&gt;extremely large neural networks&lt;/em&gt;. I’m talking hundreds of trillions of parameters. For reference, the current biggest neural net reported in the literature has only &lt;a href=&quot;https://arxiv.org/abs/2006.16668&quot;&gt;600 billion parameters&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Why do I say that? There’s lots of evidence. &lt;a href=&quot;https://www.gwern.net/GPT-3&quot;&gt;GPT-3&lt;/a&gt; has 175 billion parameters—there’s &lt;a href=&quot;https://arxiv.org/abs/2001.08361&quot;&gt;theoretical&lt;/a&gt; and &lt;a href=&quot;https://www.youtube.com/watch?v=9P_VAMyb-7k&quot;&gt;empirical&lt;/a&gt; reasons to believe that big models really are necessary—and that’s &lt;em&gt;just&lt;/em&gt; for language understanding/generation: our robot will also need motor control, audio processing, vision, and haptics. The way we do AI now, each of these modalities is handled separately; I think ultimately we will need the modalities to be &lt;a href=&quot;https://arxiv.org/abs/1706.05137&quot;&gt;handled together&lt;/a&gt;, so that your robot doesn’t think &lt;a href=&quot;https://openai.com/blog/better-language-models/&quot;&gt;a unicorn might have 4 horns&lt;/a&gt;—and so our models will need even more capacity, for understanding how the different modalities fit together. And if you like taking inspiration from human brains: ours have 100 trillion synapses—assuming that 1 synapse = 1 weight, and assuming that Nature is efficient, then we’ll need 100 trillion parameters to make AI that can do the range of things human brains can do.&lt;/p&gt;

&lt;p&gt;Running 100-trillion-parameter networks on a domestic robot is going to be hard. Just running a 600-billion-parameter network on a supercomputer was a significant engineering challenge for &lt;a href=&quot;https://arxiv.org/abs/2006.16668&quot;&gt;Google&lt;/a&gt;. Will Moore’s Law alone get us there? Could we just wait until 2040, when a Raspberry Pi has 1 PetaFLOPS of processing and 100 TB of storage? Or do we need to consider fundamentally different hardware designs?&lt;/p&gt;

&lt;p&gt;No, I’m not talking about matrix multiplication accelerators, which have been optimized to death already. I’m talking about data storage—way less sexy, but ripe for neural-net-specific optimizations. Right now, &lt;a href=&quot;https://twitter.com/hardmaru/status/1289498209320972290?s=20&quot;&gt;a 100 TB drive costs $40,000&lt;/a&gt;. The alternative, connecting up dozens of cheaper, smaller drives, might not be energy-efficient or space-efficient enough for a robot. Could we make denser, cheaper storage by taking advantage of the fact that—unlike the hard drives for storing your OS and bank account information—neural nets can withstand a little noise?&lt;/p&gt;

&lt;p&gt;Whatever the hardware for domestic robot AI looks like, conditional computation—that is, not using all 100 trillion parameters of the neural network every time the clock ticks—is certainly going to be a part of the solution. That’s the focus of my PhD right now. My current hunch is that the &lt;a href=&quot;https://arxiv.org/abs/1701.06538&quot;&gt;hard mixture-of-experts&lt;/a&gt; is a good starting point, but not the final word in conditional computation. Just as inductive biases like translational invariance can make computer vision easier to learn, there are simple inductive biases from computer architecture and psychology which I think can make the hard mixture-of-experts easier to train. The fun part of AI is figuring out how little inductive bias we can get away with—and pilfering that bias from human brains and other sources of inspiration when we do need it.&lt;/p&gt;</content><author><name>Loren Lugosch</name></author><summary type="html">I wanted to clarify to myself and others what some of my research goals are, and why I’m working on certain problems. The hope is that putting this online for the world to see will help challenge me to keep focused and working towards those goals—sort of like telling your friends that you’re going to quit smoking, or something like that.</summary></entry><entry><title type="html">Predictive coding in machines and brains</title><link href="https://lorenlugosch.github.io/posts/2020/07/predictive-coding/" rel="alternate" type="text/html" title="Predictive coding in machines and brains" /><published>2020-07-11T00:00:00-07:00</published><updated>2020-07-11T00:00:00-07:00</updated><id>https://lorenlugosch.github.io/posts/2020/07/predictive-coding</id><content type="html" xml:base="https://lorenlugosch.github.io/posts/2020/07/predictive-coding/">&lt;p&gt;The name “predictive coding” has been applied to a number of engineering techniques and scientific theories. All these techniques and theories involve predicting future observations from past observations, but what exactly is meant by “coding” differs in each case. Here is a quick tour of some flavors of “predictive coding” and how they’re related.&lt;/p&gt;

&lt;h2 id=&quot;what-is-coding&quot;&gt;What is “coding”?&lt;/h2&gt;
&lt;p&gt;In signal processing and related fields, the term “coding” generally means putting a signal into some format where the signal will be easier to handle for some task.&lt;/p&gt;

&lt;p&gt;In any coding scheme, there is an encoder, which puts the input signal into the new format, and a decoder, which puts the encoded signal back into the original format (or as close as possible to the original). The “code” is the space of possible encoded signals. For example, the &lt;a href=&quot;https://en.wikipedia.org/wiki/ASCII&quot;&gt;American Standard Code for Information Interchange (ASCII)&lt;/a&gt; consists of a set of 8-bit representations of commonly used characters. Your computer doesn’t understand characters, but it does understand bits, so we use a bit-based coding scheme to represent characters in a computer.&lt;/p&gt;

&lt;p&gt;Some other common types of coding are:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Source coding&lt;/strong&gt; (also known as “&lt;a href=&quot;http://mattmahoney.net/dc/dce.html&quot;&gt;data compression&lt;/a&gt;”). The encoder compresses the input into a small bitstream; the decoder decompresses that bitstream to get back the input. &lt;em&gt;Examples: arithmetic coding (used in JPEG), Huffman coding (used to &lt;a href=&quot;https://arxiv.org/abs/1510.00149&quot;&gt;compress neural networks&lt;/a&gt;).&lt;/em&gt;&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Channel coding&lt;/strong&gt; (also known as “&lt;a href=&quot;https://lorenlugosch.github.io/Masters_Thesis.pdf&quot;&gt;error correction&lt;/a&gt;”). The encoder adds redundancy to protect a message from noise in the channel through which it will be transmitted; the decoder maps the noisy received signal back to the most likely original message, with the help of the added redundancy. &lt;em&gt;Examples: low-density parity-check codes (used in your cell phone and your hard drive), Reed-Solomon codes (used in CDs, for those old enough to remember this technology).&lt;/em&gt;&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Encryption.&lt;/strong&gt; (Encryption could also be called “receiver coding” because it considers possible &lt;em&gt;receivers&lt;/em&gt;, but nobody uses this terminology.) The encoder encrypts the message into a format such that only a receiver with the appropriate keys can open the message; the decoder decrypts the encrypted message into the readable original. &lt;em&gt;Examples: RSA (used in HTTPS), SHA-1 (once used, but now broken).&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Aside:&lt;/em&gt; The “source coding” and “channel coding” lingo come from Claude Shannon’s wonderful 1948 paper “&lt;a href=&quot;http://people.math.harvard.edu/~ctm/home/text/others/shannon/entropy/entropy.pdf&quot;&gt;A mathematical theory of communication&lt;/a&gt;”.&lt;/p&gt;

&lt;h2 id=&quot;predictive-coding-for-data-compression&quot;&gt;Predictive coding for data compression&lt;/h2&gt;
&lt;p&gt;OK, that’s coding in general. So what’s predictive coding?&lt;/p&gt;

&lt;p&gt;The term “predictive coding” was coined in 1955 &lt;a href=&quot;https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=1055126&amp;amp;casa_token=4M_d797vjM8AAAAA:bSflmHvRRXAGjcDRj2UldgoGkmYQggI1Up7hOo3a1wUTht-92EnD89CJ8JVz-xUqqnBmXVdd&amp;amp;tag=1&quot;&gt;by Peter Elias&lt;/a&gt;. Specifically, what Elias proposed was a method called &lt;em&gt;linear&lt;/em&gt; predictive coding (LPC) for communication systems.&lt;/p&gt;

&lt;p&gt;In LPC, the next sample of a signal is predicted using a linear function of the previous $n$ samples. Then the error between the predicted sample and the actual sample is transmitted, along with the coefficients of the linear predictor. Predicting a sample from nearby previous samples works because in signals like speech, nearby samples are strongly correlated with each other.&lt;/p&gt;

&lt;p&gt;The idea behind transmitting the &lt;em&gt;error&lt;/em&gt; in LPC is that if we have a good predictor, the error will be small; thus it will require less bandwidth to transmit than the original signal. (So here “coding” specifically refers to “source coding”, or compression.)&lt;/p&gt;

&lt;p&gt;LPC has had a long history of successes for audio compression. &lt;a href=&quot;https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=6366797&amp;amp;casa_token=QMjcWbI_Ms4AAAAA:d1m6vIdXHIwFnkOtjljmqxjtR5m_S0iywLUj9rqFKPoNlvotD7y5UF9Jg0GQbwKA4jmaAFyj&amp;amp;tag=1&quot;&gt;Speak ‘n’ Spell used LPC&lt;/a&gt; to store and synthesize speech sounds. The old &lt;a href=&quot;https://en.wikipedia.org/wiki/Game_Boy_Sound_System&quot;&gt;Game Boy soundchip&lt;/a&gt; mostly used simple square wave beeps ‘n’ boops to make music, but sometimes used a special case of LPC known as &lt;em&gt;DPCM&lt;/em&gt; to store certain sounds, like Pikachu’s voice in &lt;em&gt;Pokémon Yellow&lt;/em&gt;. (See &lt;a href=&quot;https://www.youtube.com/watch?v=q_3d1x2VPxk&quot;&gt;this video&lt;/a&gt; for a great overview of this and other old-school soundchips. Audio compression was crucial back when game cartridges had limited data storage.) &lt;a href=&quot;https://en.wikipedia.org/wiki/SILK&quot;&gt;The speech codec used in Skype&lt;/a&gt; uses LPC, combined with a bunch of other gadgets.&lt;/p&gt;

&lt;p&gt;Predictive coding is also used in &lt;em&gt;video&lt;/em&gt; compression, under the name “&lt;a href=&quot;https://en.wikipedia.org/wiki/Motion_compensation&quot;&gt;motion compensation&lt;/a&gt;”. Like adjacent audio samples, adjacent video frames are strongly correlated, and so can be predicted from each other. If you haven’t already, it’s good to take a moment to &lt;a href=&quot;https://sidbala.com/h-264-is-magic/&quot;&gt;have your mind blown by H.264 video compression&lt;/a&gt;—the digital equivalent of shrinking a 3000-pound car to 0.4 pounds.&lt;/p&gt;

&lt;p&gt;And of course, &lt;em&gt;linear&lt;/em&gt; models are not the only way to do predictive coding; nonlinear models like neural networks can be used as well. &lt;a href=&quot;https://arxiv.org/pdf/1910.06464.pdf&quot;&gt;A speech codec using WaveNet&lt;/a&gt; has been reported to get lower bitrate for the same quality as some traditional speech codecs.&lt;/p&gt;

&lt;h2 id=&quot;predictive-coding-for-representation-learning&quot;&gt;Predictive coding for representation learning&lt;/h2&gt;

&lt;p&gt;Linear predictive coding and friends can be thought of as special cases of something called an &lt;em&gt;autoregressive model&lt;/em&gt;. An autoregressive model is a model that cleverly splits a complicated probability distribution over sequences into a number of chunks that are easier to handle.&lt;/p&gt;

&lt;p&gt;Let $\mathbf{x} = \{x_1, x_2, \dots\}$ denote our sequence of interest. In an autoregressive model, the joint distribution $p(\mathbf{x})$ is defined as&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*} p(\mathbf{x}) = \prod_t p(x_t|x_{t-1}, x_{t-2}, \dots) \end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;The $p(x_t|x_{t-1}, x_{t-2}, \dots)$ term can be implemented in different ways. In something simple like an $n$-gram model for text, $p(x_t|x_{t-1}, x_{t-2}, \dots)$ is just a lookup table containing the probability of the next letter being $x_t$, given that the previous letters were $x_{t-1}, x_{t-2}, \dots$. Another way to implement $p(x_t|x_{t-1}, x_{t-2}, \dots)$ is to feed $x_{t-1}, x_{t-2}, \dots$ into a neural network, which outputs a feature vector $h$, and then predict $x_t$ using a linear model on top of $h$, like a softmax classifier (for discrete $x_t$) or a linear regression model (for real-valued $x_t$).&lt;/p&gt;

&lt;p&gt;Such a model is called “autoregressive” because if you want to &lt;em&gt;sample&lt;/em&gt; from the distribution it defines, you must feed back the model’s own outputs (= “auto”) to predict the next output (= “regressive”).&lt;/p&gt;

&lt;p&gt;It turns out that if you train a neural network as an autoregressive model, the internal representations learned by the network will &lt;a href=&quot;https://arxiv.org/abs/1511.01432&quot;&gt;work really well for supervised downstream tasks&lt;/a&gt;. Work like &lt;a href=&quot;https://arxiv.org/abs/1801.06146&quot;&gt;ULMFiT&lt;/a&gt; and &lt;a href=&quot;https://worldmodels.github.io/&quot;&gt;World Models&lt;/a&gt; showed that this trick can make neural nets a lot more data-efficient.&lt;/p&gt;

&lt;p&gt;Notice that using the networks’ internal representations is somewhat different from what we did in LPC: whereas in LPC the outputs are the &lt;em&gt;errors&lt;/em&gt; (because the purpose is compression), in these autoregressive feature extractors the outputs are the &lt;em&gt;features&lt;/em&gt; (because the purpose is extracting discriminative features). So in either case, we are predicting future observations, but the “coding scheme” and purpose of predicting the future is different.&lt;/p&gt;

&lt;hr /&gt;
&lt;p&gt;&lt;em&gt;A note about audio.&lt;/em&gt; Autoregressive modeling with big neural nets operating on the raw audio signal works—&lt;a href=&quot;https://arxiv.org/abs/1609.03499&quot;&gt;WaveNet&lt;/a&gt; did this for generative modeling, and the representations could be used on a downstream task (speech recognition)—but it is expensive to do because audio signals are high-dimensional.&lt;/p&gt;

&lt;p&gt;Two methods have been developed to make it easier to do autoregressive modeling for audio: contrastive predictive coding and autoregressive predictive coding.&lt;/p&gt;

&lt;p&gt;The first method, &lt;a href=&quot;https://arxiv.org/abs/1807.03748&quot;&gt;contrastive predictive coding (CPC)&lt;/a&gt;, works by first encoding the input signal into a much lower-dimensional sequence of feature vectors using a convolutional neural network, and then training an autoregressive model on top of this sequence. Since the encoder could simply learn to output all 0s to make the objective function as easy as possible to optimize, the autoregressive component is instead trained using a &lt;em&gt;contrastive&lt;/em&gt; loss: it must guess whether a sample is actually the next sample in the sequence, or a fake, thus preventing the encoder from collapsing to a trivial representation. The technique also works well for other high-dimensional signals, like images (represented as a sequence of pixels). Note that strictly speaking CPC is not a &lt;em&gt;true&lt;/em&gt; autoregressive model because we can’t draw samples from it.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/1904.03240&quot;&gt;Autoregressive predictive coding (APC)&lt;/a&gt; takes a slightly different approach. Instead of modeling the raw audio, it extracts low-dimensional frequency domain features, and then does plain old autoregressive modeling on top of those features. The disadvantage is that 1) maybe some low-level information in the original signal gets thrown out, and 2) now you need to hand-craft some feature extraction, since what works for audio will not necessarily work for other modalities. (Incidentally, I think “autoregressive predictive coding” is not a very good name, because WaveNet is already an “autoregressive” “predictive coding” model. Engineers are not that great at naming things. Oh well.)&lt;/p&gt;

&lt;p&gt;A nice paper comparing CPC and APC was recently published at the ICML 2020 Workshop on Self-supervision in Audio and Speech—check it out &lt;a href=&quot;https://openreview.net/forum?id=cnLz5ckGs1y&quot;&gt;here&lt;/a&gt;, if you’re interested in speech models.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;The drawback of using these predictive coding models for representation learning is that the representations we get are “unidirectional”: that is, they extract information only from the past, and not the future, to represent the current input. That’s a problem because the future is often very informative for interpreting the present. If you think of a phrase like “milk the cow”, we know that “milk” is a verb, and not a noun, from the words that follow it.&lt;/p&gt;

&lt;p&gt;One way to overcome this problem is to do autoregression in both directions, and concatenate the representations from the forward and backward models, as is done in &lt;a href=&quot;https://arxiv.org/abs/1802.05365&quot;&gt;ELMo&lt;/a&gt;. Alternately, models like &lt;a href=&quot;https://arxiv.org/abs/1810.04805&quot;&gt;BERT&lt;/a&gt; use a bidirectional context and minimize a contrastive or denoising loss instead. But for tasks in which observations need to be processed in “real-time”, as is usually the case in control problems, a unidirectional context makes more sense.&lt;/p&gt;

&lt;h2 id=&quot;predictive-coding-for-computational-efficiency&quot;&gt;Predictive coding for computational efficiency&lt;/h2&gt;

&lt;p&gt;Another really neat thing forward predictive models can do is tell you roughly whether an observation is “difficult” or not. The idea is this: if an input is &lt;em&gt;surprising&lt;/em&gt;—if, for example, your autoregressive model assigns low probability to it—it contains more &lt;em&gt;information&lt;/em&gt;, and it is therefore probably worth more attention.&lt;/p&gt;

&lt;p&gt;This observation suggests yet another use for predictive coding: we can allocate less computation to more predictable inputs, a type of “&lt;a href=&quot;https://nervanasystems.github.io/distiller/conditional_computation.html&quot;&gt;conditional computation&lt;/a&gt;”. Jürgen Schmidhuber’s “&lt;a href=&quot;https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.34.1205&amp;amp;rep=rep1&amp;amp;type=pdf&quot;&gt;Neural Sequence Chunker&lt;/a&gt;” is an early instance of this idea, in which predictable inputs are ignored and not sent to a subsequent neural network, and (shameless plug!) recently I wrote &lt;a href=&quot;https://arxiv.org/abs/2006.01659&quot;&gt;a paper&lt;/a&gt; describing a slightly more general version of the idea.&lt;/p&gt;

&lt;p&gt;While unsurprising inputs merit less computation, the inverse is not necessarily true: surprising inputs do not always merit &lt;em&gt;more&lt;/em&gt; computation. Alex Graves in his &lt;a href=&quot;https://arxiv.org/pdf/1603.08983v4.pdf&quot;&gt;Adaptive Computation Time paper&lt;/a&gt; gives an excellent example of unpredictable inputs—specifically, random ID numbers in Wikipedia metadata—which a neural network model does not bother to allocate more computation to, because the model realizes that throwing more processing time at the problem just won’t help. This implies that we can do better than just using predictability by &lt;em&gt;learning&lt;/em&gt; when to use more or less computation. Still, predictability is a good inductive bias for conditional computation in neural nets—and it appears that human brains may use this inductive bias for similar purposes.&lt;/p&gt;

&lt;h2 id=&quot;predictive-coding-in-the-brain&quot;&gt;Predictive coding in the brain&lt;/h2&gt;

&lt;p&gt;Artificial neural networks in machine learning were inspired by scientific theories about the structure of the human brain. For predictive coding, it was the other way around: engineers (starting with Elias) developed predictive coding to solve certain problems in signal processing, and only afterwards did scientists realize that the brain might be doing something like what engineers had developed.&lt;/p&gt;

&lt;p&gt;One such early inkling was described in Jeffrey Elman’s classic “&lt;a href=&quot;https://crl.ucsd.edu/~elman/Papers/fsit.pdf&quot;&gt;Finding Structure in Time&lt;/a&gt;”. Elman trained a recurrent neural network to predict the next letter in a stream of letters. The streams were formed by taking sentences and removing the spaces between words. This is sort of analogous to the scenario of children learning language. Children are not told where the boundaries between words are; presumably they only hear a relatively unbroken stream of phonemes when adults speak.&lt;/p&gt;

&lt;p&gt;What Elman found was that letters with high surprisal corresponded closely to the location of word boundaries. Predictive coding might therefore be one of the ingredients for language learning in humans: children could in theory infer word boundaries by noting which phonemes have high surprisal. (This method isn’t foolproof: my uncle reports that as a kid he thought “tractorworking” was a word because he so often heard “tractor” and “working” together, e.g. “there’s a tractor working in the field”.)&lt;/p&gt;

&lt;p&gt;Another piece of evidence for predictive coding in humans comes from studies of reading time. &lt;a href=&quot;https://www.aclweb.org/anthology/W19-0101.pdf&quot;&gt;van Schijndel and Linzen&lt;/a&gt; put it nicely: “One of the most robust findings in the reading literature is that more predictable words are read faster than less predictable words… Word predictability effects fit into a picture of human cognition in which humans constantly make predictions about upcoming events and test those predictions against their perceptual input.”&lt;/p&gt;

&lt;p&gt;There is an even more general “&lt;a href=&quot;https://royalsocietypublishing.org/doi/pdf/10.1098/rstb.2005.1622&quot;&gt;predictive coding hypothesis&lt;/a&gt;” which claims that the brain does predictive coding at every level: sensory signals are predicted by the neurons that receive them, and the prediction errors become the input to other neurons, which also make predictions and errors, and so on. This hypothesis is somewhat controversial, though, as can be seen from the many responses to Andy Clark’s “&lt;a href=&quot;https://www.fil.ion.ucl.ac.uk/~karl/Whatever%20next.pdf&quot;&gt;Whatever next? Predictive brains, situated agents, and the future of cognitive science&lt;/a&gt;” (the responses go from pg. 24 onward—hear also &lt;a href=&quot;http://unsupervisedthinkingpodcast.blogspot.com/2018/05/episode-33-predictive-coding.html&quot;&gt;Grace Lindsay’s podcast episode on predictive coding&lt;/a&gt;, which discusses this essay).&lt;/p&gt;

&lt;h2 id=&quot;the-future-of-predictive-coding&quot;&gt;The future of predictive coding&lt;/h2&gt;
&lt;p&gt;Regardless of the extent to which it actually happens in human brains, predictive coding is a very powerful idea. &lt;a href=&quot;https://openai.com/blog/image-gpt/&quot;&gt;As researchers from OpenAI put it recently&lt;/a&gt;, it is “a universal unsupervised learning algorithm”. Indeed, OpenAI’s GPT-3—which is trained solely using next-step-prediction—can do all kinds of things it was never explicitly taught to do, like &lt;a href=&quot;https://www.gwern.net/GPT-3&quot;&gt;word arithmetic and Tom Swifty puns&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I suspect that more and more AI systems will have something like a predictive coding component built in. Forward predictive models are already needed for things like &lt;a href=&quot;http://rail.eecs.berkeley.edu/deeprlcourse-fa17/f17docs/lecture_9_model_based_rl.pdf&quot;&gt;model-based reinforcement learning&lt;/a&gt;, where a model of the environment is used to plan by simulating and optimizing over possible trajectories; so, why not take advantage of the rich representations learned by those predictive models, and use them as inputs to subsequent processing?&lt;/p&gt;

&lt;hr /&gt;</content><author><name>Loren Lugosch</name></author><category term="sequence modeling" /><summary type="html">The name “predictive coding” has been applied to a number of engineering techniques and scientific theories. All these techniques and theories involve predicting future observations from past observations, but what exactly is meant by “coding” differs in each case. Here is a quick tour of some flavors of “predictive coding” and how they’re related.</summary></entry><entry><title type="html">A contemplation of $\text{logsumexp}$</title><link href="https://lorenlugosch.github.io/posts/2020/06/logsumexp/" rel="alternate" type="text/html" title="A contemplation of $\text{logsumexp}$" /><published>2020-06-30T00:00:00-07:00</published><updated>2020-06-30T00:00:00-07:00</updated><id>https://lorenlugosch.github.io/posts/2020/06/logsumexp</id><content type="html" xml:base="https://lorenlugosch.github.io/posts/2020/06/logsumexp/">&lt;p&gt;$\text{logsumexp}$ is an interesting little function that shows up surprisingly often in machine learning. Join me in this post to shed some light on $\text{logsumexp}$: where it lives, how it behaves, and how to interpret it.&lt;/p&gt;

&lt;h2 id=&quot;what-is-textlogsumexp&quot;&gt;What is $\text{logsumexp}$?&lt;/h2&gt;
&lt;p&gt;Let $\mathbf{x} \in \mathbb{R}^n$. $\text{logsumexp}(\mathbf{x})$ is defined as:&lt;/p&gt;
&lt;center&gt;$$\begin{eqnarray*} \text{logsumexp}(\mathbf{x}) = \text{log} \left( \sum_{i} \text{exp}(x_i) \right). \end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;Numerically, $\text{logsumexp}$ is similar to $\text{max}$: in fact, it’s sometimes called the “smooth maximum” function. For example:&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*} \text{max}([1,2,3]) = 3 \end{eqnarray*}$$&lt;/center&gt;

&lt;center&gt;$$\begin{eqnarray*} \text{logsumexp}([1,2,3]) = 3.4076 \end{eqnarray*}$$&lt;/center&gt;

&lt;h2 id=&quot;examples-of-textlogsumexp&quot;&gt;Examples of $\text{logsumexp}$&lt;/h2&gt;

&lt;p&gt;Here are some places in machine learning where $\text{logsumexp}$ is used.&lt;/p&gt;

&lt;h3 id=&quot;softmax-classifiers&quot;&gt;Softmax classifiers&lt;/h3&gt;

&lt;p&gt;In a softmax classifier, the likelihood of label $i$ is defined as:&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*} p_{\theta}(i | \mathbf{l}) = \text{softmax}(\mathbf{l})_i, \end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;where $\mathbf{l}$ is the vector of logits (unnormalized scores for each label).&lt;/p&gt;

&lt;p&gt;Softmax classifiers are trained by minimizing the negative log-likelihood loss:

    &lt;div&gt;
	$$\begin{align*}
    -\text{log } p_{\theta}(i | \mathbf{l}) &amp;= -\text{log} \left( \text{softmax}(\mathbf{l})_i \right) \\
    &amp;= -\text{log} \left( \text{exp}(l_i) / \sum_j \text{exp}(l_j) \right) \\
    &amp;= -\text{log} \left( \text{exp}(l_i)\right) +\text{log} \left(\sum_j \text{exp}(l_j) \right) \\
    &amp;= -l_i + \text{logsumexp}(\mathbf{l}).
\end{align*}$$
    &lt;/div&gt;
&lt;/p&gt;

&lt;p&gt;Our friend $\text{logsumexp}$ appears in the last line.&lt;/p&gt;

&lt;h3 id=&quot;global-pooling&quot;&gt;Global pooling&lt;/h3&gt;

&lt;p&gt;For sequence classification tasks, it is usually necessary to map a variable length sequence of feature vectors to a single feature vector to be able to use something like a softmax classifier. To obtain a single vector, global pooling operations like max pooling or mean pooling can be used. Another aggregation method, which is less commonly used but has some of the advantages of both mean and max pooling, is $\text{logsumexp}$ pooling. (See &lt;a href=&quot;https://ronan.collobert.com/pub/matos/2016_wordaggr_interspeech.pdf&quot;&gt;this paper&lt;/a&gt; for an example.)&lt;/p&gt;

&lt;h3 id=&quot;latent-alignment-models&quot;&gt;Latent alignment models&lt;/h3&gt;

&lt;p&gt;In latent alignment models, like &lt;a href=&quot;https://dl.acm.org/doi/pdf/10.1145/1143844.1143891&quot;&gt;connectionist temporal classification (CTC)&lt;/a&gt;, dynamic programming is used to add up the probability of all possible alignments of an input sequence $\mathbf{x}$ to an output sequence $\mathbf{y}$ to train the model.&lt;/p&gt;

&lt;p&gt;The dynamic programming algorithm in CTC uses the following recursion:&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*} \alpha_{s,t} = (\alpha_{s,t-1} + \alpha_{s-1,t-1} + \alpha_{s-2,t-1}) \cdot p_{\theta}(y_s | x_t)\end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;This algorithm multiplies a long chain of probabilities, and so will underflow when aligning long sequences (like speech). Instead, we can run the algorithm in the log domain, in which case the recursion becomes:&lt;/p&gt;

    &lt;div&gt;
	$$\begin{align*}
    \text{log}(\alpha_{s,t})  &amp;=  \text{log}(\alpha_{s,t-1} + \alpha_{s-1,t-1} + \alpha_{s-2,t-1}) + \text{log } p_{\theta}(y_s | x_t)\\
    &amp;=  \text{log}(\text{exp}(\text{log }\alpha_{s,t-1}) + \text{exp}(\text{log }\alpha_{s-1,t-1}) + \text{exp}(\text{log }\alpha_{s-2,t-1})) + \text{log } p_{\theta}(y_s | x_t)\\
    &amp;=  \text{logsumexp}([\text{log }\alpha_{s,t-1},\text{log }\alpha_{s-1,t-1}, \text{log }\alpha_{s-2,t-1}]) + \text{log } p_{\theta}(y_s | x_t)\\
\end{align*}$$
    &lt;/div&gt;

&lt;p&gt;Similar recursions using $\text{logsumexp}$ can be derived for the forward-backward algorithm used in &lt;a href=&quot;https://lorenlugosch.github.io/posts/2020/01/hmm/&quot;&gt;Hidden Markov Models&lt;/a&gt; and &lt;a href=&quot;https://arxiv.org/pdf/1211.3711.pdf&quot;&gt;Transducer&lt;/a&gt; models.&lt;/p&gt;

&lt;p&gt;Fun fact: if we replace the $\text{logsumexp}$ with $\text{max}$, we get the Viterbi algorithm, which gives us the score of the single most likely alignment (&lt;a href=&quot;https://arxiv.org/pdf/2002.00876.pdf&quot;&gt;and if we backpropagate, the alignment itself&lt;/a&gt;).&lt;/p&gt;

&lt;h2 id=&quot;some-properties-of-textlogsumexp&quot;&gt;Some properties of $\text{logsumexp}$&lt;/h2&gt;

&lt;p&gt;$\text{logsumexp}$ has some useful properties. It is:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;em&gt;Convex.&lt;/em&gt; If you can pose your machine learning problem as a &lt;a href=&quot;https://web.stanford.edu/~boyd/cvxbook/&quot;&gt;convex optimization&lt;/a&gt; problem, you can solve it quickly and reliably.&lt;/li&gt;
  &lt;li&gt;&lt;em&gt;Differentiable everywhere.&lt;/em&gt; This is nice to have if your optimization algorithm is picky and doesn’t like functions with non-differentiable points, like $\text{max}$. For those picky algorithms, we can approximate $\text{max}$ using $\text{logsumexp}$.&lt;/li&gt;
  &lt;li&gt;&lt;em&gt;Associative.&lt;/em&gt; So:&lt;/li&gt;
&lt;/ul&gt;

    &lt;div&gt;
	$$\begin{align*}
    \text{logsumexp}([a,b,c,d]) &amp;= \text{logsumexp}([ \\

    &amp; \text{logsumexp}([a,b]), \\
    &amp; \text{logsumexp}([c,d]) \\
    ]). \\ 
\end{align*}$$
    &lt;/div&gt;

&lt;ul&gt;
  &lt;li&gt;(Hence, it can be computed in just $\text{log}_2(n)$ timesteps using a &lt;a href=&quot;https://en.wikipedia.org/wiki/Reduction_Operator&quot;&gt;parallel reduction&lt;/a&gt;, where $n$ is the length of the vector we’re $\text{logsumexp}$ing.)&lt;/li&gt;
  &lt;li&gt;&lt;em&gt;Close to&lt;/em&gt; $\text{max}$. In what sense “close”? It is bounded as follows:&lt;/li&gt;
&lt;/ul&gt;
&lt;center&gt;$$\begin{eqnarray*} \text{max}(\mathbf{x}) \leq \text{logsumexp}(\mathbf{x})  \leq \text{max}(\mathbf{x}) + \text{log}(n) \end{eqnarray*}$$&lt;/center&gt;

&lt;h2 id=&quot;the-textlogsumexp-trick&quot;&gt;The $\text{logsumexp}$ trick&lt;/h2&gt;

&lt;p&gt;Computing $\text{log} \left( \sum_{i} \text{exp}(x_i) \right)$ directly is numerically unstable because of the $\text{exp}$. Try the following in numpy:&lt;/p&gt;

&lt;pre style=&quot;font-size:13px&quot;&gt;
x = np.array([7000,8000,9000])
np.log(np.sum(np.exp(x)))
&lt;/pre&gt;

&lt;p&gt;From the approximation $\text{max}(\mathbf{x}) \approx \text{logsumexp}(\mathbf{x})$, we know that the result should be a little &lt;a href=&quot;https://www.youtube.com/watch?v=SiMHTK15Pik&quot;&gt;over 9000&lt;/a&gt;—but if you run this code, the result will be infinity because of overflow.&lt;/p&gt;

&lt;p&gt;Instead, to compute $\text{logsumexp}$, use the following trick:&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*}
\text{logsumexp}(\mathbf{x}) = \text{log} \left( \sum_{i} \text{exp}(x_i - \text{max}(\mathbf{x})) \right) + \text{max}(\mathbf{x}).
\end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;(See &lt;a href=&quot;https://www.xarg.org/2016/06/the-log-sum-exp-trick-in-machine-learning/&quot;&gt;this post&lt;/a&gt; for the proof that the trick works.)&lt;/p&gt;

&lt;p&gt;Now we won’t get an overflow because we’re taking the $\text{exp}$ of $[-2000,-1000,0]$ instead of $[7000,8000,9000]$. If we now run this instead:&lt;/p&gt;

&lt;pre style=&quot;font-size:13px&quot;&gt;
x = np.array([7000,8000,9000])
np.log(np.sum(np.exp(x - x.max()))) + x.max()
&lt;/pre&gt;

&lt;p&gt;we’ll get what we expect.&lt;/p&gt;

&lt;h2 id=&quot;takeaways&quot;&gt;Takeaways&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;$\text{logsumexp}$ is everywhere!&lt;/li&gt;
  &lt;li&gt;You can make certain dense equations easier to digest by identifying instances of $\text{logsumexp}$ and mentally replacing them with $\text{max}$.
    &lt;blockquote&gt;
      &lt;p&gt;For example, from the discussion of softmax classifiers above, you now know that the loss for a classifier is just the difference between the maximum of the logits (roughly) and the logit for the right answer!&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;When computing $\text{logsumexp}$, use the “logsumexp trick”. (Or just check if your library already has a numerically stable $\text{logsumexp}$ function, &lt;a href=&quot;https://pytorch.org/docs/stable/torch.html#torch.logsumexp&quot;&gt;as PyTorch does&lt;/a&gt;.)&lt;/li&gt;
&lt;/ul&gt;

&lt;hr /&gt;</content><author><name>Loren Lugosch</name></author><category term="machine learning" /><summary type="html">$\text{logsumexp}$ is an interesting little function that shows up surprisingly often in machine learning. Join me in this post to shed some light on $\text{logsumexp}$: where it lives, how it behaves, and how to interpret it.</summary></entry><entry><title type="html">Notebook: Fun with Hidden Markov Models</title><link href="https://lorenlugosch.github.io/posts/2020/01/hmm/" rel="alternate" type="text/html" title="Notebook: Fun with Hidden Markov Models" /><published>2020-01-28T00:00:00-08:00</published><updated>2020-01-28T00:00:00-08:00</updated><id>https://lorenlugosch.github.io/posts/2020/01/hmm</id><content type="html" xml:base="https://lorenlugosch.github.io/posts/2020/01/hmm/">&lt;p&gt;I’ve written a notebook introducing Hidden Markov Models (HMMs) with a PyTorch implementation of the forward algorithm, the Viterbi algorithm, and training a model on a text dataset—check it out &lt;a href=&quot;https://colab.research.google.com/drive/1IUe9lfoIiQsL49atSOgxnCmMR_zJazKI&quot;&gt;here&lt;/a&gt;!&lt;/p&gt;</content><author><name>Loren Lugosch</name></author><category term="sequence modeling" /><summary type="html">I’ve written a notebook introducing Hidden Markov Models (HMMs) with a PyTorch implementation of the forward algorithm, the Viterbi algorithm, and training a model on a text dataset—check it out here!</summary></entry><entry><title type="html">An introduction to sequence-to-sequence learning</title><link href="https://lorenlugosch.github.io/posts/2019/02/seq2seq/" rel="alternate" type="text/html" title="An introduction to sequence-to-sequence learning" /><published>2019-02-19T00:00:00-08:00</published><updated>2019-02-19T00:00:00-08:00</updated><id>https://lorenlugosch.github.io/posts/2019/02/seq2seq</id><content type="html" xml:base="https://lorenlugosch.github.io/posts/2019/02/seq2seq/">&lt;p&gt;Many interesting problems in artificial intelligence can be described in the following way:&lt;/p&gt;
&lt;blockquote&gt;
  &lt;p&gt;Map a sequence of inputs $\mathbf{x}$ to the correct sequence of outputs $\mathbf{y}$.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Speech recognition is one example: the goal is to map an audio signal $\mathbf{x}$ (a sequence of real-valued audio samples) to the correct text transcript $\mathbf{y}$ (a sequence of letters). Other examples are machine translation, image captioning, and speech synthesis.&lt;/p&gt;

&lt;p&gt;This post is a tutorial introduction to sequence-to-sequence learning, a method for using neural networks to solve these “sequence transduction” problems. In this post, I will:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;show you why these problems are interesting and challenging&lt;/li&gt;
  &lt;li&gt;give a detailed description of sequence-to-sequence learning—or “seq2seq”, as the cool kids call it&lt;/li&gt;
  &lt;li&gt;walk through an example seq2seq application (with PyTorch code)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The only background you need to read this is basic knowledge of how to train and use neural networks for classification problems.&lt;/p&gt;

&lt;h2 id=&quot;the-problem&quot;&gt;The problem&lt;/h2&gt;
&lt;p&gt;Let’s start by defining the general problem of sequence transduction a little more carefully.&lt;/p&gt;

&lt;p&gt;Let $\mathbf{x} = \{x_1, x_2, \dots x_T\}$ be the input sequence and $\mathbf{y} = \{y_1, y_2, \dots, y_U\}$ be the output sequence, where $x_t \in \mathcal{S}_x$, $y_u \in \mathcal{S}_y$, and $\mathcal{S}_x$ and $\mathcal{S}_y$ are the sets of possible things that each $x_t$ and $y_u$ can be, respectively.&lt;/p&gt;

&lt;p&gt;In speech recognition, $\mathcal{S}_x$ would be the set of real numbers, $\mathcal{S}_y$ would be the set of letters, $T$ might be on the order of thousands, and $U$ might be on the order of tens.&lt;/p&gt;

&lt;p&gt;We’ll assume that the inputs and outputs are random variables, and the value of $T$ and $U$ may vary between examples of $\mathbf{x}$ and $\mathbf{y}$—for example, not all sentences have the same length. The fact that $U$ may vary is part of what makes these problems challenging: we need to guess how long the output should be, and often this can’t simply be inferred from the input length.&lt;/p&gt;

&lt;p&gt;The goal is to find the best function $f(\mathbf{x})$ for mapping $\mathbf{x}$ to $\mathbf{y}$. But what does “best” mean here?&lt;/p&gt;

&lt;p&gt;For a classification problem, the “best” $f(\mathbf{x})$ is the one with the highest accuracy, i.e. the lowest probability of guessing an output $\mathbf{\hat{y}}$ which is not equal to the true $\mathbf{y}$.&lt;/p&gt;

&lt;p&gt;Similarly, we can use “accuracy” as a performance measure if the output is a sequence: if the guess $\mathbf{\hat{y}}$ is at all different from the true output $\mathbf{y}$, the output is incorrect. That is, if $\mathbf{y} = \text{“hello”}$ and $\mathbf{\hat{y}} = \text{“helo”}$, then $\mathbf{\hat{y}}$ is incorrect. This is also called the “$0$-$1$ loss”: $0$ if $\mathbf{\hat{y}} = \mathbf{y}$, $1$ otherwise.&lt;/p&gt;

&lt;p&gt;The simple $0$-$1$ loss is not a realistic performance measure: an output like $\text{“lskdjfl”}$ should really be regarded as less accurate than $\text{“helo”}$ if the true output is $\text{“hello”}$.&lt;/p&gt;

&lt;p&gt;In practice, we probably care more about some other performance measure, such as the word error rate (speech recognition) or the BLEU score (machine translation, image captioning). But the $0$-$1$ loss is often a good approximation or surrogate for other performance measures, and it doesn’t require any domain-specific knowledge to apply, so let’s assume that the $0$-$1$ loss is what we’re trying to optimize.&lt;/p&gt;

&lt;h2 id=&quot;the-ideal-solution&quot;&gt;The ideal solution&lt;/h2&gt;
&lt;p&gt;Suppose that we have access to a magical genie who can tell us $p(\mathbf{y}|\mathbf{x})$, the probability that the correct output sequence is $\mathbf{y}$ given that the input is $\mathbf{x}$, for any $\mathbf{x}$ and $\mathbf{y}$. In that case, what would be the best function $f(\mathbf{x})$ to use?&lt;/p&gt;

&lt;p&gt;Since we’re trying to minimize the probability of making an error (guessing a $\mathbf{\hat{y}}$ which is not equal to the correct $\mathbf{y}$), the following function is the best choice:&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*}\mathbf{\hat{y}} = f(\mathbf{x}) = \underset{\mathbf{y}}{\text{argmax }} p(\mathbf{y}|\mathbf{x})\end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;Alas! We usually don’t have a magical genie who can tell us $p(\mathbf{y}|\mathbf{x})$. What we can do instead is fit a model $p_{\theta}(\mathbf{y}|\mathbf{x})$ to some training data, and if we’re lucky, that model will be close to the true distribution $p(\mathbf{y}|\mathbf{x})$.&lt;/p&gt;

&lt;h2 id=&quot;modeling-pmathbfy&quot;&gt;Modeling $p(\mathbf{y})$&lt;/h2&gt;

&lt;p&gt;Before we consider implementing the model $p_{\theta}(\mathbf{y}|\mathbf{x})$, let’s start with a slightly easier but closely related problem: implementing a model $p_{\theta}(\mathbf{y})$ that is close to $p(\mathbf{y})$.&lt;/p&gt;

&lt;p&gt;What $p(\mathbf{y})$ represents is the probability of observing a particular output sequence $\mathbf{y}$, independent of whatever the input sequence might be.&lt;/p&gt;

&lt;p&gt;Intuitively, what does this mean? Well, consider the two following sequences of words:&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*}\mathbf{y}_1 = \text{I ate food}\end{eqnarray*}$$&lt;/center&gt;

&lt;center&gt;$$\begin{eqnarray*}\mathbf{y}_2 = \text{I eight food}\end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;The probability of observing $\mathbf{y}_1$ should be higher than the probability of observing $\mathbf{y}_2$, since $\mathbf{y}_1$ is a meaningful English sentence, and $\mathbf{y}_2$ is nonsensical. This fact could be useful in speech recognition to figure out which of the two phrases the person actually said, since both of these phrases sound the same if you say them out loud.&lt;/p&gt;

&lt;p&gt;Likewise, consider the two following sequences of letters:&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*}\mathbf{y}_1 = \text{florpy}\end{eqnarray*}$$&lt;/center&gt;

&lt;center&gt;$$\begin{eqnarray*}\mathbf{y}_2 = \text{fhqhwgads}\end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;Although these two sequences are both not real English words, the first one $\mathbf{y}_1$ really could be an English word, whereas the second one $\mathbf{y}_2$ just looks like someone banging on a keyboard. If we were to train a model $p_{\theta}(\mathbf{y})$ on English text, we would probably find in this case that $p_{\theta}(\mathbf{y}_1) &amp;gt; p_{\theta}(\mathbf{y}_2)$.&lt;/p&gt;

&lt;p&gt;To assign a probability to any sequence, we could just use a gigantic lookup table, with one entry for every possible $\mathbf{y}$. The problem is there are usually too many possible sequences for this to be feasible—in fact, there may be infinitely many possible sequences.&lt;/p&gt;

&lt;p&gt;Instead, let’s define a model that computes the probability of a sequence by dividing it into a number of simpler probabilities using the chain rule of probability.&lt;/p&gt;

&lt;p&gt;The chain rule of probability says that, for two random variables A and B, the probability of A &lt;em&gt;and&lt;/em&gt; B is equal to the probability of A &lt;em&gt;given&lt;/em&gt; B, multiplied by the probability of B:&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*} p(A,B) = p(A|B) \cdot p(B) \end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;That’s one way to factorize $p(A,B)$—we could also factorize it like this:&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*} p(A,B) = p(B|A) \cdot p(A) \end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;Moreover, you can apply the chain rule for as many random variables as you want:&lt;/p&gt;

    &lt;div&gt;
	$$\begin{align*} 
	p(A,B,C) &amp;= p(A|B,C) \cdot p(B,C) \\ 
	&amp;= p(A|B,C) \cdot p(B|C) \cdot p(C) 
	\end{align*}$$
    &lt;/div&gt;

&lt;p&gt;Remember that our sequence $\mathbf{y}$ is a collection of random variables $y_1, y_2, y_3, \dots, y_U$. So let’s use the chain rule to write out $p_{\theta}(\mathbf{y})$ as the product of the probability of each of these variables, given the previous ones:&lt;/p&gt;

    &lt;div&gt;
$$\begin{eqnarray*} p_{\theta}(\mathbf{y}) &amp;=&amp; p_{\theta}(y_1, y_2, y_3, \dots, y_U)\\ &amp;=&amp; \overset{U}{\underset{u=1}{\prod}} p_{\theta}(y_u|y_{u-1},y_{u-2},\dots,y_{1}) \end{eqnarray*}$$
    &lt;/div&gt;

&lt;p&gt;This $p_{\theta}(y_u|y_{u-1},y_{u-2},\dots,y_{1})$ term means “the probability of the next element, given what came before”. A model that implements $p_{\theta}(y_u|y_{u-1},y_{u-2},\dots,y_{1})$ is sometimes called a “next step predictor” or “language model”. When you are texting someone, and your phone suggests the next word for you as you are typing, it uses a model like this.&lt;/p&gt;

&lt;p&gt;We can implement this “next step” probability $p_{\theta}(y_u|y_{u-1},y_{u-2},\dots,y_{1})$ using a neural network&lt;sup id=&quot;fnref:NPLM&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:NPLM&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;. The neural network takes as input the previous $y_u$’s and predicts the next $y_u$. This is shown in the figure below:&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/seq2seq/RNN_model.png&quot; style=&quot;max-width:50%&quot; /&gt;&lt;/center&gt;

&lt;p&gt;For example, if $\mathcal{S}_y$ is the 26 letters of the alphabet, then $y_{u-1},y_{u-2},\dots,y_{1}$ would be the previous letters in the sequence, and the neural network would have a softmax output of size 26 representing the probability of the next letter given the previous letters.&lt;/p&gt;

&lt;p&gt;A common choice for the neural network is a recurrent neural network (RNN), which is what is shown in the diagram here.&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;An RNN, if you are not familiar, is a neural network with memory. At each timestep, the RNN takes in an input vector $i$ and its current state vector $h$, and outputs an updated state vector $h := f(i, h)$. Through the state $h$, the RNN can remember what inputs it has seen so far ($i_1, i_2, \dots$). The RNN can assign probabilities to different classes based on its state using a softmax classifier: $p(\text{class }c) = \text{softmax}(Wh + b)_c$&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;To get the probability of the entire sequence, the chain rule tells us to multiply each $p_{\theta}(y_u|y_{u-1},y_{u-2},\dots,y_{1})$ term together. In other words, just multiply the softmax outputs together.&lt;/p&gt;

&lt;p&gt;For example, let’s say that $\mathcal{S}_y$ is the three letters $\{a,b,c\}$, and we want to calculate the probability of the sequence $ba$. In this case, the neural network would have a softmax output of size 3. If we have:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;$p_{\theta}(y_1) = [0.3, 0.4, 0.3]$, and&lt;/li&gt;
  &lt;li&gt;$p_{\theta}(y_2|y_1=b) = [0.5, 0.3, 0.2]$,&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;then $p_{\theta}(ba) = p(a|b) p(b) = 0.5 \cdot 0.4 = 0.2$.&lt;/p&gt;

&lt;p&gt;One final ingredient we need for a complete model $p_{\theta}(\mathbf{y})$ is a special element called the “end-of-sequence” element. Each sequence needs to end with “end-of-sequence”. Predicting “end-of-sequence” along with all the other elements of $\mathcal{S}_y$ allows the model to implicitly predict the &lt;em&gt;length&lt;/em&gt; of the sequence.&lt;/p&gt;

&lt;h2 id=&quot;modeling-pmathbfymathbfx&quot;&gt;Modeling $p(\mathbf{y}|\mathbf{x})$&lt;/h2&gt;

&lt;p&gt;We’re not just interested in $p(\mathbf{y})$: what we really want to model is $p(\mathbf{y}|\mathbf{x})$, the probability of an output sequence &lt;em&gt;given&lt;/em&gt; a particular input sequence.&lt;/p&gt;

&lt;p&gt;The simplest way to condition the output on the input is to split the model into an encoder RNN and a decoder RNN, where the encoder RNN converts the input sequence into a single vector that is used to “program” the decoder RNN.&lt;sup id=&quot;fnref:EncDec&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:EncDec&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;The encoder RNN reads the input sequence element-by-element. As the encoder reads each input element, it updates its state. The final state of the encoder after it has read the entire input sequence represents a fixed-length encoding of the input sequence. This encoding then becomes the initial state of the decoder RNN:&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/seq2seq/encoder_decoder.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;

&lt;p&gt;Hopefully, the encoding (a single vector) will contain all the information in the input needed for the decoder to accurately model the correct output. Then we just proceed to calculate $p_{\theta}(\mathbf{y}|\mathbf{x})$ as we would calculate $p_{\theta}(\mathbf{y})$, by multiplying the probabilities of all the $y_u$’s together:&lt;/p&gt;

    &lt;div&gt;
$$\begin{eqnarray*} p_{\theta}(\mathbf{y}|\mathbf{x}) &amp;=&amp; p_{\theta}(y_1, y_2, y_3, \dots, y_U|\mathbf{x})\\ &amp;=&amp; \overset{U}{\underset{u=1}{\prod}} p_{\theta}(y_u|y_{u-1},y_{u-2},\dots,y_{1},\mathbf{x}) \end{eqnarray*}$$
    &lt;/div&gt;

&lt;p&gt;Here’s a question you might ask, looking at the diagram of the model: Given that an RNN has state, and can remember what it has already outputted, why we do we need to feed $y_1, y_2, \dots$ into the decoder to compute the probability of the next output $y_u$?&lt;/p&gt;

&lt;p&gt;There are two good reasons.&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;First, if we don’t make the output explicitly conditioned on the previous outputs, we are implicitly saying that the outputs are independent, which may not be a good assumption. Consider transcribing a recording of someone saying “triple A”. There are two valid transcriptions: $\text{AAA}$ and $\text{triple A}$. If the first output $y_1$ is $\text{A}$, then we can be certain that the second output $y_2$ will be $\text{A}$.&lt;sup id=&quot;fnref:AAA&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:AAA&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;ul&gt;
  &lt;li&gt;Second, feeding in previous outputs allows us to use a feedforward model (which does not have state) for the decoder. A feedforward model can be much faster to train, since each $p_{\theta}(y_u|y_{u-1},y_{u-2},\dots,y_{1},\mathbf{x})$ term can be computed in parallel.&lt;sup id=&quot;fnref:convS2S&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:convS2S&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Although we’ve been writing $p_{\theta}(\mathbf{y}|\mathbf{x})$, it’s usually better to work with the log probability $\text{log } p_{\theta}(\mathbf{y}|\mathbf{x})$ instead, for a few reasons. First, it is often easier to work with sums than it is to work with products. Recall that $\text{log }(a \cdot b) = \text{log }a + \text{log }b$. Thus, if you take the log of $p_{\theta}(\mathbf{y}|\mathbf{x})$, it becomes a sum instead of a product:&lt;/p&gt;

    &lt;div&gt;
$$\begin{eqnarray*} \text{log } p_{\theta}(\mathbf{y}|\mathbf{x}) &amp;=&amp; \text{log } p_{\theta}(y_1, y_2, y_3, \dots, y_U|\mathbf{x})\\ &amp;=&amp; \overset{U}{\underset{u=1}{\sum}} \text{log } p_{\theta}(y_u|y_{u-1},y_{u-2},\dots,y_{1},\mathbf{x}) \end{eqnarray*}$$
    &lt;/div&gt;

&lt;p&gt;This prevents multiplying together a bunch of numbers smaller than 1, which could cause an underflow. It also makes training the model using gradient descent easier, since minimizing a sum of terms is easier than minimizing a product of terms.&lt;sup id=&quot;fnref:product&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:product&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;5&lt;/a&gt;&lt;/sup&gt; And finally, since many useful probability distributions have an $\text{exp}(\cdot)$ term (including softmax, Gaussian, and Poisson), and $\text{log }\text{exp}(a) = a$, taking the log may transform $p_{\theta}$ into a simpler form.&lt;/p&gt;

&lt;h2 id=&quot;learning&quot;&gt;Learning&lt;/h2&gt;
&lt;p&gt;Now that we have a way of computing $p_{\theta}(\mathbf{y}|\mathbf{x})$ using a neural network, we need to train the model, i.e. find $\theta$ such that the model distribution $p_{\theta}(\mathbf{y}|\mathbf{x})$ matches the true distribution $p(\mathbf{y}|\mathbf{x})$. As usual with neural networks, we can do that by minimizing a loss function.&lt;/p&gt;

&lt;p&gt;Many seq2seq papers don’t explicitly write out a loss function, though. Instead, they will just say something like  “we use maximum likelihood” or “we minimize the negative log likelihood”. Here, we will see how this translates to a particular loss function that you can implement.&lt;/p&gt;

&lt;p&gt;Suppose that you have a training set $\mathcal{T}$ composed of ($\mathbf{x}^i, \mathbf{y}^i$) pairs. If the training examples are considered to be fixed, and you think of $p_{\theta}(\mathbf{y}|\mathbf{x})$ as a function of the parameters $\theta$, then we call $p_{\theta}(\mathbf{y}^i|\mathbf{x}^i)$ the “likelihood” of the training example ($\mathbf{x}^i, \mathbf{y}^i$). In maximum likelihood estimation, the model parameters $\theta$ are learned by maximizing the likelihood of the entire training set, $L_{\theta}(\mathcal{T})$, which is the product of the likelihoods of all the training examples:&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*}L_{\theta}(\mathcal{T}) = \overset{|\mathcal{T}|}{\underset{i=1}{\prod}} p_{\theta}(\mathbf{y}^i|\mathbf{x}^i)\end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;If these ($\mathbf{x}^i, \mathbf{y}^i$) samples are independent and identically distributed (i.i.d.), then as the number of samples increases, maximum likelihood estimation gives you a model that is closer and closer to the true distribution.&lt;/p&gt;

&lt;p&gt;As mentioned earlier, it’s easier to maximize a sum than it is to maximize a product, so we take the log, and maximize that instead. (The log function is monotonic—that is, $\text{log }a &amp;gt; \text{log }b$ implies that $a &amp;gt; b$—so maximizing the log likelihood is equivalent to maximizing the likelihood.) Also, in machine learning it’s often more natural to think of &lt;em&gt;minimizing&lt;/em&gt; a loss function, so we minimize the &lt;em&gt;negative&lt;/em&gt; log likelihood:&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*}-\text{log } L_{\theta}(\mathcal{T}) = -\overset{|\mathcal{T}|}{\underset{i=1}{\sum}} \text{log } p_{\theta}(\mathbf{y}^i|\mathbf{x}^i)\end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;If we expand the summed term, we get:&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*}-\text{log } L_{\theta}(\mathcal{T}) = -\overset{|\mathcal{T}|}{\underset{i=1}{\sum}} \overset{U^i}{\underset{u=1}{\sum}} \text{log } p_{\theta}(y_{u}^i|y_{u-1}^i,y_{u-2}^i,\dots,y_{1}^i,\mathbf{x}^i)\end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;In other words, for every example in the dataset, we sum up the negative log probability of the correct output at each timestep, given the previous correct outputs.&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Feeding the previous &lt;strong&gt;correct&lt;/strong&gt; outputs into the model during training, as opposed to the model’s own predictions, is called “teacher forcing”. A long time ago (in deep learning years, which are like dog years), it was thought that teacher forcing is bad, and you should sometimes sample previous outputs from the model’s output distribution.&lt;sup id=&quot;fnref:Scheduled&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:Scheduled&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;6&lt;/a&gt;&lt;/sup&gt; Nowadays, this is less common, and big parallelizable seq2seq models like the Transformer&lt;sup id=&quot;fnref:Transformer&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:Transformer&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;7&lt;/a&gt;&lt;/sup&gt; rely on teacher forcing to go fast. Also, the name “teacher forcing” makes it sound like a hack, but really it’s the right way to apply maximum likelihood!&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;To minimize the negative log likelihood loss, you can use stochastic gradient descent (SGD), just like in a regular classification problem.&lt;/p&gt;

&lt;h2 id=&quot;inference&quot;&gt;Inference&lt;/h2&gt;
&lt;p&gt;So far, we’ve described how you can use a neural network to compute $p_{\theta}(\mathbf{y}|\mathbf{x})$, and how to train the model using maximum likelihood.&lt;/p&gt;

&lt;p&gt;The question remains: How do we generate an output? That is, how do we find $\underset{\mathbf{y}}{\text{argmax }} p_{\theta}(\mathbf{y}|\mathbf{x})$, given a new $\mathbf{x}$?&lt;/p&gt;

&lt;p&gt;The brute force solution is an exhaustive search: just compute $p_{\theta}(\mathbf{y}|\mathbf{x})$ for every possible $\mathbf{y}$ and pick the $\mathbf{y}$ with the highest probability. That’s what exactly what we do for classification problems.&lt;/p&gt;

&lt;p&gt;However, unlike a typical classification problem, where you might have a thousand classes, in any practical sequence prediction problem there will be astronomically many possible output sequences, so an exhaustive search is infeasible.&lt;/p&gt;

&lt;p&gt;That means we need a new ingredient: an efficient search algorithm. The goal of the search is to approximately find $\mathbf{y}^* \approx \underset{\mathbf{y}}{\text{argmax }} p_{\theta}(\mathbf{y}|\mathbf{x})$.&lt;/p&gt;

&lt;p&gt;(Note: for historical reasons&lt;sup id=&quot;fnref:Jelinek&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:Jelinek&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;8&lt;/a&gt;&lt;/sup&gt;, the search process is also often called “decoding”. I’m not a big fan of this terminology because “decoding” already means several other things in machine learning.)&lt;/p&gt;

&lt;p&gt;We will consider two search algorithms:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;greedy search&lt;/li&gt;
  &lt;li&gt;beam search&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;1) Greedy search.&lt;/strong&gt; A greedy search works as follows: at each step, pick the top output of the network, and feed this output back into the network. In other words, for each timestep $u$, pick $y_u^* = \underset{y_u}{\text{argmax }}  p_{\theta}(y_u | y_{u-1}^*, y_{u-2}^*, \dots, y_{1}^*, \mathbf{x})$.&lt;/p&gt;

&lt;p&gt;The search can continue until an “end-of-sequence” is predicted, or until a maximum number of steps has occurred. An example of a greedy search is shown in the diagram below:&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/seq2seq/greedy_search.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;

&lt;p&gt;Because inference requires making a prediction, and feeding it back in to make the next prediction, this type of model is called “autoregressive” (“auto”=”self”, “regress”=”predict”).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2) Beam search.&lt;/strong&gt; Greedy searching is fast, but you can show that it will not always find the most likely output sequence. As an example, suppose that:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;$p_{\theta}(a|\mathbf{x}) = 0.4$&lt;/li&gt;
  &lt;li&gt;$p_{\theta}(b|\mathbf{x}) = 0.6$&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;and&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;$p_{\theta}(aa|\mathbf{x}) = 0.4$&lt;/li&gt;
  &lt;li&gt;$p_{\theta}(ab|\mathbf{x}) = 0.0$&lt;/li&gt;
  &lt;li&gt;$p_{\theta}(ba|\mathbf{x}) = 0.35$&lt;/li&gt;
  &lt;li&gt;$p_{\theta}(bb|\mathbf{x}) = 0.25$&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here, a greedy search would pick $b$, then $a$, and thus return $ba$, which has probability $0.35$. But the most likely sequence is actually $aa$, which has a probability of $0.4$.&lt;/p&gt;

&lt;p&gt;We can get better results if we delay decisions about keeping a particular output until we have considered some future outputs. Beam search is one way of doing this.&lt;/p&gt;

&lt;p&gt;In a beam search, we maintain a list (“beam”) of $B$ likely sequences (“hypotheses”). At each step, for each hypothesis, we compute the top $B$ outputs, and append them to the hypothesis. Now we have $B^2$ hypotheses. Of these, we prune the beam down to the top $B$, and then we continue to the next step.&lt;/p&gt;

&lt;p&gt;An example of a beam search with $B=3$ is shown below.&lt;sup id=&quot;fnref:CTC&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:CTC&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;9&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;center&gt;&lt;img src=&quot;https://lorenlugosch.github.io/images/seq2seq/beam_search.png&quot; style=&quot;max-width:75%&quot; /&gt;&lt;/center&gt;

&lt;p&gt;The algorithm returns the top $B$ hypotheses found. If your goal is just to estimate $\mathbf{y}$, you would just keep the top hypothesis, but for some applications it may also be useful to keep the rest of the hypotheses.&lt;/p&gt;

&lt;p&gt;Notice that if $B = 1$, the beam search is equivalent to a greedy search. Also, if $B = \infty$, the beam search becomes an exhaustive search.&lt;/p&gt;

&lt;p&gt;Beam search, too, is not guaranteed to find the most likely output sequence, but the wider you make the beam, the smaller the chance of a search error.&lt;sup id=&quot;fnref:searcherror&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:searcherror&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;10&lt;/a&gt;&lt;/sup&gt; The tradeoff is that you must re-run the neural network $B$ times, since you need to feed outputs back in.&lt;/p&gt;

&lt;h2 id=&quot;attention&quot;&gt;Attention&lt;/h2&gt;
&lt;p&gt;We now have a complete method for doing sequence-to-sequence learning.&lt;/p&gt;

&lt;p&gt;Unfortunately, if you apply the method exactly as described above—using an encoder RNN to map the input to a fixed-length vector consumed by the decoder RNN—it will not work for long sequences. The problem is that it is difficult to compress the entire input sequence into a single fixed-length vector.&lt;sup id=&quot;fnref:Cho&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:Cho&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;11&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;Attention&lt;sup id=&quot;fnref:Bahdanau&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:Bahdanau&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;12&lt;/a&gt;&lt;/sup&gt; is a mechanism that removes this fixed-length bottleneck. With attention, the decoder does not rely on a single vector to represent the input; instead, at every decoding step, it “looks” at a different part of the input using a weighted sum.&lt;/p&gt;

&lt;p&gt;A more detailed introduction to the various forms of attention can be found &lt;a href=&quot;https://lilianweng.github.io/lil-log/2018/06/24/attention-attention.html&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;

&lt;h2 id=&quot;a-toy-task&quot;&gt;A toy task&lt;/h2&gt;
&lt;p&gt;Let’s use a toy task to test out our method. Consider the following string:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;“Mst ppl hv lttl dffclty rdng ths sntnc”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This string is an example&lt;sup id=&quot;fnref:Shannon&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:Shannon&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;13&lt;/a&gt;&lt;/sup&gt; of how natural language is highly redundant or predictable: it’s generated by removing all the vowels from a normal sentence, and you can still understand the meaning.&lt;/p&gt;

&lt;p&gt;If human intelligence can infer the missing vowels, then maybe artificial intelligence can as well! We will train a sequence-to-sequence model to take as input a vowel-less sentence (“Mst ppl”) and output the sentence with the correct vowels re-inserted (“Most people”).&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;This toy task is a bit easier to work with than a task like speech recognition or translation, in which you need lots of labelled data and lots of tricks to get something to work. A tutorial which covers the actual useful task of translating from French to English using PyTorch can be found &lt;a href=&quot;https://pytorch.org/tutorials/intermediate/seq2seq_translation_tutorial.html&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2 id=&quot;running-the-experiment&quot;&gt;Running the experiment&lt;/h2&gt;
&lt;p&gt;We can easily generate a dataset for the task of inferring missing vowels using existing text. I used the text of “War and Peace”&lt;sup id=&quot;fnref:Karpathy&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:Karpathy&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;14&lt;/a&gt;&lt;/sup&gt;: the input sequences are lines from the text with all the vowels removed, and the target output sequences are just the original lines. I also added the Penn Treebank (PTB) dataset, a commonly used dataset of news articles for language modeling experiments, to give the training data a little variety.&lt;/p&gt;

&lt;p&gt;To run the code for yourself, or to train a model for filling in missing vowels on a new dataset, the code and a pre-trained model for this experiment can be found &lt;a href=&quot;https://github.com/lorenlugosch/infer_missing_vowels&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;We will first try training the simple encoder-decoder described above, without an attention mechanism. Here’s the result when we run a new sentence through the model:&lt;/p&gt;

&lt;pre style=&quot;font-size:13px&quot;&gt;
&lt;b&gt;input:&lt;/b&gt; Mst ppl hv lttl dffclty rdng ths sntnc.
&lt;b&gt;truth:&lt;/b&gt; Most people have little difficulty reading this sentence.
&lt;b&gt;guess:&lt;/b&gt; Mostov played with Prince Andrew and strengthers and
&lt;/pre&gt;

&lt;p&gt;Oh dear! The search starts off strong—it correctly outputs “Most”—but then it gets distracted and tries to fill in the name “Rostov” (the name of a character in “War and Peace”). The next word, “played”, at least starts with the right letter, the “p” in “people”, but after that, the output really goes off the rails.&lt;/p&gt;

&lt;p&gt;If we let the simple encoder-decoder model train a lot longer, the results get a bit better, but we still find bizarre mistakes like this:&lt;/p&gt;

&lt;pre style=&quot;font-size:13px&quot;&gt;
&lt;b&gt;input:&lt;/b&gt; th dy bfr, nmly, tht th cmmndr-n-chf
&lt;b&gt;truth:&lt;/b&gt; the day before, namely, that the commander-in-chief
&lt;b&gt;guess:&lt;/b&gt; the day before, manling that the mimicinacying Freemason
&lt;/pre&gt;

&lt;p&gt;So let’s add in that fancy attention mechanism I mentioned and see if that helps:&lt;/p&gt;

&lt;pre style=&quot;font-size:13px&quot;&gt;
&lt;b&gt;input:&lt;/b&gt; Mst ppl hv lttl dffclty rdng ths sntnc.
&lt;b&gt;truth:&lt;/b&gt; Most people have little difficulty reading this sentence.
&lt;b&gt;guess:&lt;/b&gt; Most people have little difficulty riding this sentence.
&lt;/pre&gt;

&lt;p&gt;Much better! But still not completely correct.&lt;/p&gt;

&lt;p&gt;Let’s look at the beam (the $B$ hypotheses found by the beam search) and the beam scores (the hypotheses’ log probabilities):&lt;/p&gt;
&lt;pre style=&quot;font-size:13px&quot;&gt;
Most people have little difficulty riding this sentence.   | -2.98
&lt;b&gt;Most people have little difficulty reading this sentence.  | -3.28&lt;/b&gt;
Most people have little difficulty roading this sentence.  | -3.79
Most people have little difficulty riding those sentence.  | -3.81
Most people have little difficulty reading those sentence. | -4.11
Most people have little difficulty riding these sentence.  | -4.16
Most people have little difficulty reading these sentence. | -4.45
Most people have little difficulty roading those sentence. | -4.60
&lt;/pre&gt;

&lt;p&gt;The correct answer does in fact appear in the beam (the 2nd hypothesis), but the model incorrectly assigns a higher probability to the hypothesis with “riding” instead of “reading”. Maybe with more/better training data this error would not occur, since “reading this sentence” ought to be a lot more probable in the training data than “riding this sentence”.&lt;/p&gt;

&lt;p&gt;Notice another peculiar aspect of the beam: it is ordered (roughly) from shortest to longest. Autoregressive models are biased towards shorter output sequences!&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Why? At each timestep, to compute the probability of a sequence, we multiply it (or add it, in the log domain) by the probability of the next output, which is always less than 1 (less than 0, in the log domain), so the probability of the complete sequence keeps getting smaller. In fact, the only thing keeping the search from producing outputs of length 0 is the fact that the model needs to predict the “end-of-sequence” token, and from the training data the model learns to assign low probability to “end-of-sequence” until it makes sense. It’s easy for the model to learn that a sentence like “The.” is very unlikely, but comparing two plausible sequences of roughly the same length seems to be harder.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;How do people deal with the short-sequence bias in practice? Google Translate uses a variation&lt;sup id=&quot;fnref:GNMT&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:GNMT&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;15&lt;/a&gt;&lt;/sup&gt; of beam search in which the log probability is divided by a “length penalty” (“lp”), with a hyperparameter $\alpha$, that is computed as follows:&lt;/p&gt;

&lt;center&gt;$$\begin{eqnarray*} \text{lp}(\mathbf{y}) = \frac{(5 + |\mathbf{y}|)^{\alpha}}{(5 + 1)^{\alpha}}\end{eqnarray*}$$&lt;/center&gt;

&lt;p&gt;Yikes. How many TPU hours did they burn finding that formula? I hope we find a better way to mitigate the short-sequence bias!&lt;/p&gt;

&lt;h2 id=&quot;the-end&quot;&gt;The End&lt;/h2&gt;

&lt;p&gt;If you have any questions or if you find something wrong with this tutorial, please let me know.&lt;/p&gt;

&lt;p&gt;Check out the code and try it out! It’s fun to feed the model random inputs, like your name, and see what stuff it comes up with trying to fill in the gaps. The code is also written in such a way that it should not be too hard to adapt it to a new task.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;em&gt;Thanks to Christoph Conrads and Mirco Ravanelli for their feedback on the draft of this post.&lt;/em&gt;&lt;/p&gt;

&lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
  &lt;ol&gt;
    &lt;li id=&quot;fn:NPLM&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;See: Yoshua Bengio, Réjean Ducharme, Pascal Vincent, Christian Jauvin, “&lt;a href=&quot;http://www.jmlr.org/papers/volume3/bengio03a/bengio03a.pdf&quot;&gt;A Neural Probabilistic Language Model&lt;/a&gt;”, Journal of Machine Learning Research 3 (2003), 1137–1155. &lt;a href=&quot;#fnref:NPLM&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:EncDec&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Like many great ideas, the encoder-decoder model was invented independently and simultaneously by multiple groups. The paper that is usually cited, which invented the name “sequence-to-sequence learning”, is this one: Ilya Sutskever, Oriol Vinyals, and Quoc V. Le, &lt;a href=&quot;https://papers.nips.cc/paper/5346-sequence-to-sequence-learning-with-neural-networks.pdf&quot;&gt;“Sequence to sequence learning with neural networks”&lt;/a&gt;, NeurIPS 2014. &lt;a href=&quot;#fnref:EncDec&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:AAA&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;The “triple A” example comes from: William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals. “&lt;a href=&quot;https://arxiv.org/abs/1508.01211&quot;&gt;Listen, attend and spell: A neural network for large vocabulary conversational speech recognition&lt;/a&gt;”, ICASSP 2016. &lt;a href=&quot;#fnref:AAA&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:convS2S&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;For an example of this, see: Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. “&lt;a href=&quot;https://arxiv.org/abs/1705.03122&quot;&gt;Convolutional sequence to sequence learning&lt;/a&gt;”, ICML 2017. &lt;a href=&quot;#fnref:convS2S&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:product&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Why? Consider minimizing $y_{sum} = x_1 + x_2$ and $y_{prod} = x_1 \cdot x_2$. The derivative $\frac{dy_{sum}}{dx_1}$ is just $1$ (independent of $x_2$), whereas the derivative $\frac{dy_{prod}}{dx_1}$ is equal to $x_2$. If $x_2$ is very small, $x_1$ will be “held back” from changing easily if you try to minimize $y_{prod}$—imagine running a race if you are tied to a slow person by a rope. &lt;a href=&quot;#fnref:product&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:Scheduled&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;See: Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer, “&lt;a href=&quot;https://arxiv.org/abs/1506.03099&quot;&gt;Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks&lt;/a&gt;”, NeurIPS 2015. &lt;a href=&quot;#fnref:Scheduled&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:Transformer&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;See: Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin, “&lt;a href=&quot;https://arxiv.org/abs/1706.03762&quot;&gt;Attention is all you need&lt;/a&gt;”, NeurIPS 2017. &lt;a href=&quot;#fnref:Transformer&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:Jelinek&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Some of the early pioneers in speech recognition and machine translation, like Fred Jelinek, originally worked on digital communications and error-correcting codes, where it really does make sense to refer to searching for the best output as “decoding”. They didn’t bother to find a better word and kept saying “decoding” when they started working on speech recognition. It may not be entirely surprising that Jelinek once said “every time I fire a linguist, the performance of my speech recognizer goes up.” &lt;a href=&quot;#fnref:Jelinek&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:CTC&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Taken from: Awni Hannun, “&lt;a href=&quot;https://distill.pub/2017/ctc/&quot;&gt;Sequence Modeling with CTC&lt;/a&gt;”, Distill, 2017. &lt;a href=&quot;#fnref:CTC&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:searcherror&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;A search error happens when $p_{\theta}(\mathbf{y}|\mathbf{x}) &amp;gt; p_{\theta}(\mathbf{\hat{y}}|\mathbf{x})$, but $\mathbf{y}$ wasn’t found during the search. In other words, the model correctly assigns more probability to the correct output sequence than an incorrect output sequence, but the search just didn’t get a chance to evaluate the correct sequence. Another type of error is when the model assigns more probability to an incorrect sequence and picks that sequence as a result. &lt;a href=&quot;#fnref:searcherror&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:Cho&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;This paper discovered and diagnosed the problem: Kyunghyun Cho, Bart van Merrienboer, Dzmitry Bahdanau, Yoshua Bengio, “&lt;a href=&quot;https://arxiv.org/abs/1409.1259&quot;&gt;On the Properties of Neural Machine Translation: Encoder–Decoder Approaches&lt;/a&gt;”, SSST-8, 2014. &lt;a href=&quot;#fnref:Cho&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:Bahdanau&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;See: Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio, “&lt;a href=&quot;http://arxiv.org/abs/1409.0473&quot;&gt;Neural Machine Translation by Jointly Learning to Align and Translate&lt;/a&gt;”, ICLR 2015. &lt;a href=&quot;#fnref:Bahdanau&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:Shannon&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Supposedly given by Claude Shannon, although I haven’t been able to find the original reference where he wrote it. &lt;a href=&quot;#fnref:Shannon&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:Karpathy&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Provided by Andrej Karpathy in the code for &lt;a href=&quot;https://karpathy.github.io/2015/05/21/rnn-effectiveness/&quot;&gt;The Unreasonable Effectiveness of Recurrent Neural Networks&lt;/a&gt;. &lt;a href=&quot;#fnref:Karpathy&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:GNMT&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;See: Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, Jeffrey Dean, “&lt;a href=&quot;https://arxiv.org/abs/1609.08144&quot;&gt;Google’s neural machine translation system: Bridging the gap between human and machine translation&lt;/a&gt;”, arXiv preprint arXiv:1609.08144, 2016. &lt;a href=&quot;#fnref:GNMT&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
&lt;/div&gt;</content><author><name>Loren Lugosch</name></author><category term="sequence modeling" /><summary type="html">Many interesting problems in artificial intelligence can be described in the following way: Map a sequence of inputs $\mathbf{x}$ to the correct sequence of outputs $\mathbf{y}$.</summary></entry></feed>