<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Reinforcement Learning | Learning, Intelligence &#43; Signal Processing Lab</title>
    <link>http://lisplab.host.dartmouth.edu/tag/reinforcement-learning/</link>
      <atom:link href="http://lisplab.host.dartmouth.edu/tag/reinforcement-learning/index.xml" rel="self" type="application/rss+xml" />
    <description>Reinforcement Learning</description>
    <generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><lastBuildDate>Wed, 03 Aug 2022 22:57:42 -0500</lastBuildDate>
    <image>
      <url>http://lisplab.host.dartmouth.edu/media/sharing.png</url>
      <title>Reinforcement Learning</title>
      <link>http://lisplab.host.dartmouth.edu/tag/reinforcement-learning/</link>
    </image>
    
    <item>
      <title>nFlip : Deep Reinforcement Learning in Multiplayer FlipIt</title>
      <link>http://lisplab.host.dartmouth.edu/project/nflip-deep-reinforcement-learning-in-multiplayer-flipit/</link>
      <pubDate>Wed, 03 Aug 2022 22:43:03 -0500</pubDate>
      <guid>http://lisplab.host.dartmouth.edu/project/nflip-deep-reinforcement-learning-in-multiplayer-flipit/</guid>
      <description>&lt;p&gt;Reinforcement learning has shown much success in games such as chess, backgammon and Go. However, in most of these games, agents have full knowledge of the environment at all times. We describe a deep learning model that successfully maximizes its score using reinforcement learning in a game with incomplete and imperfect information. We apply our model to FlipIt &lt;sup id=&#34;fnref:1&#34;&gt;&lt;a href=&#34;#fn:1&#34; class=&#34;footnote-ref&#34; role=&#34;doc-noteref&#34;&gt;1&lt;/a&gt;&lt;/sup&gt;, a two-player game in which both players, the attacker and the defender, compete for ownership of a shared resource and only receive information on the current state upon making a move. Our model is a deep neural network combined with Q-learning and is trained to maximize the defender’s time of ownership of the resource. We extend FlipIt to a larger action-spaced game with the introduction of a new lower-cost move and generalize the model to multiplayer FlipIt.&lt;/p&gt;
&lt;div class=&#34;footnotes&#34; role=&#34;doc-endnotes&#34;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&#34;fn:1&#34;&gt;
&lt;p&gt;van Dijk, M., Juels, A., Oprea, A., Rivest, R.L. FlipIt : The Game of “Stealthy Takeover”. Journal of Cryptology 26,655-713 (2013).&amp;#160;&lt;a href=&#34;#fnref:1&#34; class=&#34;footnote-backref&#34; role=&#34;doc-backlink&#34;&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</description>
    </item>
    
    <item>
      <title>Using Game Theory and Reinforcement Learning to Predict the Future</title>
      <link>http://lisplab.host.dartmouth.edu/project/using-game-theory-and-reinforcement-learning-to-predict-the-future/</link>
      <pubDate>Wed, 03 Aug 2022 22:57:42 -0500</pubDate>
      <guid>http://lisplab.host.dartmouth.edu/project/using-game-theory-and-reinforcement-learning-to-predict-the-future/</guid>
      <description>&lt;p&gt;Baseball is a well known, repeated, finite, adversarial, stochastic game that has a massive amount of available data. On the other hand, Reinforcement Learning (RL) models take significant time and resources to train. By fusing Game Theory and RL, we are answering interesting questions such as &amp;ldquo;given a video of a pitch, can we compute the utility of a pitch given the desired location, resulting location, and setting?&amp;rdquo;&lt;/p&gt;
</description>
    </item>
    
  </channel>
</rss>
