<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <id>https://reganmian.net/</id>
  <title>Random Stuff that Matters</title>
  <updated>2020-01-31T14:27:28Z</updated>
  <link rel="alternate" href="https://reganmian.net/"/>
  <link rel="self" href="https://reganmian.net/blog/feed/index.xml"/>
  <author>
    <name>Stian Håklev</name>
    <uri>https://reganmian.net/blog/about-me/</uri>
  </author>
  <entry>
    <id>tag:reganmian.net,2020-01-31:/blog/blog/2020/01/31/note-taking-with-roam/</id>
    <title type="html">Taking notes on Tiago Forte's Just-in-Time Project Management</title>
    <published>2020-01-31T14:27:28Z</published>
    <updated>2020-01-31T14:27:28Z</updated>
    <link rel="alternate" href="https://reganmian.net/blog/2020/01/31/note-taking-with-roam/"/>
    <content type="html">&lt;h2&gt;Note taking and Roam&lt;/h2&gt;

&lt;p&gt;I was heavily invested in a note-taking infrastructure during my PhD, using Dokuwiki with a bunch of my own extensions to enable things like &lt;a href="/blog/2012/06/13/tag-extract-a-tool-to-automatically-restructure-textoutline-using-tags/"&gt;extracting and reorganizing content by tags&lt;/a&gt; and my own &lt;a href="/blog/2012/05/10/using-web-clipping-and-sidewiki-to-gather-and-synthesize-information/"&gt;web clipper&lt;/a&gt;. However, this workflow was so brittle that I couldn't even manage to transfer it easily to a new laptop, and in the years since that intense bout of literature review (early stage PhD), I've muddled through with a variety of tools, DevonThink, Google Docs, NValt etc. &lt;/p&gt;

&lt;p&gt;Recently, I came across Roam Research, and was very excited by how they are combining some of the best ideas of wikis, refining my early idea of reorganizing content through tags, and hierarchical outliners (which I had seen, but never really paid that much attention to - turns out this plus backlinks enables really powerful things). For me, it was &lt;a href="https://www.youtube.com/watch?v=Hw2kJF_kxjE"&gt;an interview&lt;/a&gt; that Tiago Forte did with the founder of Roam, Conor White-Sullivan (&lt;a href="https://roamresearch.com/#/app/stian-research/page/Ank-ovFDT"&gt;my notes&lt;/a&gt;), which "sealed the deal". I was very impressed by his vision, and have been digging into Roam and the community around it, ever since. &lt;/p&gt;

&lt;p&gt;Since then, I've been thinking about the best ways of using Roam, inspired by some of the use cases people have shared (see some in &lt;a href="https://github.com/roam-unofficial/awesome-roam"&gt;Awesome Roam&lt;/a&gt;). I've tried &lt;a href="https://pauljacobson.me/2018/03/10/getting-stuff-done-with-interstitial-journaling/"&gt;interstitial journaling&lt;/a&gt;, keeping track of projects and goals, capturing interesting resources, etc. &lt;/p&gt;

&lt;h2&gt;Tiago Forte's concept of Just-in-Time Project Management&lt;/h2&gt;

&lt;p&gt;I've been following Tiago Forte for a few months now, through his prolific blog posts, newsletters, podcast appearances and tweets. He clearly embodies a "productivity first" or "learning through writing" approach, but he has a lot of useful things to say. One of the concepts I kept circling around was "intermediate deliverables" or "intermediate packages". One of the ways Roam works (possibly mirroring how our brains work) is that you can begin "circling in" on a concept without having to really define it. If you keep tagging a certain concept whenever you hear about it, once you finally visit that page, even though you've never written anything, you can see all the different contexts that you've referred to it, and in this way already get what someone called a "pointilistic portrait" of the concept. &lt;/p&gt;

&lt;p&gt;So I've been thinking about getting back to blogging or at least sharing more informal writing publicly (this is the first blog post in three years), and also about how I could leverage better all the writing that happens anyway (in long e-mails, notes, etc). I decided that I wanted to properly understand what Tiago meant by this concept, and came across a link to his &lt;a href="https://praxis.fortelabs.co/series/just-in-time-project-management/"&gt;series on Just-in-Time Project Management&lt;/a&gt;. This is a list of about 22 blog posts, which are only for subscribers, so I paid 10$, and began reading.
Since I really want to retain and organize this information, I opened one window in Roam and one window to read, and began taking notes. As I went along, I came to really enjoy the way Roam let's you collapse and expand sections, which enables you to keep an overview of all the different sections, even though it's actually many many pages, and made it easy to see how the subsequent information fit in, allowing me to go back and modify or add to previous structures etc, mirroring how I want to understand the ideas, rather than the order in which it is presented.&lt;/p&gt;

&lt;h2&gt;Recording the process of taking notes&lt;/h2&gt;

&lt;p&gt;After having processed about five blog posts, I had the idea to just turn on a screen recorder, and let it watch me work. I thought it might be fun to go back and look at my process, and perhaps others would find it fun - at least in a sped-up version. It was also helpful to keep focused, since I didn't want to have to go back and edit later, so I really resisted the urge to switch context, check email etc. 
I have also looked into, and regularly use, many tools for capturing content, such as Hypothes.is for web annotation, Scrivener as a "reading buffer", etc. But this time, I really enjoyed doing all of the processing synchronously. I know some people like to dump all of the text into Roam first, and then process, but I prefer doing at least a pre-selection (using annotations), and in this case, I did most of the processing synchronously. 
As I read, I change wordings, use bold and italic, and try to organize the notes so that it will be easy to come back to. I add lot's of links to other concepts (both ones I have pages on already, and ones I don't). I also add my own thoughts and questions, which I usually highlight in yellow to make them stand out.&lt;/p&gt;

&lt;p&gt;When I had finished, I did one last quick pass over the whole text, re-reading to see if I could detect new connections etc. I could have done much more work to reorganize or integrate here, but I'll leave that for the next time I revisit this text. 
haring the process and the outcome&lt;/p&gt;

&lt;p&gt;I've made all of &lt;a href="https://roamresearch.com/#/app/stian-research/page/b6izvXt8X"&gt;my notes&lt;/a&gt; public (note that this is not the same database as I use privately, so many of the links which will have content in my private db will go to empty pages here), and if you are curious, you can also see the process by which these notes were made. Below is a 1000x sped up version (one minute), with links to a 100x and a real-time version.&lt;/p&gt;

&lt;iframe width="560" height="315" src="https://www.youtube.com/embed/8iOWZSQmF2A" frameborder="0" allow="accelerometer; autoplay; encrypted-media; gyroscope; picture-in-picture" allowfullscreen&gt;&lt;/iframe&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.youtube.com/watch?v=tMP9zgGhpM8"&gt;Full length version (1h 37 min)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.youtube.com/watch?v=8iOWZSQmF2A"&gt;100x speedup&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;Take-aways from Tiago's writings&lt;/h2&gt;

&lt;p&gt;Since I've already taken so detailed notes, and interleaved them with many personal observations, I will not try to summarize them here (although perhaps that would have been useful for my learning/retention). Suffice to say that I really like some of his approaches, even though I do sometimes feel like a lot of productivity advice comes from people who are self-employed as Youtubers, bloggers etc, and they naturally have a very different way of organizing their time, and evaluating "Return-on-Attention" than others. &lt;/p&gt;

&lt;p&gt;I started out struggling with note taking as a student, first as a bachelor student primarily occupied with learning and passing exams, as well as doing just enough reading to be able to finish a paper (typically something like five sources), but also doing an Honour's Thesis with a large amount of material. Later I did an MA degree, with mostly seminar classes, and again a large research task at the end, and finally a PhD, where you are of course not only doing readings to prepare you for a momentous task several years in the future, but also thinking about publishing papers, participating in the scholarly community, etc.&lt;/p&gt;

&lt;p&gt;(At that time, I also spent a lot of time thinking about how academics should collaborate, better formats for academic publishing, etc.)
And currently, I'm employed at the Minerva Project, where I have an interesting job combining pure software development, with product work (specs, architecture, pedagogical design), and even some pure research. I still need to keep up with the literature, yet goals and criteria for evaluation are very different.&lt;/p&gt;

&lt;p&gt;That said, I think I would have very much appreciated reading this back when I was a PhD student (and certainly finding Roam!), but it's possible that the time I spent in a tech company learning about modern management approaches also helps me better appreciate some of his writings.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;PS: This blog post was &lt;a href="https://roamresearch.com/#/app/stian-research/page/nOw3XeKDN"&gt;authored in Roam&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <id>tag:reganmian.net,2017-01-05:/blog/blog/2017/01/05/from-seminar-to-lecture-to-mooc/</id>
    <title type="html">From Seminar to Lecture to MOOC</title>
    <published>2017-01-05T16:41:48Z</published>
    <updated>2017-01-05T16:41:48Z</updated>
    <link rel="alternate" href="https://reganmian.net/blog/2017/01/05/from-seminar-to-lecture-to-mooc/"/>
    <content type="html">&lt;p&gt;This summer, I successfully defended my PhD thesis &lt;a href="https://infoscience.epfl.ch/record/224081/files/From%20seminar%20to%20lecture%20to%20MOOC%20-%20Ha%CC%8Aklev%202016.pdf"&gt;"From Seminar to Lecture to MOOC: Scripting and Orchestration at Scale"&lt;/a&gt; at Ontario Institute Studies in Education, University of Toronto. &lt;/p&gt;

&lt;p&gt;I have always been interested in different ways of publishing and sharing my research. With my &lt;a href="/blog/2008/09/20/mencerdaskan-bangsa-an-inquiry-into-the-phenomenon-of-taman-bacaan-in-indonesia/"&gt;BA thesis on community libraries in Indonesia&lt;/a&gt;, I had it translated to Indonesian, and also made it into an ebook.&lt;/p&gt;

&lt;p&gt;I have an &lt;a href="/blog/the-chinese-national-top-level-courses-project/"&gt;extensive page documenting my experiments&lt;/a&gt; with my MA thesis on Chinese Open Educational Courses. I made the thesis available in a number of formats, serialized it on my blog, had it translated into Chinese, and also published all of my raw notes (open notebook).&lt;/p&gt;

&lt;p&gt;For my PhD thesis, I had big ambitions about having the finished thesis be a live document linking to my notes on all the papers cited (using my &lt;a href="/wiki/researchr/start/"&gt;Researchr system&lt;/a&gt;), but although I wrote my initial drafts in Markdown (with Scrivener), they ended up in Word for collaborative editing and track changes.&lt;/p&gt;

&lt;p&gt;I then wanted to create a fancy landing page, with links to resources from the thesis, multiple formats, talks, etc. But after landing in Lausanne, where I am doing a post-doc at EPFL, I never got around to it. Now it has been half a year since I defended my thesis, it is still not public at T-Space, the institutional repository in Toronto, and I decided it's time to set it loose.&lt;/p&gt;

&lt;p&gt;I still hope to come back to add more information, formats, details etc. But for now, here are some essential links:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://infoscience.epfl.ch/record/224081"&gt;The thesis itself&lt;/a&gt; (&lt;a href="https://infoscience.epfl.ch/record/224081/files/From%20seminar%20to%20lecture%20to%20MOOC%20-%20Ha%CC%8Aklev%202016.pdf"&gt;PDF&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;&lt;a href="/files/INQ101%20poster%20MOOC%20camp.pdf"&gt;Poster I presented at EPFL MOOC camp&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.slideshare.net/StianHklev/a-principled-approach-to-the-design-of-collaborative-mooc-curricula"&gt;Slides from talk at EMOOCS 2017&lt;/a&gt;, May 22, 2017&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://infoscience.epfl.ch/record/226317/files/EMOOCs_2017_paper_54.pdf"&gt;A principled approach to the design of collaborative MOOC curricula&lt;/a&gt;, paper for EMOOCS 2017&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.youtube.com/watch?v=DuLrm-9N3lQ"&gt;Recording of talk given at UPF Barcelona on thesis, and current Orchestration Graph project&lt;/a&gt;, April 4, 2017&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://mediaserver.unige.ch/play/97162"&gt;Recording of talk given in French at TECFA, University of Geneva&lt;/a&gt;, October 18, 2106&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.youtube.com/watch?v=n31GfyfGyts"&gt;Recording of talk given at Knowledge Media Design Institute in Toronto&lt;/a&gt;, July 15, 2016&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.youtube.com/watch?v=Laf8WibtaGY"&gt;Recording of online talk given to DANCE network&lt;/a&gt;, July 7, 2016&lt;/li&gt;
&lt;/ul&gt;
</content>
  </entry>
  <entry>
    <id>tag:reganmian.net,2015-09-03:/blog/blog/2015/09/03/sending-and-receiving-email-with-elixir/</id>
    <title type="html">Sending and receiving email with Elixir</title>
    <published>2015-09-03T16:41:48Z</published>
    <updated>2015-09-03T16:41:48Z</updated>
    <link rel="alternate" href="https://reganmian.net/blog/2015/09/03/sending-and-receiving-email-with-elixir/"/>
    <content type="html">&lt;p&gt;Yesterday there was a Twitter discussion about adding a default mailer to Elixir Phoenix (&lt;a href="https://storify.com/houshuang/mail-in-phoenix"&gt;Storify&lt;/a&gt;). I've never used the Ruby on Rails mailer, but I got inspired to share my experience with both sending and receiving emails in a recent project.&lt;/p&gt;

&lt;p&gt;This summer, I ran &lt;a href="https://github.com/houshuang/survey"&gt;a Phoenix server&lt;/a&gt; that provided interactive content for &lt;a href="https://www.edx.org/course/teaching-technology-inquiry-open-course-university-torontox-inq101x"&gt;an EdX MOOC&lt;/a&gt;. One peculiarity of the setup was that authentication happened through the LTI connection with EdX, so there was no need for a sign-up/login form, confirmation of email addresses etc. However, we used email for a number of other purposes.&lt;/p&gt;

&lt;h2&gt;Sending email&lt;/h2&gt;

&lt;p&gt;To send email, you need an SMTP server. While Elixir is perfectly capable of handling this on its own, I had heard a lot of stories of how email from private domains might hit spam-filters, and be difficult to configure correctly, so we looked around for an external provider. &lt;a href="https://aws.amazon.com/ses/"&gt;Amazon Simple Email Service&lt;/a&gt; is a no-frills service with a great price-point (currently 0.10$ per 1000 emails). It might not have all the bells and whistles of services like &lt;a href="http://www.mailgun.com/"&gt;Mailgun&lt;/a&gt; and &lt;a href="https://sendgrid.com/"&gt;Sendgrid&lt;/a&gt;, but for our purposes it worked great.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-09-03-sending-and-receiving-email-with-elixir_-_whole-01.png" alt=""&gt;
Amazon SES has a web API, but the easiest way to use it is simply to configure it as an outgoing SMTP server. Looking around, I found the &lt;a href="https://github.com/kamilc/mailman"&gt;Mailman library&lt;/a&gt;, which was easy enough to configure (&lt;a href="https://github.com/houshuang/survey/blob/master/lib/mailer.ex"&gt;my code&lt;/a&gt;). You simply define a struct like this (storing credentials in a config file):&lt;/p&gt;

&lt;pre&gt;&lt;code class="elixir"&gt;  def config do
    %Mailman.Context{
      config: %Mailman.SmtpConfig{
        username: Application.get_env(:mailer, :username),
        password: Application.get_env(:mailer, :password),
        relay: Application.get_env(:mailer, :relay),
        port: 587,
        tls: :always,
        auth: :always}
    }
  end
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;For each email to send, you construct a message struct like this (&lt;a href="https://github.com/houshuang/survey/blob/master/lib/mail.ex"&gt;many examples&lt;/a&gt;):&lt;/p&gt;

&lt;pre&gt;&lt;code class="elixir"&gt;  %Mailman.Email{
      subject: "#{entered} entered the collaborative workbench",
      from: "noreply@mooc.encorelab.org",
      to: [email],
      text: text,
      html: html }
  end
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;and then use &lt;code&gt;Mailman.deliver(email, config)&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;To generate the email contents, I used EEx templates which I manually precompiled. You can see &lt;a href="https://github.com/houshuang/survey/tree/master/data/mailtemplates"&gt;a list of templates&lt;/a&gt;, and &lt;a href="https://github.com/houshuang/survey/blob/master/lib/mail/templates.ex"&gt;the module that precompiles them&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;Asynchronicity and error handling&lt;/h2&gt;

&lt;p&gt;This worked beautifully when I was testing it with individual emails, but when I wanted to send out a large number of emails (for example personalized weekly updates), I found that some emails were silently dropped. It turns out that Amazon SES has a daily limit, which for me was lifted to 50,000 emails (more than I would ever need), but also a rate-limitation, which for me was 14 emails per second. Apparently Elixir is just too fast, and Amazon would just return an error whenever it went above that rate.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-09-03-sending-and-receiving-email-with-elixir_-_whole-02.png" alt=""&gt;&lt;/p&gt;

&lt;p&gt;To fix this, the first issue was to correctly detect errors, but Mailman seemed to return &lt;code&gt;:ok&lt;/code&gt; tuple no matter what. &lt;a href="https://github.com/kamilc/mailman/blob/30715434e20b6a06528d350df11ddb143def8e18/lib/mailman/external_smtp_adapter.ex#L15:L22"&gt;Looking into their code&lt;/a&gt;, I found that requests to deliver led to some processing of the structs, and then passing them on to &lt;code&gt;:gen_smtp_client.send_blocking&lt;/code&gt; using Task.async, and then immediately returning an &lt;code&gt;:ok&lt;/code&gt; tuple. Personally, I think that libraries should let users decide on their own concurrency strategy, and hiding the errors is a big problem. I have a &lt;a href="https://github.com/kamilc/mailman/pull/16/files"&gt;pull request&lt;/a&gt; pending which changes the function to call the &lt;code&gt;:gen_smtp_client&lt;/code&gt; function directly, and in the meantime I am using my own fork of Mailman.&lt;/p&gt;

&lt;p&gt;With this change, we do indeed get access to the error messages from Amazon, but we still need to handle catching errors, retrying etc. To solve this problem, I began writing my own task queue library, with the idea that it should be possible to specify different "Quality of Service" levels, including number of retries, max requests per second, time to wait per retry etc. My initial experiments need to be rewritten almost completely, and extracted into its own package, but even this simplified approach was used to send thousands of emails, scrape hundreds of websites every hours, etc. (&lt;a href="https://github.com/houshuang/survey/blob/master/web/models/job.ex"&gt;1&lt;/a&gt;, &lt;a href="https://github.com/houshuang/survey/blob/master/lib/job_worker.ex"&gt;2&lt;/a&gt;)&lt;/p&gt;

&lt;h2&gt;User-specific links&lt;/h2&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-09-03-sending-and-receiving-email-with-elixir_-_whole-03.png" alt=""&gt;&lt;/p&gt;

&lt;p&gt;As mentioned above, the application had no login screen, so we wanted to include links in the emails that let the users directly access the content linked to. I had already written &lt;a href="https://github.com/houshuang/survey/blob/master/web/models/cache.ex"&gt;a generic cache function&lt;/a&gt;, which stores any arbitrary Erlang term (using a &lt;a href="https://github.com/houshuang/survey/blob/master/lib/extensions/ecto_term.ex"&gt;custom Ecto type&lt;/a&gt;) returning a simple index (and checking for uniqueness). I used this, together with the &lt;a href="https://github.com/alco/hashids-elixir"&gt;hashids&lt;/a&gt; library to generate shorturls for specific URLs and user ids (&lt;a href="https://github.com/houshuang/survey/blob/master/lib/mail.ex"&gt;source&lt;/a&gt;):&lt;/p&gt;

&lt;pre&gt;&lt;code class="elixir"&gt;  def gen_url(id, url) do
    term = %{url: url, userid: id}
    id = Survey.Cache.store(term)
    hash = Hashids.encode(@hashid, id)
    @basename &amp;lt;&amp;gt; "/email/" &amp;lt;&amp;gt; hash
  end
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;When a user reached the /email/ controller, I simply extracted the id from the hashid, looked up in the cache store, retrieved the user from the database, and set the appropriate session variables to "log in" the user, before redirecting to the appropriate URL (&lt;a href="https://github.com/houshuang/survey/blob/master/web/controllers/email_controller.ex"&gt;source&lt;/a&gt;):&lt;/p&gt;

&lt;pre&gt;&lt;code class="elixir"&gt;def redirect(conn, %{"hash" =&amp;gt; hash}) do
    {:ok, [id]} = Hashids.decode(@hashid, String.strip(hash))
    struct = Survey.Cache.get(id)
    hash = (from f in Survey.User,
    where: f.id == ^struct.userid,
    select: f.hash) |&amp;gt; Repo.one
    conn
    |&amp;gt; put_session(:repo_id, struct.userid)
    |&amp;gt; put_session(:lti_userid, hash)
    |&amp;gt; ParamSession.redirect(struct.url)
  end
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;In each email, I also included a link to unsubscribe from all email notifications, or from the specific type of notification (for example, you might still want to receive weekly updates, but not notifications every time someone posted something), and I checked
whether someone had unsubscribed before sending out notifications. These links also had the userid encoded, and just displayed a page showing "Success!", rather than a form asking the user to enter email and password, etc.&lt;/p&gt;

&lt;p&gt;In addition to user-specific links, we could also customize the content of the emails based on user data in the database. Below is an example of an email that is both using a fancy HTML template (easy to do with EEx), and containing user-specific information:&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-09-03-sending-and-receiving-email-with-elixir_-_whole-04.png" alt=""&gt;&lt;/p&gt;

&lt;h2&gt;Receiving emails&lt;/h2&gt;

&lt;p&gt;I was very pleased with how well Elixir and Mailman was able to handle the sending of emails, and I managed to use quite a lot of Elixir features in implementing the tasks above (parallelism, worker queues, genservers, Ecto and custom types, EEx templates, etc). However, just sending emails isn't in itself anything extraordinary, any framework should support it at some level. Receiving emails seems like it would add a whole other level of complexity, though.&lt;/p&gt;

&lt;p&gt;In our project, we wanted to support small-group collaboration and communication between members who had never met, and were often separated by timezones. We already had embedded Etherpad, wiki and live chat, and we made it more likely that people would meet each other online, by sending out notifications whenever a group member entered the online environment. In addition, I had the idea of generating ad-hoc mailing lists for group discussion. However, one limitation was that we could not share users email addresses with anyone, due to privacy concerns.&lt;/p&gt;

&lt;p&gt;I looked into whether there were some mailing list system that we could install on the server, with an API that would let us dynamically add mailing lists and members etc. However, when I came across &lt;a href="https://github.com/Vagabond/gen_smtp"&gt;gen_smtp&lt;/a&gt; (which Mailman is built on top of), I realized that Erlang could handle this all by itself - and in a very simple manner.&lt;/p&gt;

&lt;p&gt;Since Mailman doesn't wrap the receiving functionality of gen_smtp, I had to dig into the Erlang code and documentation to understand how to set it up. While I love the flexibility you get by defining your own callback-based SMPT server, it does seem that most people simply want to receive emails and do something with the contents. Luckily, the repository has &lt;a href="https://github.com/Vagabond/gen_smtp/blob/master/src/gen_smtp_server.erl"&gt;an example implementation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I had a few hiccups figuring out how to set up the server. First, I somehow spent an unreasonable amount of time figuring out how to start the server from my Application configuration (translating the Erlang invocation into Elixir), although the result is very simple (&lt;a href="https://github.com/houshuang/survey/blob/master/lib/survey.ex"&gt;source&lt;/a&gt;):&lt;/p&gt;

&lt;pre&gt;&lt;code&gt; worker(:gen_smtp_server, [Mail.SMTPServer,
        Application.get_env(:mailer, :smtp_opts)]),
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;where the relevant line from the config file is:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;  smtp_opts: [[port: 3000, sessionoptions: [certfile: 'server.crt', keyfile: 'server.key']]]
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;The keyfiles are for TLS verification. The reason for putting the port number in a config file, rather than directly in the Application invocation, is that only one process can listen to the port at the same time - so if you have a production instance running, and you want to fire up an IEx to experiment with some code, they will both try to listen to the same code, and the second process will fail.&lt;/p&gt;

&lt;p&gt;The second issue was accessing port 25 on the server, which is the standard SMTP port. I didn't want to run my BEAM process as root, but by default, only root can bind to ports below 1024. I finally found the obscure answer in a &lt;a href="http://stackoverflow.com/questions/413807/is-there-a-way-for-non-root-processes-to-bind-to-privileged-ports-1024-on-l"&gt;Stackoverflow post&lt;/a&gt;, which is that executing &lt;code&gt;sudo /sbin/setcap 'cap_net_bind_service=ep' BINARY&lt;/code&gt; on a BINARY gives that binary the privilege of listening to ports below 1024. I ran this command against the BEAM binary, and everything worked well.&lt;/p&gt;

&lt;p&gt;I initially just tweaked the example module in Erlang to fit my purposes, but then decided to translate the file to Elixir, while removing a lot of the documentation comments and edge cases (which were mostly there to show the possibilities). The result was a very lean module that looks a lot less scary. You can see the whole &lt;a href="https://github.com/houshuang/survey/blob/master/lib/mail/smtp_server.ex"&gt;source&lt;/a&gt;, but I reorganized it so that the callbacks you are most likely to want to modify are at the top (shown below), and the callbacks that you should probably leave alone, are at the bottom (see source):&lt;/p&gt;

&lt;pre&gt;&lt;code class="elixir"&gt;defmodule Mail.SMTPServer do
    require Logger
    @behaviour :gen_smtp_server_session

    def init(hostname, session_count, address, options) do
        if session_count &amp;gt; 40 do
            Logger.warn('SMTP server connection limit exceeded')
            {:stop, :normal, ['421 ', hostname, ' is too busy to accept mail right now']}
        else
            banner = [hostname, ' ESMTP']
            state = %{}
            {:ok, banner, state}
        end
    end

    # possibility of rejecting based on _from_ address
    def handle_MAIL(from, state) do
        {:ok, state}
    end

    # possibility of rejecting based on _to_ address
    def handle_RCPT(to, state) do
        {:ok, state}
    end

    # getting the actual mail. all the relevant stuff is in data.
    def handle_DATA(from, to, data, state) do
        Mail.Receive.receive_message(from, to, data)
        {:ok, UUID.uuid5(:dns, "mooc.encorelab.org", :default), state}
    end
    [...]
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;All this does is give you the option of rejecting based on &lt;code&gt;from&lt;/code&gt;, and &lt;code&gt;to&lt;/code&gt; addresses, and then pass the entire message to a separate function for processing.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-09-03-sending-and-receiving-email-with-elixir_-_whole-06.png" alt="" class="img-left"&gt;
The way we used this is to include a form in the web interface to initiate an email conversation.&lt;/p&gt;

&lt;p&gt;Any emails sent through this form would be forwarded to all members of the group (who had not unsubscribed), with a unique &lt;code&gt;from&lt;/code&gt; address, which encoded the group id using hashids. Replies to this email would reach the server, and be resent to all group members, without revealing the email address of the original sender.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-09-03-sending-and-receiving-email-with-elixir_-_whole-05.png" alt=""&gt;&lt;/p&gt;

&lt;h2&gt;Conclusion&lt;/h2&gt;

&lt;p&gt;So there you have it. Social notifications, personalized emails, &lt;a href="/blog/2015/08/26/email-notifications-about-errors-in-elixir/"&gt;error messages by email&lt;/a&gt;, URLs that automatically log you in, ad-hoc mailing lists, and all with pure Elixir and Erlang. This whole experience really made me appreciate the Erlang/Elixir ecosystem. And for sending 13,500 emails in a month, I paid less than for a cup of coffee...&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-09-03-sending-and-receiving-email-with-elixir_-_whole-07.png" alt=""&gt;&lt;/p&gt;

&lt;p&gt;It would be interesting to discuss how we can improve the tooling around emails even more. Perhaps Mailman can be extended to also cover the server aspects of &lt;code&gt;gen_smtp&lt;/code&gt;, perhaps we need better documentation (hopefully this blog post can be a modest contribution), or better integration with other libraries. But I think we're building on a great foundation!&lt;/p&gt;
</content>
  </entry>
  <entry>
    <id>tag:reganmian.net,2015-08-26:/blog/blog/2015/08/26/email-notifications-about-errors-in-elixir/</id>
    <title type="html">Email notifications about errors in Elixir</title>
    <published>2015-08-26T16:41:48Z</published>
    <updated>2015-08-26T16:41:48Z</updated>
    <link rel="alternate" href="https://reganmian.net/blog/2015/08/26/email-notifications-about-errors-in-elixir/"/>
    <content type="html">&lt;p&gt;This summer, I ran &lt;a href="https://github.com/houshuang/survey"&gt;a Phoenix server&lt;/a&gt; that provided interactive content for &lt;a href="https://www.edx.org/course/teaching-technology-inquiry-open-course-university-torontox-inq101x"&gt;an EdX MOOC&lt;/a&gt;. Given the geographical distribution of students, the server was getting hit 24/7, and I wanted a quick way to get notified about any error messages. Elixir comes with a built-in &lt;a href="http://elixir-lang.org/docs/v1.0/logger/"&gt;logging framework&lt;/a&gt; that has several levels of logging (debug, info, warn, error). Any processes crashing will emit error logs, and you can also emit them manually from your own code.&lt;/p&gt;

&lt;p&gt;The stock logger only comes with a console backend. I ran the server through Ubuntu's &lt;code&gt;upstart&lt;/code&gt; system with a very simple configuration script, which simply set some environment variables and paths, and then used &lt;code&gt;PORT=4000 MIX_ENV=prod mix phoenix.server&lt;/code&gt; (I used Nginx to proxy the local port, and add SSL). I used &lt;code&gt;mosh&lt;/code&gt; and &lt;code&gt;tmux&lt;/code&gt; to keep a persistent connection, and &lt;code&gt;tail -f&lt;/code&gt; to watch the upstart log. On the screenshot below, I had several panes tracking both the access log (using &lt;a href="https://github.com/mneudert/plug_accesslog"&gt;plug_accesslog&lt;/a&gt;), the Phoenix log, and also a &lt;code&gt;grep&lt;/code&gt; on &lt;code&gt;error&lt;/code&gt; to see the latest error messages.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-08-26-email-notifications-about-errors-in-elixir_-_whole-01.png" alt=""&gt;&lt;/p&gt;

&lt;p&gt;Due to BEAM (the Erlang virtual machine) and &lt;a href="http://elixir-lang.org/getting-started/mix-otp/supervisor-and-application.html"&gt;OTP&lt;/a&gt;, Phoenix is remarkably stable. The only time the whole server went down would be because of a serious issue in the datacenter rendering the entire virtual machine inaccessible. Errors were typically bugs that would occur given a certain edge case (for example trying to render the group of a user who had not selected a group yet). They would crash the connection of the user who hit that edge case, but would not propagate up the stack. Thus it would be quite possible for an error to occur and not be noticed, since everything else would continue happily running.&lt;/p&gt;

&lt;p&gt;I decided I'd like to get an email every time an error was triggered. Luckily, the Elixir logger was built from the start as a pluggable system, and there are already a number of possible backends, feeding into &lt;a href="https://github.com/smpallen99/syslog"&gt;syslog&lt;/a&gt;, or several "monitoring as a service" systems like &lt;a href="https://github.com/honeybadger-io/honeybadger-elixir"&gt;Honeybadger&lt;/a&gt; and &lt;a href="https://github.com/crashdumpio/fink-elixir"&gt;Crashdump.IO&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I had already integrated e-mail into my system, using &lt;a href="https://github.com/crashdumpio/fink-elixir"&gt; Mailman &lt;/a&gt; to send email through &lt;a href="https://aws.amazon.com/ses/"&gt; AWS Simple Email Service &lt;/a&gt; for notifications, and looking at the examples in Elixir, and the external loggers, writing a custom back-end that sends email notifications turned out to be remarkably simple.&lt;/p&gt;

&lt;p&gt;Copying some code from the console backend, I ended up with &lt;a href="https://github.com/houshuang/survey/blob/master/lib/error_mail.ex"&gt;this module&lt;/a&gt;, with only 50+ lines of code, most of which handles config and setup. The real work is done in these two functions, first a &lt;code&gt;gen_event&lt;/code&gt; callback whenever a message with level &lt;code&gt;error&lt;/code&gt; is emitted, and which simply dispatches to the log_event function:&lt;/p&gt;

&lt;pre&gt;&lt;code class="elixir"&gt;def handle_event({:error, _gl, {_, msg, ts, md}}, state) do
  log_event(msg, ts, md, state)
  {:ok, state}
end
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;then the log_event function, which uses the &lt;a href="http://elixir-lang.org/docs/v1.0/logger/Logger.Formatter.html"&gt;Logger.Formatter&lt;/a&gt; module to get the string to mail out, briefly checks against two kinds of errors which I did not want logged by email, and then composes and sends an email, using exactly the same functionality that I use for sending a new user notification.&lt;/p&gt;

&lt;pre&gt;&lt;code class="elixir"&gt;defp log_event(msg, ts, md, {from, to_list, format, metadata}) do
  msg = Logger.Formatter.format(format, :error, msg, ts, Dict.take(md, metadata))
        |&amp;gt; IO.iodata_to_binary
  if !String.contains?(inspect(msg), "GenServer :job_worker") &amp;amp;&amp;amp;
    !String.contains?(inspect(msg), "bad_charset") do
    %Mailman.Email{
      from: from,
      to: to_list,
      text: msg}
    |&amp;gt; Survey.Mailer.deliver
  end
end
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;This worked perfectly.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-08-26-email-notifications-about-errors-in-elixir_-_whole-02.png" alt=""&gt;&lt;/p&gt;

&lt;p&gt;A trick when developing something like this, is to set the level to &lt;code&gt;warn&lt;/code&gt; while developing. Otherwise, if you have an error in your logging code, emitting an error log message from IEx to test your code, will trigger an error log message about the error in your logger, which will trigger another one, etc. This also means that in general, it is very important to make sure that the error logging code cannot itself trigger errors. If sending the email might cause an error, it might be better for the mailer to be a separate gen_event, and for the logger to simple send a message to it.&lt;/p&gt;

&lt;p&gt;There is no checking for duplicate errors in this system, which means that in a situation where a lot of errors are emitted, you will receive a lot of emails.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-08-26-email-notifications-about-errors-in-elixir_-_whole-03.png" alt=""&gt;&lt;/p&gt;

&lt;p&gt;This sometimes happened when I forgot to do a &lt;code&gt;mix deps.get&lt;/code&gt; after pushing new code that had an added dependency (I didn't use any kind of release management, but simply pushed code to Github, pulled it onto the server, and restarted the server). It once happened that an incoming e-mail triggered a panic with &lt;code&gt;gen_smtp&lt;/code&gt; because of a failing charset - this e-mail was resent every ten minutes for 24 hours, which resulted in many error messages. It also happened that the server was rebooted into single-user mode because of a datacenter failure, which meant it could not connect to the database. Upstart will automatically restart the system on a crash, again resulting in a stream of error mails.&lt;/p&gt;

&lt;p&gt;I'm not sure how one could handle this, because I think it makes sense to keep the error logging system as simple and robust as possible. A &lt;code&gt;gen_server&lt;/code&gt; which handled e-mailing and kept counters would help to some extent, but not in the case of multiple restarts. In that situation, you would need an external service which aggregated emails and had conditional dispatch, and then you are getting into the level of building up a monitoring and reporting infrastructure, either internally or as a service.&lt;/p&gt;

&lt;p&gt;However, this was a great first try - it did work very reliably, and it's also helpful to see that there is very little magic to &lt;code&gt;Logger&lt;/code&gt; and the logger backends.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <id>tag:reganmian.net,2015-08-21:/blog/blog/2015/08/21/elixir-prelude-packaging-up-utility-functions/</id>
    <title type="html">Elixir Prelude: Packing up utility functions</title>
    <published>2015-08-21T16:41:48Z</published>
    <updated>2015-08-21T16:41:48Z</updated>
    <link rel="alternate" href="https://reganmian.net/blog/2015/08/21/elixir-prelude-packaging-up-utility-functions/"/>
    <content type="html">&lt;h2&gt;Introduction&lt;/h2&gt;

&lt;p&gt;I've been spending the summer in China, head-deep in &lt;a href="http://elixir-lang.org/"&gt;Elixir&lt;/a&gt;, writing interactive scripts for an &lt;a href="https://www.edx.org/course/teaching-technology-inquiry-open-course-university-torontox-inq101x"&gt;EdX MOOC on inquiry and technology for teachers&lt;/a&gt; (&lt;a href="https://imgur.com/a/rAXVz"&gt;screenshots&lt;/a&gt;). I will have much more to share about that project later, but now I wanted to share a small Elixir package I released on Github called &lt;a href="https://github.com/houshuang/elixir-prelude"&gt;Prelude&lt;/a&gt; (&lt;a href="https://houshuang.github.io/prelude"&gt;documentation&lt;/a&gt;). This is a collection of utility functions for Elixir that I've extracted out of my code base.&lt;/p&gt;

&lt;h2&gt;Organization&lt;/h2&gt;

&lt;p&gt;I began by scattering these utility functions all over as small &lt;code&gt;defp&lt;/code&gt;s, but when I wanted to reuse them in other modules, I began gathering them. Initially I had a single file called Prelude - modelled after the Haskell Prelude, just because it was the only file I'd ever wholesale &lt;code&gt;import&lt;/code&gt; into my other modules.&lt;/p&gt;

&lt;p&gt;Eventually I began separating the code out into sub-modules, which meant I went from&lt;/p&gt;

&lt;pre&gt;&lt;code class="elixir"&gt;import Prelude

map_atomify(...)
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;to&lt;/p&gt;

&lt;pre&gt;&lt;code class="elixir"&gt;Prelude.Map.atomify(...)
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;which seems cleaner.&lt;/p&gt;

&lt;h2&gt;The point of gathering "trivial" functions&lt;/h2&gt;

&lt;p&gt;Many of these functions are very simple, but yet things you use all the time - gathering them together both means less repetition and greater legibility. Compare:&lt;/p&gt;

&lt;pre&gt;&lt;code class="elixir"&gt;x
|&amp;gt; Enum.map(fn {k, v} -&amp;gt; {v, k} end)
|&amp;gt; Enum.into(%{})
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;to&lt;/p&gt;

&lt;pre&gt;&lt;code class="elixir"&gt;x
|&amp;gt; Prelude.Map.switch
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;The code in the first version is short enough that I wouldn't bother making it a separate function, but the intent of the second snippet is much clearer. I might also forget to do &lt;code&gt;Enum.into(%{})&lt;/code&gt;, and somehow I always forget the &lt;code&gt;end&lt;/code&gt; in anonymous functions. So, some quick time saving.&lt;/p&gt;

&lt;h2&gt;Map functions&lt;/h2&gt;

&lt;p&gt;Some of the functions are a bit more involved, mainly found in the &lt;a href="http://houshuang.github.io/prelude/Prelude.Map.html"&gt;Prelude.Map&lt;/a&gt; module. I was inspired by the way Clojurists often work with deeply nested maps, sometimes keeping the entire application state in one &lt;code&gt;atom&lt;/code&gt;. This has also become relevant in Javascript-land with experiments around immutable state such as &lt;a href="https://github.com/Yomguithereal/baobab"&gt;Baobab&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I see deeply nested maps much less often in Elixir. Part of it might be the architecture (we use gen_servers, message passing, and ETS tables, rather than atom-swapping), but part of it might also be because of the syntax (&lt;code&gt;%{%{%{}}}&lt;/code&gt; is not pretty).&lt;/p&gt;

&lt;p&gt;However, Elixir already has &lt;code&gt;get_in&lt;/code&gt; and &lt;code&gt;put_in&lt;/code&gt; to reach into deeply nested objects. Prelude provides &lt;code&gt;deep_put&lt;/code&gt;, which is a parallel to &lt;code&gt;mkdir -p&lt;/code&gt;, it will insert an object and create any part of the missing path. If you put to a path that already exists, it will turn it into an array.&lt;/p&gt;

&lt;p&gt;Built on top of &lt;code&gt;deep_put&lt;/code&gt; is &lt;code&gt;group_by&lt;/code&gt;. Building the Phoenix app, I often had to extract data from the database using Ecto, and render it in templates. Let's say you have a table of students, listing their group affiliation, and which class their in. You wish to list all students, sorted by class, and then by group. With &lt;code&gt;group_by&lt;/code&gt;, you can get the data from Ecto like this:&lt;/p&gt;

&lt;pre&gt;&lt;code class="elixir"&gt;MyRepo.all(Student)
|&amp;gt; Prelude.Map.group_by([:class, :group])
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;The result will be a map of classes. Each class will point to a map of groups, and each group will point to a list of students in that group. Then in an EEx template, you can display the list like this:&lt;/p&gt;

&lt;pre&gt;&lt;code class="elixir"&gt;&amp;lt;%= for class &amp;lt;- @classes %&amp;gt;
  &amp;lt;h2&amp;gt;&amp;lt;%= class %&amp;gt;&amp;lt;/h2&amp;gt;
  &amp;lt;%= for group &amp;lt;- class %&amp;gt;
    &amp;lt;h3&amp;gt;&amp;lt;%= group %&amp;gt;&amp;lt;/hr&amp;gt;
    &amp;lt;%= for student &amp;lt;- group %&amp;gt;
      &amp;lt;li&amp;gt;&amp;lt;%= student %&amp;gt;&amp;lt;/li&amp;gt;
    &amp;lt;% end %&amp;gt;
  &amp;lt;% end %&amp;gt;
&amp;lt;% end %&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;h2&gt;Wrapping up&lt;/h2&gt;

&lt;p&gt;Anyway, the code is available, so feel free to browse around. I will be using this in my future projects, and probably adding functionality that is general enough to stand on its own. I'd love to add tests, and am happy to think about naming, APIs etc, if anyone else think this is interesting.&lt;/p&gt;

&lt;p&gt;It's also my first time to use &lt;a href="https://github.com/elixir-lang/ex_doc"&gt;ExDoc&lt;/a&gt; and &lt;a href="https://pages.github.com/"&gt;Github Pages&lt;/a&gt;, but both were a breeze.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <id>tag:reganmian.net,2015-04-19:/blog/blog/2015/05/19/fun-with-terminal-and-vim-macros/</id>
    <title type="html">Fun with terminal and Vim macros</title>
    <published>2015-04-19T16:41:48Z</published>
    <updated>2015-04-19T16:41:48Z</updated>
    <link rel="alternate" href="https://reganmian.net/blog/2015/05/19/fun-with-terminal-and-vim-macros/"/>
    <content type="html">&lt;p&gt;I always enjoy finding new ways of automating boring tasks, and today I had a few different small tasks that I was able to solve using some terminal-fu, and Vim macros. I really enjoyed how well it came together, so I thought I'd put together a short screencast to demonstrate.&lt;/p&gt;

&lt;p&gt;I had generated a number of discussion forum graphs for different courses. The processed data is in a json file in a directory named after the course. My task was to put an identical index.html file into each directory, create an index of the courses, with links, and modify all the landing pages to reflect the names of the courses. Not a hugely complicated task, but something you could run into day to day, and which would be quite cumbersome to do manually.&lt;/p&gt;

&lt;iframe width="600" height="500" src="https://www.youtube.com/embed/HesocMwPuJc" frameborder="0" allowfullscreen&gt;&lt;/iframe&gt;

&lt;p&gt;The video should be quite self-explanatory, but I've included some "show-notes" below.&lt;/p&gt;

&lt;p&gt;As far as I know, &lt;code&gt;ls&lt;/code&gt; doesn't have a built-in way to list only directories, so I found this snippet somewhere and added it to my .zshrc:&lt;/p&gt;

&lt;pre&gt;&lt;code class="bash"&gt;alias ld="ls -lht | grep '^d'"
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;&lt;code&gt;awk&lt;/code&gt; is a very powerful string manipulation tool, but in this case we're just using it to get the 9th column of each line (&lt;code&gt;awk '{ print $9 }'&lt;/code&gt;)&lt;/p&gt;

&lt;p&gt;&lt;code&gt;xargs&lt;/code&gt; is very useful to run commands on input arguments. We need &lt;code&gt;-n 1&lt;/code&gt;, otherwise it would put use all the arguments in the same command. With &lt;code&gt;-n 1&lt;/code&gt;, it executes the command once for each input argument. If we didn't want to put the arguments last, we could do something like &lt;code&gt;echo 1 2 3 | xargs -n 1 - J % echo % hi&lt;/code&gt;.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <id>tag:reganmian.net,2015-02-20:/blog/blog/2015/02/20/openedx-and-lti-pedagogical-scripts-and-sso/</id>
    <title type="html">OpenEdX and LTI: Pedagogical scripts and SSO</title>
    <published>2015-02-20T16:41:48Z</published>
    <updated>2015-02-20T16:41:48Z</updated>
    <link rel="alternate" href="https://reganmian.net/blog/2015/02/20/openedx-and-lti-pedagogical-scripts-and-sso/"/>
    <content type="html">&lt;div class="toc"&gt;&lt;ul style="margin-top: 0em;margin-bottom: 1.5em;"&gt;&lt;li&gt;&lt;a href="#0"&gt;Installation&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#1"&gt;Cohorts and single-sign on&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#2"&gt;LTI to the rescue&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#3"&gt;Early experiment, group-based Etherpad selection&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#4"&gt;Confluence integration and templates&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#5"&gt;See it in action&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;&lt;/ul&gt;&lt;/div&gt;&lt;p&gt;&lt;strong&gt;I came up with a neat way to embed external web tools into EdX courses using an LTI interstitial, which also allows for rich pedagogical scripting and group formation. Illustrated with &lt;a href="https://www.youtube.com/watch?v=-OY9UPT4dK8"&gt;a brief screencast&lt;/a&gt;.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Last year, I &lt;a href="/blog/2014/10/03/a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content/"&gt;experimented&lt;/a&gt; with &lt;a href="/blog/2014/10/03/supporting-idea-convergence-through-pedagogical-scripts/"&gt;implementing pedagogical scripts&lt;/a&gt; spanning different collaborative Web 2.0 tools, like Etherpad and Confluence wiki, in two hybrid university courses.&lt;/p&gt;

&lt;p&gt;&lt;a href="/blog/2014/10/03/supporting-idea-convergence-through-pedagogical-scripts/"&gt;&lt;img src="/blog/images/2014-10-03-a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content_-_whole-04.png" alt=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Coordinating collaboration and group formation in a class of 85 students was already challenging, but currently we are planning a MOOC built on the same principles. Before we begin to design the pedagogical scripts and the flow of the course, a key question was how much we could possible implement technically given the MOOC platform. Happily, it turns out that it will be easier and more seamless than I had feared.&lt;/p&gt;

&lt;p&gt;We chose to go with EdX (University of Toronto has agreements with both EdX and Coursera), which is based on an open-source platform. Open source can mean many different things, there are platforms that are mainly developed privately in a company, and released to the public as periodic "code dumps", and where installation is very complex, and the product is unlikely to be used by anyone else. I didn't know much about the EdX code, but the fact that the platform is already used by &lt;a href="https://github.com/edx/edx-platform/wiki/Sites-powered-by-Open-edX"&gt;a very impressive number of other installations&lt;/a&gt;, such as China's &lt;a href="https://www.xuetangx.com/"&gt;XuetangX&lt;/a&gt; and &lt;a href="https://www.france-universite-numerique-mooc.fr/cours/"&gt;France's Université Numérique&lt;/a&gt; was already a very positive sign. &lt;/p&gt;

&lt;!-- more --&gt;

&lt;h2 id="0"&gt;Installation&lt;/h2&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-02-10-openedx-and-lti-pedagogical-scripts-and-sso_-_half-01.png" alt="" class="img-right"&gt;
The OpenEdX code is available on GitHub, and seems to be a very active community -- more than 28k commits by 199 contributors, &amp;gt;100 pull requests, 621 branches... But the most important thing is that it is very easy to install. I remember when we were building P2PU's Lernanta platform, also on Django, and it was quite difficult to install the code locally (having the right Python version, pulling in all the right libraries, etc), even to make some tiny changes. OpenEdX, which is much more complex, with a number of different services that need to work together, offers several pre-built packages. I chose to use Vagrant to run it locally, and once you've installed Vagrant and VirtualBox, simply using the commands listed &lt;a href="https://github.com/edx/configuration/wiki/edx-Full-stack--installation-using-Vagrant-Virtualbox"&gt;here&lt;/a&gt; was enough to get a fully functioning local OpenEdX instance:&lt;/p&gt;

&lt;pre&gt;&lt;code class="bash"&gt;mkdir fullstack
cd fullstack
curl -L https://raw.githubusercontent.com/edx/configuration/master/vagrant/release/fullstack/Vagrantfile &amp;gt; Vagrantfile
vagrant plugin install vagrant-hostsupdater
vagrant up
&lt;/code&gt;&lt;/pre&gt;

&lt;h2 id="1"&gt;Cohorts and single-sign on&lt;/h2&gt;

&lt;p&gt;Now that I had a functional version of OpenEdX running locally, it was time to look at our challenges. We were interested in two separate but connected pieces of functionality. &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;enabling single-sign on with external tools, so that students can use an external wiki, for example, without having to create a new account and remember a new password&lt;/li&gt;
&lt;li&gt;ways of "scripting" the collaboration process, particularly grouping students into smaller groups, and sending students to various resources based on time, group affiliation, and what they had already done&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-02-10-openedx-and-lti-pedagogical-scripts-and-sso_-_half-02.png" alt="" class="img-right"&gt;
For the second challenge, EdX already has a "cohort-feature", which let's you divide students into groups, and make certain parts of the discussion forum specific to each group. However, we were interested in much more detailed scripts than that. In our existing course, we used APIs to automatically insert templates into each group's space, take content from different groups and aggregate together, conditionally send students to different tools depending on their previous input, etc. Since EdX does not have an API that gives an instructor live access to the discussion forum, we couldn't for example create 300 cohort-groups and automatically populate each forum thread with a question post, then extract their answers, and do something with them. We would thus need to use external tools for this aspect of the course, bringing us back to the single-sign on issue.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-02-10-openedx-and-lti-pedagogical-scripts-and-sso_-_half-04.png" alt="" class="img-right"&gt;
Inspired in particular by the recent EdX MOOC &lt;a href="https://www.edx.org/course/data-analytics-learning-utarlingtonx-link5-10x"&gt;Data, Analytics and Learning&lt;/a&gt;, and their inclusion of external social and collaborative tools, I began investigating our options. The simplest way of including external tools, beyond simply linking to them, is to use an IFrame. However, a tool included in an IFrame receives no information from EdX about which student is visiting, and can send no information back.&lt;/p&gt;

&lt;p&gt;An intriguing option that the &lt;a href="https://linkresearchlab.org/dalmooc/"&gt;DALMOOC&lt;/a&gt; used, is to let the EdX platform function as an OpenID provider, and having a "Login with EdX" button, similar to the many sites that feature "Login with Google/Facebook/Twitter, etc". However, I could not find any information about how to enable this on EdX.org, and it would also mean that the students still would have to do a round-trip back to EdX to allow the login, making it less seamless.&lt;/p&gt;

&lt;h2 id="2"&gt;LTI to the rescue&lt;/h2&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-02-10-openedx-and-lti-pedagogical-scripts-and-sso_-_half-03.png" alt="" class="img-right"&gt;
Although I've been working with hybrid and online learning for many years, I've always been drawn to existing Web 2.0 tools, using APIs and RSS feeds to tie things together. Because of this, I don't have much experience with traditional LMSes, and things like SCORM and LTI -- in fact, they always sounded quite cumbersome and complex. However, it turns out that LTI (&lt;a href="http://www.imsglobal.org/toolsinteroperability2.cfm"&gt;Learning Tools Interoperability&lt;/a&gt;) is a very simple protocol; it's basically a simple web request (either as an IFrame, or as a link opening a separate window) with three extra pieces: &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Authentication&lt;/strong&gt;&lt;br&gt;
The authentication ensures that the request really comes from the OpenEdX platform (in this case), and that I cannot easily fake a request with someone else's user ID.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Student info&lt;/strong&gt;&lt;br&gt;
The student info in the case of EdX is not any actual student info, but a persistent hashed identifier. This means I can always recognize when the same student accesses the LTI object, and I can also later connect learning analytics from the LTI object with EdX data (there is a hash table connecting the hash IDs with the actual student IDs in the data dumps).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Callback URL&lt;/strong&gt;&lt;br&gt;
 The callback URL is for graded LTI objects, to deliver a simple numeric grade back to the EdX platform.&lt;br&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I first spent some time researching an LTI integration directly with &lt;a href="http://confluence.atlassian.com"&gt;Confluence&lt;/a&gt;, but I could not find one that was currently maintained. Limiting ourselves only to web apps that offer LTI integration would also severly restrict our choices, and even if the tools did offer LTI integration, that would only give us something like single-sign on, not any opportunity to implement pedagogical scripts. This led to idea of using an interstitial LTI script. This script would receive the request from OpenEdX, keep track of the relevant information about the student ID, and forward to various web services, providing automatic logon or deep-linking where appropriate/possible.&lt;/p&gt;

&lt;h2 id="3"&gt;Early experiment, group-based Etherpad selection&lt;/h2&gt;

&lt;p&gt;I found &lt;a href="https://github.com/instructure/ims-lti.git"&gt;an LTI library for Ruby&lt;/a&gt;, which had a great little &lt;a href="https://github.com/instructure/lti_tool_provider_example"&gt;minimal example application&lt;/a&gt;. The application is simple -- it prompts the user for a grade, and then sends that grade back to the platform -- but the fact that I could get that to work with OpenEdX using &lt;a href="http://edx-partner-course-staff.readthedocs.org/en/latest/exercises_tools/lti_component.html"&gt;the provided instructions&lt;/a&gt;, gave me the confidence to keep experimenting. I decided to use &lt;a href="http://redis.io/"&gt;Redis&lt;/a&gt; as a backing database, since it is very easy to experiment with, and would also provide easy scaling if needed in the future. &lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-02-10-openedx-and-lti-pedagogical-scripts-and-sso_-_half-05.png" alt="" class="img-right"&gt;Since we don't know anything about the student arriving, other than the hashed persistent id (there might be ways of getting more information from the system in the future), my first step was a simple form where students would specify their nickname, and choose a group (this could in the future be a long list of groups, or even some complex group formation algorithm).&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-02-10-openedx-and-lti-pedagogical-scripts-and-sso_-_whole-01.png" alt=""&gt;
The next access point then simply forwards to an Etherpad, however it chooses the Etherpad URL based on the group already select (and stored in Redis). This way, we could have three hundred groups, and funnel students into 300 different Etherpads without changing anything in the EdX course setup. We could even have an on-demand feature which puts entrants into a "waiting room" until a pre-determined number of students appear, upon which a new Etherpad room is generated, students are forwarded to that room, and a new waiting room is created. &lt;/p&gt;

&lt;h2 id="4"&gt;Confluence integration and templates&lt;/h2&gt;

&lt;p&gt;Since Etherpad does not use any authentication, it was a much simpler thing to integrate than Confluence. Luckily, Confluence has a rich API, which I've already used extensively in previous courses, however I did not know if it would be possible to automatically log students in. In the end, it turns out that by accessing a particular URL, you can offer as arguments the username, password, and the page to be redirected to:&lt;/p&gt;

&lt;pre&gt;&lt;code class="ruby"&gt;url = "https://www.WIKIURL.org/dologin.action?os_username=#{wiki_username}" + 
"&amp;amp;os_password=#{wiki_pwd}&amp;amp;os_destination=/display/IN/#{page}".
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;When a user tries to access a wiki endpoint, I first check in Redis whether the user already has a wiki password. If not, I randomly generate a wiki username and password, and create that user with the appropriate permissions, using the Confluence API. I then redirect the user's browser to the URL above, with the username and password (which they never need to see). The result is that they are seemlessly logged in, and redirected to the page I want.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-02-10-openedx-and-lti-pedagogical-scripts-and-sso_-_half-07.png" alt="" class="img-right"&gt;
However, having thousands of users simultaneously edit an unstructured wiki can be chaotic. Our idea is to make it easy for students to supply information, and for a smaller group of engaged users to engage with editing and enriching that information. As an example, I created a simple form which then populates a wiki page. In the current example, upon visiting this LTI object, the student automatically gets a wiki account (if he/she doesn't already have one), the page is created, and the student is automatically logged in and sent to the newly created page. In an actual course, the page could be generated by information submitted by numerous students, and accessing the page wouldn't have to follow immediately after filling out the template.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2015-02-10-openedx-and-lti-pedagogical-scripts-and-sso_-_whole-02.png" alt=""&gt;&lt;/p&gt;

&lt;h2 id="5"&gt;See it in action&lt;/h2&gt;

&lt;iframe width="420" height="315" src="https://www.youtube.com/embed/-OY9UPT4dK8" frameborder="0" allowfullscreen&gt;&lt;/iframe&gt;

&lt;p&gt;Although this example is quite simple, I find it very promising. I am very happy that due to the OpenEdX platform being so easy to install, and the Ruby LTI library comes with an example app, I was able to get this running in an afternoon. Now comes the hard work -- designing meaningful scripts and interaction patterns for students with very different reasons to participate, and levels of engagement, but at least we know that we can probably implement what we want to. We also need to think about scaling, and whether our local services can handle the amount of users we are expecting.&lt;/p&gt;

&lt;p&gt;I'd love to hear about others experimenting with pushing on the interactive and collaborative features in MOOCs and large courses. It would also be interesting to think about how to make LTI tools even easier to integrate -- I'd love to see a catalogue of external LTI tools, ranging from simple individual widgets, to groupware and collaboration tools, which could all be easily slotted into an EdX, Instructure, or even Blackboard course.&lt;/p&gt;

&lt;p&gt;All the example code &lt;a href="https://github.com/houshuang/inquirymooc"&gt;is on GitHub&lt;/a&gt;.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <id>tag:reganmian.net,2015-01-23:/blog/blog/2014/10/17/random-german-stuff-that-matters---great-content-auf-deutsch/</id>
    <title type="html">Random German Stuff that Matters - great content auf Deutsch</title>
    <published>2015-01-23T16:42:31Z</published>
    <updated>2015-01-23T16:42:31Z</updated>
    <link rel="alternate" href="https://reganmian.net/blog/2014/10/17/random-german-stuff-that-matters---great-content-auf-deutsch/"/>
    <content type="html">&lt;div class="toc"&gt;&lt;ul style="margin-top: 0em;margin-bottom: 1.5em;"&gt;&lt;li&gt;&lt;a href="#0"&gt;Why introduce German content in English?&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#1"&gt;Podcasts and radio&lt;/a&gt;&lt;/li&gt;&lt;ul style="margin-top: 0em;margin-bottom: 1.5em;"&gt;&lt;li&gt;&lt;a href="#2"&gt;Radio programs I like&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#3"&gt;Podcasts I like&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;&lt;li&gt;&lt;a href="#4"&gt;TV, movies and series&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#5"&gt;Literature&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;&lt;/div&gt;&lt;h1 id="0"&gt;Why introduce German content in English?&lt;/h1&gt;

&lt;p&gt;One issue that I have not heard much discussion about, is the challenge of finding interesting material in foreign languages. In your own culture and language, you have gotten to know authors through school, through popular media, recommendations from friends, etc. But when you begin exploring a foreign language, it's often like starting from scratch again. I remember walking around in the libraries in Italy, having no idea where to start, or what kinds of authors I might enjoy. &lt;/p&gt;

&lt;p&gt;This is made worse by the fact that I am typically much slower at reading in a foreign language, and particularly in "skimming". This means that even though I can really enjoy sitting down with a nice novel, and slowly working through it, actually keeping up with a bunch of blogs, skimming the newspaper every day, or even doing a Google-search and flitting from one page to the next, becomes much more difficult. &lt;/p&gt;

&lt;p&gt;Some things are easier these days - it's typically easier to get access to media over the Internet, whether it be movies, podcasts, or ebooks (although this differs a lot between different languages, it is still difficult finding Indonesian movies or TV-shows online, and even Norwegian e-books are hard to come by online), but navigating is still difficult. When I began learning Russian again, I was amazed at the number of very high-quality English blogs that talk about Russian modern literature (places like &lt;a href="http://lizoksbooks.blogspot.ca/"&gt;Lizok's bookshelf&lt;/a&gt;, &lt;a href="http://xixvek.wordpress.com/"&gt;XIX век&lt;/a&gt; and &lt;a href="http://languagehat.com/"&gt;Languagehat&lt;/a&gt;), which really inspired me to continue working on my Russian, to be able to read the books they so enthusiastically discussed.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-17-random-german-stuff-that-matters---great-content-auf-deutsch_-_whole-08.png" alt="" class="img-right"&gt;
These last two years have been a bit of a renaissance for my own usage of German. It's a language I learnt in school in Norway, I read &lt;a href="/blog/2008/07/19/celebrating-first-book-read-in-a-foreign-language/"&gt;my first novel&lt;/a&gt; in it when I was 16, and have visited the country a number of times. However, for the past 10-15 years, I did not have many opportunities to use it, and did not seek out German novels, podcasts or anything like that. However, spending a very enjoyable month visiting a lab at the University of Munich gave me a chance to reawaken my language skills and interest. I then did a short visit last winter, and spent about three months in Berlin this summer. &lt;/p&gt;

&lt;p&gt;In addition to the visits, what got me in to reading was ironically Karl May. Not well known in North America, he is the all-time best-selling German author from the 19th Century, writing a large amount of fantastical tales of travel and adventure around the world, from places he had never visited. I had heard about him, but never read anything, and decided that I would try it out for fun. I found &lt;a href="http://www.karl-may-gesellschaft.de/kmg/primlit/roman/herzen/index.htm"&gt;"Deutsche Herzen, deutsche Helden"&lt;/a&gt; as a free download and began. It was certainly a fantastic tale, taking us from Istanbul, to Egypt, through the Wild West and ending up in Siberia. &lt;/p&gt;

&lt;p&gt;Long like a Bollywood-movie, it felt like each part could have been it's own substantial book. Full of handsome and brave Germans, it was certainly not always politically correct (although I have read far worse things in English from the colonial times), but quite enjoyable, and at the end, I realized I had read almost 3000 pages (these things happen on a Kindle), and that my German reading speed had improved measurably as a result.&lt;/p&gt;

&lt;p&gt;Since then, I have found a great number of enjoyable German novels and podcasts, and I wanted to share some of these with you. If you speak German as a foreign language, but were not quite sure where to start, perhaps some of this will be useful. And if you don't speak German, perhaps you'll be inspired to learn. Of course, this is just a sampling of what I've randomly come across, not an exhaustive survey, and it unapologetically is stuff that I like and enjoy, no attempt at objective critique.&lt;/p&gt;

&lt;h1 id="1"&gt;Podcasts and radio&lt;/h1&gt;

&lt;p&gt;I love going for walks and listening to podcasts and audio books. I've also found that audio supports language learning and maintenance in a different way than reading; in the same way that reading a lot can make you a better writer, I find myself speaking more fluently and confidently after listening to audio in a language, even if I haven't actually spoken it for a long time. A good example was when I met someone speaking German in a Moscow airport a few years ago. At that time, I hadn't been speaking  German for years, but just finished listening to &lt;a href="https://en.wikipedia.org/wiki/Momo_(novel)"&gt;Momo&lt;/a&gt;, and was surprised by my fluency.&lt;/p&gt;

&lt;p&gt;Germany has a very vibrant private podcasting scene, and also very high-quality TV and radio stations. Most of these have so-called Mediathek, where you can easily find programs and links to feeds (like &lt;a href="http://www.ardmediathek.de/tv"&gt;ARD's Mediathek&lt;/a&gt;).&lt;/p&gt;

&lt;h2 id="2"&gt;Radio programs I like&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;DRadio Wissen Einhundert&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="http://dradiowissen.de/einhundert"&gt;&lt;img src="/blog/images/2014-10-17-random-german-stuff-that-matters---great-content-auf-deutsch_-_whole-01.png" alt=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="http://dradiowissen.de/einhundert"&gt;Einhundert&lt;/a&gt;&lt;/strong&gt; is my favorite German radio program. Every week, they choose a theme, which can be very broad ("Starting over", "Hitting the road"), but also more specific (a program about when half a million Russian soldiers left Eastern Germany). In each case, they find people with interesting stories that somehow relate to the theme (often in very different ways).&lt;/p&gt;

&lt;p&gt;It's a bit similar to &lt;a href="http://www.thisamericanlife.org/"&gt;This American Life&lt;/a&gt; or &lt;a href="http://www.radiolab.org/"&gt;RadioLab&lt;/a&gt; (the latter of which I love as well), but they have managed to find their own voice. It is difficult to put the finger on what is so brilliant about this program, but somehow the combination of interesting stories and good editing makes for great listening. I often tell my wife about stories I head in this podcast, two of the more memorable ones:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="http://dradiowissen.de/beitrag/wie-aus-papa-abui-wurde"&gt;A son learns Arabic as an adult, and begins to understand his father in a whole new way&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="http://dradiowissen.de/beitrag/iran-die-stellvertreter-reise-in-meine-heimat"&gt;A girl has never been to her homeland, Iran, and is not allowed to travel there. Her boyfriend goes in her place, as her eyes and ears&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;News&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-17-random-german-stuff-that-matters---great-content-auf-deutsch_-_half-01.png" alt="" class="img-right"&gt;You can watch or download the video of the main evening news in Germany, &lt;strong&gt;&lt;a href="http://www.tagesschau.de/"&gt;Die Tagesschau&lt;/a&gt;&lt;/strong&gt;, and sometimes I watched the news on my iPad while doing dishes. However, my preference is to download the audio and listen while walking. Surprisingly, I find I never really miss the visuals. I appreciate the focus on domestic and international politics, rather than "a man was hit by a bus yesterday".&lt;/p&gt;

&lt;p&gt;Listening regularly, you get a good understanding of the main political issues in Germany and the EU, and how they are developing. It's been particularly interesting following the Ukraine-conflict through German eyes. I also often listen to &lt;strong&gt;&lt;a href="http://www.tagesschau.de/sendung/tagesthemen/"&gt;Tagesthemen&lt;/a&gt;&lt;/strong&gt;, which tends to go a bit more in depth. There is also &lt;strong&gt;&lt;a href="http://www.tagesschau.de/bab/"&gt;Bericht aus Berlin&lt;/a&gt;&lt;/strong&gt;, which is a weekly political magazine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Other academic programs&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="http://dradiowissen.de/beitrag/musik-digitale-reproduzierbarkeit-beat-mp3"&gt;&lt;img src="/blog/images/2014-10-17-random-german-stuff-that-matters---great-content-auf-deutsch_-_half-02.png" alt="" class="img-right"&gt;&lt;/a&gt;
&lt;strong&gt;&lt;a href="http://dradiowissen.de/hoersaal"&gt;DRadio Hörsaal&lt;/a&gt;&lt;/strong&gt; is an interesting format, where they first interview a researcher about their research, followed by a public lecture made by the same person. The topics vary widely, and I remember hearing &lt;a href="http://dradiowissen.de/beitrag/musik-digitale-reproduzierbarkeit-beat-mp3"&gt;a fascinating talk&lt;/a&gt; about how we now play classical music faster, than when it was authored.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="http://www.swr.de/swr2/wissen/"&gt;SWR Wissen&lt;/a&gt;&lt;/strong&gt; is a series of traditional radio documentaries, where a journalist tells a story, intermixed with some interviews. Topics vary very widely, but somehow the good story-tellers are able to make many topics that don't sound very enticing into interesting experiences. (Looking at the site now, I see that they also make &lt;a href="http://www.swr.de/swr2/wissen/epub/-/id=661224/nid=661224/did=9072786/1w7y7sx/index.html"&gt;transcripts of all programs available in epub-format&lt;/a&gt; - quite innovative.)&lt;/p&gt;

&lt;h2 id="3"&gt;Podcasts I like&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Chaosradio&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="http://www.ccc.de/en/"&gt;&lt;img src="/blog/images/2014-10-17-random-german-stuff-that-matters---great-content-auf-deutsch_-_whole-02.png" alt="" class="img-right"&gt;&lt;/a&gt;
&lt;a href="http://www.ccc.de/en/"&gt;Chaos Computer Club&lt;/a&gt; is a historically famous network of hacker-clubs in Germany, and their Berlin chapter produce a monthly radio program/podcast called &lt;strong&gt;&lt;a href="http://chaosradio.ccc.de/chaosradio.html"&gt;Chaosradio&lt;/a&gt;&lt;/strong&gt;. The goal is to explain technical concepts to a more lay audience, so they usually have members of the club moderated by a radio moderator. The shows last something like three hours, and it can be quite enjoyable to hear the friendly bantering of the participants. For example, I quite enjoyed the show last winter &lt;a href="http://chaosradio.ccc.de/cr197.html"&gt;about the NSA surveillance scandal&lt;/a&gt; and how mass-digital surveillance works.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CRE: Technik, Kultur, Gesellschaft&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="http://cre.fm"&gt;&lt;img src="/blog/images/2014-10-17-random-german-stuff-that-matters---great-content-auf-deutsch_-_half-03.png" alt="" class="img-right"&gt;&lt;/a&gt;
Through Chaosradio, I came across &lt;strong&gt;&lt;a href="http://cre.fm/"&gt;CRE&lt;/a&gt;&lt;/strong&gt;, which began as an extension to Chaosradio, but now has a life of its own. It's run by &lt;a href="http://metaebene.me/timpritlove/"&gt;Tim Pritlove&lt;/a&gt;, who is an incredibly prolific German podcaster, and the format is basically Tim sitting down with someone who is a specialist on a certain topic, and just talking. Sometimes the conversations will go over 3-4 hours, and somehow listening to two very intelligent people in a nice conversation can be quite interesting and calming. Some of my favorite episodes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="http://cre.fm/cre205-wikidata"&gt;Wikidata&lt;/a&gt;, the new semantic Wikimedia project, and how we think about ontologies and hierarchies to organize knowledge&lt;/li&gt;
&lt;li&gt;
&lt;a href="http://cre.fm/cre200-stadtplanung"&gt;Stadtplanung&lt;/a&gt;, city planning from Alexandria and Babylon to Russian suburbs, Syria and Berlin&lt;/li&gt;
&lt;li&gt;
&lt;a href="http://cre.fm/cre184-bundeswehr"&gt;Bundeswehr&lt;/a&gt;, the role of the German army in international conflicts&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="http://fokus-europa.de/"&gt;&lt;img src="/blog/images/2014-10-17-random-german-stuff-that-matters---great-content-auf-deutsch_-_half-05.png" alt="" class="img-right"&gt;&lt;/a&gt;
Tim also hosts a few other noteworthy podcasts, including &lt;strong&gt;&lt;a href="http://logbuch-netzpolitik.de/"&gt;Logbuch:Netzpolitik&lt;/a&gt;&lt;/strong&gt;, which reminds me of many English-language podcasts, where two people sit around every week and chat about "what's new" related to Internet and technology politics. A more structured interview-podcast is &lt;strong&gt;&lt;a href="http://fokus-europa.de/"&gt;Fokus Europa&lt;/a&gt;&lt;/strong&gt; which attempts to explain the structure of modern Europe, focused mostly on the European Union. I found the first two episodes &lt;a href="http://fokus-europa.de/podcast/fe001-geschichte-der-europaeischen-einigung/"&gt;Geschichte der Europäischen Einigung&lt;/a&gt; and &lt;a href="http://fokus-europa.de/podcast/fe002-deutschland-und-europa/"&gt;Deutschland und Europa&lt;/a&gt; quite interesting.&lt;/p&gt;

&lt;p&gt;&lt;a href="http://www.staatsbuergerkunde-podcast.de"&gt;&lt;img src="/blog/images/2014-10-17-random-german-stuff-that-matters---great-content-auf-deutsch_-_half-04.png" alt="" class="img-right"&gt;&lt;/a&gt;
&lt;strong&gt;&lt;a href="http://www.staatsbuergerkunde-podcast.de"&gt;Staatsbürgerkunde&lt;/a&gt;&lt;/strong&gt; is a great example of oral history. Martin Lutz interviews his parents Martin and Christine Lutz about their life in the DDR, mixing general knowledge and history with their own personal recollections and anecdotes. I particularly enjoyed the episode on &lt;a href="http://www.staatsbuergerkunde-podcast.de/2013/11/23/sbk030-weisensee/"&gt;Weissensee&lt;/a&gt;, the drama series I mention below.&lt;/p&gt;

&lt;h1 id="4"&gt;TV, movies and series&lt;/h1&gt;

&lt;p&gt;This will unfortunately be a very short section, and I would love further suggestions. I enjoy watching German movies when I come across them, but I have no good online source to find them (Netflix only has a few). I have also been looking for good German drama series, and would love to find something like the great Danish crime dramas (like &lt;a href="https://en.wikipedia.org/wiki/The_Bridge_(Danish/Swedish_TV_series)"&gt;The Bridge&lt;/a&gt;). However, I'll mention one great drama series and a serie of documentaries.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://de.wikipedia.org/wiki/Terra_X:_Deutschland_von_oben"&gt;&lt;img src="/blog/images/2014-10-17-random-german-stuff-that-matters---great-content-auf-deutsch_-_whole-03.png" alt=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://de.wikipedia.org/wiki/Terra_X:_Deutschland_von_oben"&gt;Deutschland von oben&lt;/a&gt;&lt;/strong&gt; is a series of documentaries using a combination of satellite and aerial footage to show different aspects of the geography of Germany. There are several themes, "City", "Land" and "Water", and we go from looking at the majestic golden eagles in Berchtesgaden, to how current street patterns are the result of the structure of the initial Roman settlements. It gave me a much better understanding of the diversity of German landscapes and regions, and although it can get tiring, the dramatic music and voices go well with the sweeping vistas.
&lt;a href="http://www.daserste.de/unterhaltung/serie/weissensee/index.html"&gt;&lt;img src="/blog/images/2014-10-17-random-german-stuff-that-matters---great-content-auf-deutsch_-_half-06.png" alt="" class="img-right"&gt;&lt;/a&gt;
All the episodes are available in streaming HD, and there is also a feature-length cinema version.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="http://www.daserste.de/unterhaltung/serie/weissensee/index.html"&gt;Weissensee&lt;/a&gt;&lt;/strong&gt; is a mini-series in two seasons about life in East-Berlin during the DDR-times. The series focus on two families, a Stasi-family and a rebellious artist family, whose lives interact in various ways. There are definitively aspects of Romeo and Juliette, and cliches from other mini-series, but it's very well done, and apparently very period-authentic. I very much enjoyed both seasons, and also &lt;a href="http://www.staatsbuergerkunde-podcast.de/2013/11/23/sbk030-weisensee/"&gt;the discussion about the series&lt;/a&gt; on the Staatsbürgerkunde podcast mentioned above.&lt;/p&gt;

&lt;h1 id="5"&gt;Literature&lt;/h1&gt;

&lt;p&gt;Reading fiction is one of my main ways of approaching a foreign language, and German is a very welcoming host. It has become much easier now that we have access to e-books, and literature portals. I remember when I began learning German and Italian 15 years ago, I would go into a library and be completely lost as to where to start. These days, we have great sites like &lt;a href="http://www.krimi-couch.de/"&gt;Krimi-couch&lt;/a&gt;, or even &lt;a href="http://www.zeit.de/kultur/literatur/"&gt;the literature section in Die Zeit&lt;/a&gt; for more high-brow literature.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-17-random-german-stuff-that-matters---great-content-auf-deutsch_-_half-07.png" alt="" class="img-right"&gt;
In general I quite enjoy crime novels, and I found a German crime writer tradition that felt quite similar to the &lt;a href="http://www.scandinaviancrimefiction.com/"&gt;Scandinavian crime&lt;/a&gt; that I grew up with. It tends to be tightly bound to a specific region or town, and to spend just as much time on the personal lives of the investigators, as well as the perpetrators and victims of the crime, sometimes also touching upon societal trends and challenges. This is not the hard-boiled crime noir from smokey San Francisco PI-offices, but rather a great way to examine different tensions in society, different occupational groups, life in small and big cities, etc.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-17-random-german-stuff-that-matters---great-content-auf-deutsch_-_half-08.png" alt="" class="img-right"&gt;
One of the things I have really missed in Canada, is regionally based crime stories. I think Toronto has many stories waiting to be told, and not just Toronto, but specific areas - Scarborough, the Junction, white-collar crime on Bay street, corruption at Queens Park, etc. In Germany, the concept "regiokrimi" (regional crime) even has it's &lt;a href="http://www.regiokrimi.de/"&gt;own website&lt;/a&gt;. Apparently there are even small towns that commission crime novels set in their area for the town's anniversary. This has led to some critical voices, complaining that literary aspects are overlooked in an attempt to stuff as many local references as possible, but I've come across a number of authors that I really enjoyed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="http://www.klauspeterwolf.de/"&gt;Klaus-Peter Wolf&lt;/a&gt;&lt;/strong&gt; has a fascinating history on &lt;a href="http://de.wikipedia.org/wiki/Klaus-Peter_Wolf"&gt;Wikipedia&lt;/a&gt;. Apparently he almost works like an investigative journalist, often diving deep into the various thematics that he then fictionalizes, going so far as to set up a fake mailorder-marriage company before he wrote a book about the trade with women, and living with a youth gang, before writing a book about youth criminality and violence. However, my favorite is his series of novels from Ost-Friesland, an area I had never heard about before, and now feel that I know well enough that I could navigate the area without a map. &lt;/p&gt;

&lt;p&gt;Another set of great novels set in the area come from &lt;strong&gt;&lt;a href="http://www.sven-koch.com/"&gt;Sven Koch&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-17-random-german-stuff-that-matters---great-content-auf-deutsch_-_whole-07.png" alt="" class="img-right"&gt;
&lt;strong&gt;&lt;a href="https://de.wikipedia.org/wiki/Jacques_Berndorf"&gt;Jaques Berndorf&lt;/a&gt;&lt;/strong&gt; writes about another area of Germany that I was not familiar with, the Eifel. In this case, the protagonist is an investigative journalist, rather than an investigator, and somehow he always gets tangled into complex issues, often involving the state and the secret services. A cranky old man living alone in the countryside, there always seems to be an intriguing female showing up just in time to join him in his investigations. &lt;/p&gt;

&lt;p&gt;Reminds me of some earlier (80's-90's) Swedish crime, which also talked about investigative journalists, the Swedish secret-services and government, and their connections to foreign powers, etc. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="http://www.jenksaborowski.com/"&gt;Jenk Saborowski&lt;/a&gt;&lt;/strong&gt; writes about the fictional European Federal Police, and agent Solveigh Lang, who travels Europe to catch criminals. Intriguing fiction, which reflects a modern Europe connected with high-speed rail and the absence of borders, but still divided by different bureaucracies, languages and cultures.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-17-random-german-stuff-that-matters---great-content-auf-deutsch_-_whole-06.png" alt="" class="img-right"&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="http://www.tomliehr.de/"&gt;Tom Liehr&lt;/a&gt;&lt;/strong&gt; has been compared to Nick Hornby, and it's a comparison that makes a lot of sense. He writes about young men, fixated on music, DJing, being a radio host, or going on crappy chartertours. Great writing, and very rich characters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;There are so many more authors I could list, here are some very quick introductions&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Nele Neuhaus&lt;/strong&gt;: Decent crime set in "Vordertaunus", rural area near Frankfurt. Somehow always ends up with a situation in which a large number of people could have had motive and possibility to kill the victim.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sebastian Fitzek&lt;/strong&gt;: "Psychodrama"-bestseller. Really enjoyed &lt;em&gt;Passagier 23&lt;/em&gt; and &lt;em&gt;Amokspiel&lt;/em&gt;, but others of his novels are too "psychological", with the I-voice seemingly going crazy, unable to believe anything he sees/remembers, etc. 
&lt;img src="/blog/images/2014-10-17-random-german-stuff-that-matters---great-content-auf-deutsch_-_whole-05.png" alt="" class="img-right"&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Petra Durst-Benning&lt;/strong&gt;: Interesting and enjoyable historical fiction from various periods in Germany, usually centered around strong women. &lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Birgit Schlieper&lt;/strong&gt;: Very interesting and well written novels about difficult issues facing young people, with a strong psychological aspect. Powerful and thoughtful.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Katharina Peters&lt;/strong&gt;: Fascinating crime novels from the other German coast, Rügen, often involving the Nazi-history of the area.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Martin Suter&lt;/strong&gt;: A series of crime novels, and some free-standing novels, usually dealing with the Swiss upperclass. Charmingly written. Best free-standing novel is &lt;em&gt;Der Koch&lt;/em&gt;, amazing writing (although it looses some of the steam after the beginning).
&lt;img src="/blog/images/2014-10-17-random-german-stuff-that-matters---great-content-auf-deutsch_-_whole-04.png" alt="" class="img-right"&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sigrid Ramge&lt;/strong&gt;: OK crime stories set in Baden-Württemberg, but felt like very "regio-krimi", almost like she wants to give you a detailed tour of the area, rather than just building it into the story.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Manuela Kuck&lt;/strong&gt;: Well-written crime from the Wolfsburg-region. &lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sebastian Lehmann&lt;/strong&gt; with &lt;em&gt;Genau mein Beutelschema&lt;/em&gt;, hilarious and well-written ironic take on the different areas and subcultures in Berlin. Perfect reading while staying in Neukölln last summer.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Gute Vergnügen!&lt;/p&gt;
</content>
  </entry>
  <entry>
    <id>tag:reganmian.net,2014-10-14:/blog/blog/2014/10/14/starting-data-analysiswrangling-with-r-things-i-wish-id-been-told/</id>
    <title type="html">Starting data analysis/wrangling with R: Things I wish I'd been told</title>
    <published>2014-10-14T16:17:31Z</published>
    <updated>2014-10-14T16:17:31Z</updated>
    <link rel="alternate" href="https://reganmian.net/blog/2014/10/14/starting-data-analysiswrangling-with-r-things-i-wish-id-been-told/"/>
    <content type="html">&lt;div class="toc"&gt;&lt;ul style="margin-top: 0em;margin-bottom: 1.5em;"&gt;&lt;li&gt;&lt;a href="#0"&gt;Use RStudio&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#1"&gt;Use Knitr&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#2"&gt;Learn from others&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#3"&gt;Separate cleaning and organizing from analysis&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#4"&gt;Use version control&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#5"&gt;Learn the DSLs in R&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#6"&gt;DRY - Don't repeat yourself&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#7"&gt;Keeping track of types in R&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#8"&gt;Asking questions - providing reproducible examples&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="#9"&gt;Packages to use&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;&lt;/ul&gt;&lt;/div&gt;&lt;p&gt;&lt;a href="http://www.r-project.org/"&gt;R&lt;/a&gt; is a very powerful open source environment for data analysis, statistics and graphing, with thousands of &lt;a href="http://cran.r-project.org/web/packages/"&gt;packages&lt;/a&gt; available. After my previous blog post about &lt;a href="/blog/2013/10/02/likert-graphs-in-r-embedding-metadata-for-easier-plotting"&gt;likert-scales and metadata in R&lt;/a&gt;, a few of my colleagues mentioned that they were learning R through &lt;a href="https://class.coursera.org/compdata-003/class"&gt;a Coursera course on data analysis&lt;/a&gt;. I have been working quite intensively with R for the last half year, and thought I'd try to document and share a few tricks, and things I wish I'd have known when I started out.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2013-10-09-starting-data-analysiswrangling-with-r-things-i-wish-id-been-told_-_whole-05.png" alt=""&gt;&lt;/p&gt;

&lt;p&gt;I don't pretend to be a statistics whiz – I don't have a strong background in math, and much of my training in statistics was of the social science &lt;em&gt;"click here, then here in SPSS"&lt;/em&gt; kind, using flowcharts to divine which tests to run, given the kinds of variables you wanted to compare. I'm eager to learn more, but the fact is that running complex statistical functions in R is typically quite easy. The difficult part is acquiring data, cleaning it up, combining different data sources, and preparing it for analysis (they say 90% of a data scientist's job is data wrangling). &lt;em&gt;Of course, knowing which tests to run, and how to analyze the results is also a challenge, but that is general statistical knowledge that applies to all statistics packages.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;So here are some of my suggestions and "lessons learnt", in no particular order. Some will find the code samples scary, others will find the suggestion to use for-loops far too basic, but hopefully you will find something useful here.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;h2 id="0"&gt;Use RStudio&lt;/h2&gt;

&lt;p&gt;&lt;img src="/blog/images/2013-10-09-starting-data-analysiswrangling-with-r-things-i-wish-id-been-told-_-whole-01.png" alt=""&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="http://www.rstudio.com/"&gt;RStudio&lt;/a&gt; is an great open source &lt;em&gt;integrated development environment&lt;/em&gt; for R. It is free and open-source, and integrates&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a project manager&lt;/li&gt;
&lt;li&gt;a text editor with syntax highlighting and tab-completion&lt;/li&gt;
&lt;li&gt;package management&lt;/li&gt;
&lt;li&gt;previewing plots&lt;/li&gt;
&lt;li&gt;previewing tables (see columns and rows in your dataset)&lt;/li&gt;
&lt;li&gt;integrated help (click F1 on any function)&lt;/li&gt;
&lt;li&gt;jump to source (click F2 on any function)&lt;/li&gt;
&lt;li&gt;and &lt;a href="http://yihui.name/knitr/"&gt;Knitr&lt;/a&gt; (write reports in &lt;a href="http://daringfireball.net/projects/markdown/"&gt;Markdown&lt;/a&gt;, see below)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There's even a version of RStudio &lt;a href="http://www.rstudio.com/ide/docs/server/getting_started"&gt;that runs in the browser&lt;/a&gt; – we're currently using it to coordinate data analysis among a geographically dispersed team on many gigabytes of data. Keeping it on a central server, and letting people run analyses directly on the data is much more convenient and secure than having everyone store tens of gigabytes of data on their personal computers.&lt;/p&gt;

&lt;h2 id="1"&gt;Use Knitr&lt;/h2&gt;

&lt;p&gt;&lt;a href="http://en.wikipedia.org/wiki/Literate_programming"&gt;Literate programming&lt;/a&gt; is the idea of mixing executable code with documentation in the same document. Knitr brings this functionality to R, and it's integrated beautifully with RStudio. By default, all the text you write in a Knitr document is interpreted as &lt;a href="http://daringfireball.net/projects/markdown/"&gt;Markdown&lt;/a&gt; (a light-weight markup language, which I'm also using to author this blog). Press &lt;code&gt;Alt+Cmd I&lt;/code&gt; to insert an R code block. You can run the code either by pressing &lt;code&gt;Cmd+Enter&lt;/code&gt; on a single line, or &lt;code&gt;Alt+Cmd C&lt;/code&gt; to execute an entire code block. When you're done, you can press "Knit HTML", which executes the whole document and produces a report.&lt;/p&gt;

&lt;p&gt;Here's an example from some recent work analyzing a questionnaire, we're introducing the graph, and then adding the code that will produce the graph (&lt;a href="/blog/2013/10/02/likert-graphs-in-r-embedding-metadata-for-easier-plotting"&gt;see my previous blog post on likert-graphs&lt;/a&gt;):&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2013-10-09-starting-data-analysiswrangling-with-r-things-i-wish-id-been-told_-_whole-04.png" alt=""&gt;&lt;/p&gt;

&lt;p&gt;Running &lt;code&gt;Knit HTML&lt;/code&gt; combines the text, formatted nicely, the code used to generate the graph, and the actual graph:&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2013-10-09-starting-data-analysiswrangling-with-r-things-i-wish-id-been-told_-_whole-03.png" alt=""&gt;&lt;/p&gt;

&lt;p&gt;For the final report, you might choose to hide all the code segments with this invocation:&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;library(knitr)
opts_chunk$set(echo=F, warning=F,message=F,results="asis", prompt=F, error=F)
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Not only is this a great way of writing reports (you can also export to PDF, or even write in LaTeX instead of Markdown), but it's a very nice way of organizing your code. I now do all my development in this mode, even for scripts where I'm not interested in the final report. I like the ease of documenting, the clear visual separation of the different code blocks, and the ease of pressing &lt;code&gt;Alt+Cmd C&lt;/code&gt; to execute a single code block.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;I recently gave a small demo to a research group, showing them RStudio, using Markdown to write knitr reports, and Shiny to make interactive webpages:&lt;/em&gt;&lt;/p&gt;

&lt;iframe width="459" height="344" src="https://www.youtube.com/embed/LIIC8cViC54?feature=oembed" frameborder="0" allowfullscreen&gt;&lt;/iframe&gt;

&lt;h2 id="2"&gt;Learn from others&lt;/h2&gt;

&lt;p&gt;There are lot's of R textbooks and documentation out there. Two other great sources of ideas are &lt;a href="http://rpubs.com/"&gt;RPubs&lt;/a&gt; and &lt;a href="http://www.r-bloggers.com/"&gt;r-bloggers&lt;/a&gt;. Using knitr to write reports in RStudio, as described above, you have the option to upload your report to RPubs. You can also see all the reports that others have written. These are great to learn from, since they typically contain both all the code, and the output and graphs, tying it all together into an analytical narrative. Unfortunately, there is no good way of searching the site, so you will find people's homework nestled between great expository writing, but it's still well worth a visit. You can also visit &lt;a href="http://yihui.name/knitr/demo/showcase/"&gt;the knitr notebook showcase&lt;/a&gt; to see some select examples.&lt;/p&gt;

&lt;p&gt;&lt;a href="http://www.r-bloggers.com/"&gt;R-bloggers&lt;/a&gt; is an aggregator for blog-posts about R and statistics, and it's a great way to discover new packages or R features, and see how other people attack various data analysis challenges, often using publicly available datasets.&lt;/p&gt;

&lt;p&gt;&lt;a href="http://rpubs.com/"&gt;&lt;img src="/blog/images/2014-10-14-starting-data-analysiswrangling-with-r-things-i-wish-id-been-told_-_whole-01.png" alt=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2 id="3"&gt;Separate cleaning and organizing from analysis&lt;/h2&gt;

&lt;p&gt;Although R is a fully-fledged (although a bit crufty) programming language with object-orientation and functions, the code that users typically write is very different from an R package. Users usually write very imperative code, &lt;em&gt;"load this file, then transform the second column, then add the third column, then graph it"&lt;/em&gt;. However, acquiring some good habits of organizing the code and working with the data, might save you a lot of time in the long run. &lt;em&gt;(Also see the section on DRY below)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I usually separate data importing and cleaning from the analysis. My goal is to leave the raw data completely unchanged, and do all the transformation in code, which can be rerun at any time. While I'm writing the scripts, I'm often jumping around, selectively executing individual lines or code blocks, running commands to inspect the data in the REPL (read-evaluate-print-loop, where each command is executed as soon as you type enter, in the picture above it's the pane to the right), etc. But I try to make sure that when I finish up, the script is runnable by itself.&lt;/p&gt;

&lt;p&gt;Knitr helps impose this - when you choose &lt;code&gt;Knit HTML&lt;/code&gt;, it begins with a clean slate. When you are working in RStudio, you might have objects lying around from calculations you did a while ago (with code that you've already changed), but if the "knitting" is successful, you know that the current code is valid and produces exactly what you see.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2013-10-09-starting-data-analysiswrangling-with-r-things-i-wish-id-been-told_-_whole-06.png" alt=""&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example of preparing data&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In one example, I had gotten spreadsheets from several students who had helped me enter data from a large survey. I could have opened these in Excel, and copied and pasted the information into one sheet, but instead I left the files as they were, and read them into R. I tag the columns from each spreadsheet with provenance, if I want to run any quick tests, or even more formal &lt;a href="http://en.wikipedia.org/wiki/Inter-rater_reliability"&gt;interrater reliability tests&lt;/a&gt;, and merge them into one table. In some spreadsheets, there was an extra empty column, so I remove that programmatically (rather than editing the Excel spreadsheet). First I load &lt;code&gt;xlsx&lt;/code&gt; to read the Excel spreadsheets, and &lt;code&gt;plyr&lt;/code&gt; to rename fields later.&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;library(xlsx)
library(plyr)
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Then I read in and join the spreadsheet files:&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;ana &amp;lt;- read.xlsx(file="ana.xlsx", 1,stringsAsFactors=FALSE)
ana$by &amp;lt;- "Ana"

chad &amp;lt;- read.xlsx(file="chad.xlsx", 1,stringsAsFactors=FALSE)
chad[["NA."]] &amp;lt;- NULL
chad$by &amp;lt;- "Chad"

dd &amp;lt;- read.xlsx(file="DD.xls", 1,stringsAsFactors=FALSE)
dd$by &amp;lt;- "DD"

db &amp;lt;- rbind(ana, chad, dd)
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;I noticed that some of the spreadsheets had a bunch of empty rows at the bottom. This might not be the most elegant way, but I used this function to remove all the spreadsheets with all empty values. The way this works is that is.na(x) produces a list of "TRUE TRUE FALSE FALSE TRUE" depending on which of the columns has an NA for that particular row, and sum() adds up the TRUE's as 1, and FALSE's as 0.&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;db2 &amp;lt;- db[apply(db,1, function(x) {
  sum(is.na(x)) &amp;lt; 43}),]
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Then I move on to turn all the various versions of "empty cell" into NA:&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;db &amp;lt;- as.data.frame(lapply(db2, function(x){
  x &amp;lt;- replace(x, x %in% c("n", "N", ""), NA)
  x &amp;lt;- as.factor(x)}))
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Some of the columns are demographic values, so I'll change these from numeric to categorical variables with the actual values:&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;db$gender &amp;lt;- revalue(db$X1, c("1"="Male", "2"="Female", "3"="Trans", "4"="Other", "5"=NA))
db$major &amp;lt;- revalue(db$X2, c("1"="History", "2"="Religion", "3"="Other"))
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;And the rest of the questions are likert-style questions with the same categories, so I'll both rename them and order them in one swoop:&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;likertcat &amp;lt;- c("1"="Not at all", "2"="To a small extent", "3"="To some extent",
  "4"="To a moderate extent", "5"="To a large extent")

for(e in names(db[,9:44])) {
  db[[e]] &amp;lt;- revalue(db[[e]], likertcat)
  db[[e]] &amp;lt;- ordered(db[[e]], levels= c("Not at all","To a small extent",
    "To some extent","To a moderate extent","To a large extent"))
}
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;I also add categories and groupings using my own addition, which &lt;a href="/blog/2013/10/02/likert-graphs-in-r-embedding-metadata-for-easier-plotting"&gt;you can read more about&lt;/a&gt;, and finally I save the new table as an RData file:&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;save(db, file="db.RData")
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;I should always be able to re-run this script, and re-generate the exactly same RData file, but typically the only time I'll re-run the script is if I get additional data (let's say two other students send me their data entry files).&lt;/p&gt;

&lt;p&gt;I then begin the data analysis script with&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;load(file="db.RData")
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;And we have a nicely prepared table that we can begin exploring.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2013-10-09-starting-data-analysiswrangling-with-r-things-i-wish-id-been-told_-_whole-07.png" alt=""&gt;&lt;/p&gt;

&lt;h2 id="4"&gt;Use version control&lt;/h2&gt;

&lt;p&gt;&lt;img src="/blog/images/2013-10-09-starting-data-analysiswrangling-with-r-things-i-wish-id-been-told_-_half-02.png" alt="version control" class="img-left"&gt;
RStudio comes with support for &lt;a href="http://git-scm.com/"&gt;git&lt;/a&gt; baked in, and it's a great practice to use it. When you create a new project, check the box for &lt;em&gt;"Create a git repository for this project"&lt;/em&gt;, and you get all this functionality for free. Once you have a functioning version of your script, commit it to Git, and when you later make changes, you can easily track those changes back in history, restore earlier versions, etc. This is great for people collaborating, but it's actually a great idea to start doing it even on solo-projects, and can save you a lot of grief in the future. RStudio has a good &lt;a href="http://www.rstudio.com/ide/docs/version_control/overview"&gt;writeup&lt;/a&gt; of this functionality.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2013-10-09-starting-data-analysiswrangling-with-r-things-i-wish-id-been-told_-_whole-08.png" alt=""&gt;&lt;/p&gt;

&lt;p&gt;Of course, if you are able to do so, it would be great if you could share your scripts (and data) on &lt;a href="https://github.com/"&gt;GitHub&lt;/a&gt; or another public repository. Even in cases where you can't share your raw data (because of ethics, etc), your code could be useful to others. (I have shared some of the code we used &lt;a href="http://www.ocw.utoronto.ca/demographic-reports/"&gt;to analyze Coursera MOOC data&lt;/a&gt; in &lt;a href="https://github.com/houshuang/coursera-scripts"&gt;a GitHub repository&lt;/a&gt;, but I could probably go through my folders and find more snippets of code to share).&lt;/p&gt;

&lt;h2 id="5"&gt;Learn the DSLs in R&lt;/h2&gt;

&lt;p&gt;DSL stands for domain-specific language, and the key to being efficient in R (and one of the reasons beginners might feel that the learning curve is very steep), is to realize that R is actually composed of several different embedded languages. Most end-users of R actually don't need to know very much about the R programming language, mostly people write scripts in a very imperative way "first do this, then do this, then do that". There are probably data analysts that have used R for years, and never written a function or a loop statement.&lt;/p&gt;

&lt;p&gt;However, if you want to plot graphics with &lt;a href="http://ggplot2.org/"&gt;ggplot&lt;/a&gt;, you will have to learn an entirely new way of thinking. It is based on the &lt;a href="http://vita.had.co.nz/papers/layered-grammar.html"&gt;Grammar of Graphics&lt;/a&gt;, and makes it possible to intuitively build up exactly the kind of graph you want, based on layers, mapping data to geometries such as bar-charts, etc. However, if you don't understand the underlying logic, even making the simplest plot will seem like a strange incantation.&lt;/p&gt;

&lt;p&gt;Similarly, for data wrangling, there are several options. The &lt;a href="http://www.stat.ubc.ca/%7Ejenny/STAT545A/block04_dataAggregation.html"&gt;plyr&lt;/a&gt; family of functions is very powerful, for example here is how you would take a data.frame with per-country data, and calculate per-continent summary statistics:&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;ddply(gDat, ~ continent, summarize,
      minLifeExp = min(lifeExp), maxLifeExp = max(lifeExp),
      medGdpPercap = median(gdpPercap))
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;An alternative is to use &lt;a href="http://blog.yhathq.com/posts/fast-summary-statistics-with-data-dot-table.html"&gt;data.table&lt;/a&gt;. Data.table is an alternative to the built-in data.frame, which is much more performant for large datasets, and also offers powerful indexing, transformation/grouping, and merging/joining. It has a very dense syntax, which is powerful once you understand the logic. Here is an example of calculating a number of statistics on a football dataset:&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;dt.statsPlayer &amp;lt;- dt.boxscore.subset[,list(minutes=sum(minutes),
                                           mean=mean(pp48),
                                           min=min(pp48),
                                           lower=quantile(pp48, .25, na.rm=TRUE),
                                           middle=quantile(pp48, .50, na.rm=TRUE),
                                           upper=quantile(pp48, .75, na.rm=TRUE),
                                           max=max(pp48)),
                                     by='player']
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;(Note also the use of white-space, R often let's you space expression out over several lines, which can help a lot with readability).&lt;/p&gt;

&lt;p&gt;The key to getting proficient in R is to choose one main way of doing things, for example ggplot2 for all graphics (and not worrying about learning grid and lattice plotting), and either data.table or plyr, and focusing on understanding their underlying logic, and using them frequently.&lt;/p&gt;

&lt;h2 id="6"&gt;DRY - Don't repeat yourself&lt;/h2&gt;

&lt;p&gt;This used to be a mantra when programming Ruby, but is often overlooked in R code. Since many people think of doing analysis in R as simply writing a number of instructions to be executed linearly, we forget that the language let's us be much more efficient. I did say above that you don't need to understand the R programming language in depth to use R profitably, however, a few features are very useful.&lt;/p&gt;

&lt;p&gt;The advantage of avoiding repetition is that your code becomes easier to read (because your intention stands out), and it becomes much easier to reuse code, and to update. Let's say you are plotting five or six similar plots, and in each case you have a big blob of ggplot2 code, with layers, themes, etc. Now your boss tells you to change the theme of all the plots. Rather than jumping around and editing a bunch of lines, perhaps introducing new errors, you could try to abstract out the plotting code to a function, which you'd only need to edit once.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Turning often used code into functions&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="/blog/images/2013-10-09-starting-data-analysiswrangling-with-r-things-i-wish-id-been-told_-_half-03-orig.png"&gt;&lt;img src="/blog/images/2013-10-09-starting-data-analysiswrangling-with-r-things-i-wish-id-been-told_-_half-03.png" alt="" class="img-right"&gt;&lt;/a&gt;
Here's an example of a simple plot I used frequently in an exploratory report I wrote about MOOC data. I first clean out the NAs from the dataframe, and then display it as a colored bar chart, where each bar corresponds to a region, and the colors correspond to age groups. The first line defines the plot, and the two subsequent lines formats it (flipping the axis, removing some clutter, adding a legend title). &lt;em&gt;(Click the image to the right to see the full graph.)&lt;/em&gt;&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;dbgen &amp;lt;- db[!is.na(db$gender), ]
ggplot(data=dbgen, aes(region, fill=age)) + geom_bar(position="fill") +
  coord_flip() + scale_y_continuous(labels = percent) + ylab("") + xlab("") +
  guides(fill=guide_legend(title="Age"))
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Since I needed a few graphs, the initial impulse would be to copy it and modify a few parameters. But how can we instead make this into a generic function? Turns out that in general we'd want to remove NAs from both the grouping and the fill column (the region column just didn't happen to have any NAs, so we didn't need to check that above), and we also need to switch to using aes_string because of how ggplot implements its "DSL", otherwise it's quite simple.&lt;/p&gt;

&lt;p&gt;We define a function &lt;code&gt;my_sideplot&lt;/code&gt;, which takes the full dataframe, the grouping variable, the fill variable, and an optional title for the fill legend, and generates a plot similar to the one above.&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;my_sideplot = function(db, group, fill, fill_title="") {
  db = db[!is.na(db[[fill]]) &amp;amp; !is.na(db[[group]]),]
  ggplot(data=db, aes_string(group, fill=fill)) +
    geom_bar(position="fill") + coord_flip() +
    scale_y_continuous(labels = percent) + ylab("") + xlab("") +
    guides(fill=guide_legend(title=fill_title))
}
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;The advantage is that I can now very easily experiment with different versions. Perhaps using the age variable for grouping, and the regions for fill is better?&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;my_sideplot(db, "age", "region", "Regions")
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;This makes my code much easier to read, since I see right away that this is doing exactly the same as the code above (otherwise you need to look very closely at the ggplot code to see what is different), and if I have a knitr report with 10-15 of these graphs, and I want them all to use &lt;code&gt;theme_bw()&lt;/code&gt; – a black-and-white theme – I can simply modify the function, and rerun the graph.&lt;/p&gt;

&lt;p&gt;Functions can of course be much more advanced than this, with branches and sanity checks, etc. However, as a first approximation, just taking chunks of code that are repeated with very few changes, identifying the changes and making them into parameters, and turning the code chunk into a function, can already be very useful.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Using lists and iteration&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The second approach to cleaning up code is to use lists and for-loops. For example, we might want to generate plots of six different categorical variables grouped by region. Or perhaps we want them grouped by region, age, and English-skill? Now we are talking about a total of 20 graphs. We could copy and paste the one-line &lt;code&gt;my_sideplot&lt;/code&gt; invocation above, which is already a big improvement, but we can do much better.&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;group_vars = c("region", "age", "english.skill")
fill_vars = c("usercat", "education.level", "degree.program")
for (groupv in group_vars) {
  for (fillv in fill_vars) {
    my_sideplot(db, groupv, fillv)
  }
}
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;This code will iterate through the &lt;code&gt;group_vars&lt;/code&gt;, and for each &lt;code&gt;group_var&lt;/code&gt;, it will iterate through the &lt;code&gt;fill_vars&lt;/code&gt;. For each &lt;code&gt;fill_var&lt;/code&gt;, it will render an appropriate plot.&lt;/p&gt;

&lt;p&gt;You can also iterate through for example the columns in a data.frame. Let's say we want to turn all the character columns into factors:&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;for (e in names(db)) {
  if (is.character(db[[e]])) { db[[e]] &amp;lt;- as.factor(db[[e]])}
}
&lt;/code&gt;&lt;/pre&gt;

&lt;h2 id="7"&gt;Keeping track of types in R&lt;/h2&gt;

&lt;p&gt;Again breaking against my statement that you don't need to know much of the R language, but &lt;a href="http://www.statmethods.net/input/datatypes.html"&gt;the various data types&lt;/a&gt; is one of the things that has tripped me up the most. This is because certain operations can return different data types than you were expecting.&lt;/p&gt;

&lt;p&gt;There are many R tutorials, and I am not going to give a full introduction here. The most basic datatype is the vector, and one of the peculiarities of R is that even single numbers or strings, are vectors of size one.&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;identical(c("hi"), "hi") # =&amp;gt; TRUE
identical(c(1), 1) # =&amp;gt; TRUE
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;A dataframe is a collection of vectors. Sometimes when you do an operation on a data.frame, you might get a vector back, or a matrix. You can usually cast one data type to another with &lt;code&gt;as.*&lt;/code&gt;, for example &lt;code&gt;as.data.frame()&lt;/code&gt;&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;&amp;gt; as.data.frame(list("A" = c(1,2,3), "B" = c(2,3,4)))
  A B
1 1 2
2 2 3
3 3 4
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;You can check the class of an object with &lt;code&gt;class()&lt;/code&gt;,&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;&amp;gt; class(db)
[1] "data.frame"
&amp;gt; class(db$age)
[1] "ordered" "factor"
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;If you run into weird bugs, make sure that every step of the pipeline returns the data type that the next function is expecting (or explicitly cast). Also remember that certain functions work differently depending on different data types. For example, &lt;code&gt;length()&lt;/code&gt; on a list or vector returns the number of elements. &lt;code&gt;length()&lt;/code&gt; on a data.frame returns the number of columns, if you want the number of rows, you need &lt;code&gt;nrow()&lt;/code&gt;.&lt;/p&gt;

&lt;h2 id="8"&gt;Asking questions - providing reproducible examples&lt;/h2&gt;

&lt;p&gt;There are many great resources for getting help with R, including all the books I showed pictures of at the beginning of this post. An advantage of R being text-oriented is that it is easy to paste exactly the code that is misbehaving, and for others to help out. &lt;a href="http://stackoverflow.com/questions/tagged/r"&gt;StackOverflow&lt;/a&gt; has a very active R community, however you have a much higher chance of getting help if you make your problem reproducible.&lt;/p&gt;

&lt;p&gt;Sometimes you can use the built-in data sets in R, which you can list with the command &lt;code&gt;data()&lt;/code&gt;. For example, with the command &lt;code&gt;data(swiss)&lt;/code&gt;, I load a dataset with data on fertility and education in Swiss cantons. I can use this to illustrate a problem, and anyone else using R will have the same dataset, and can perfectly reproduce my problem, or check that their solution works correctly.&lt;/p&gt;

&lt;p&gt;If none of the built-in datasets work, you can create a minimal example dataset that illustrates your problem. Let's say we had a problem with the &lt;code&gt;a&lt;/code&gt; data.frame we constructed above. The command &lt;code&gt;dput(a)&lt;/code&gt; gives us&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;structure(list(A = c(1, 2, 3), B = c(2, 3, 4)), .Names = c("A",
"B"), row.names = c(NA, -3L), class = "data.frame")
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;We can now construct our example:&lt;/p&gt;

&lt;pre&gt;&lt;code class="r"&gt;a = structure(list(A = c(1, 2, 3), B = c(2, 3, 4)), .Names = c("A",
"B"), row.names = c(NA, -3L), class = "data.frame")

ggplot(a, aes(x=A, y=B)) + geom_point()
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;...and anyone can paste this into their R interpreter, and see the same output that you get. &lt;a href="http://stackoverflow.com/questions/5963269/how-to-make-a-great-r-reproducible-example"&gt;Longer explanation&lt;/a&gt; and here's &lt;a href="http://stackoverflow.com/questions/1299871/how-to-join-data-frames-in-r-inner-outer-left-right"&gt;a small example&lt;/a&gt;.&lt;/p&gt;

&lt;h2 id="9"&gt;Packages to use&lt;/h2&gt;

&lt;p&gt;One of the strengths of R is the incredible selection of packages available, but this can also be quite disorienting to a beginner. Here are just a few packages that are often useful:&lt;/p&gt;

&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th style="text-align: left"&gt;Package&lt;/th&gt;
&lt;th style="text-align: left"&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="text-align: left"&gt;&lt;a href="http://plyr.had.co.nz/"&gt;plyr&lt;/a&gt;&lt;/td&gt;
&lt;td style="text-align: left"&gt;Plyr applies the "split-apply-combine" approach to R data. You might want to calculate the average height of players, grouped by nationality and year of birth - a one-liner in Plyr. Hadley Wickham is currently developing &lt;a href="https://github.com/hadley/dplyr"&gt;dplyr&lt;/a&gt;, a faster and more powerful version&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: left"&gt;&lt;a href="http://shiny.rstudio.com/"&gt;Shiny&lt;/a&gt;&lt;/td&gt;
&lt;td style="text-align: left"&gt;Shiny let's you make interactive graphical websites with R (similar to the way we created the function above, by substituting the parts of the code that vary with variables, and letting the end-user choose the variables). Also see the &lt;a href="https://www.youtube.com/watch?v=LIIC8cViC54"&gt;video&lt;/a&gt; above&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: left"&gt;&lt;a href="http://cran.r-project.org/package=data.table"&gt;data.table&lt;/a&gt;&lt;/td&gt;
&lt;td style="text-align: left"&gt;As mentioned above, a much faster version of data.frame, very well suite to larger datasets, with powerful split+apply, merge, etc. functionality offering an alternative to plyr&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: left"&gt;&lt;a href="http://jason.bryer.org/likert/"&gt;likert&lt;/a&gt;&lt;/td&gt;
&lt;td style="text-align: left"&gt;If you ever deal with likert-style questionnaire data, this package makes it very easily to visualize the results. See also &lt;a href="/blog/2013/10/02/likert-graphs-in-r-embedding-metadata-for-easier-plotting"&gt;my addition&lt;/a&gt; to this package&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Hope that was useful. Once you want more, you could do worse than checking out Hadley Wickham's &lt;a href="http://adv-r.had.co.nz/"&gt;Advanced R&lt;/a&gt; book. There's also an incredible amount of good textbooks, examples, blog posts, etc.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <id>tag:reganmian.net,2014-10-03:/blog/blog/2014/10/03/a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content/</id>
    <title type="html">A pedagogical script for idea convergence through tagging Etherpad content</title>
    <published>2014-10-03T21:17:26Z</published>
    <updated>2014-10-03T21:17:26Z</updated>
    <link rel="alternate" href="https://reganmian.net/blog/2014/10/03/a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content/"/>
    <content type="html">&lt;p&gt;&lt;strong&gt;I describe a script aimed at supporting student idea convergence through tagging Etherpad content, and discuss how it went when I implemented it in a class&lt;/strong&gt;&lt;/p&gt;

&lt;h1&gt;Background&lt;/h1&gt;

&lt;p&gt;In &lt;strong&gt;&lt;a href="/blog/2014/10/03/supporting-idea-convergence-through-pedagogical-scripts/"&gt;an earlier blog post&lt;/a&gt;&lt;/strong&gt; I introduced the idea of pedagogical scripting, as well as implementing scripts in computer code. I discussed my desire to make ideas more "moveable", and support deeper work on ideas, and talked about the idea of using tags to support this. Finally, I introduced the tool &lt;a href="/blog/2012/06/13/tag-extract-a-tool-to-automatically-restructure-textoutline-using-tags/"&gt;tag-extract&lt;/a&gt;, which I developed to work on &lt;a href="/wiki/draft_literature_review_open_courses/"&gt;a literature review&lt;/a&gt;.&lt;/p&gt;

&lt;h1&gt;Context&lt;/h1&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-03-a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content_-_half-02.png" alt="" class="img-right"&gt;I am currently teaching a course on Knowledge and Communication for Development (&lt;a href="http://idsb10.pbworks.com/w/page/37652952/FrontPage"&gt;earlier open source syllabi&lt;/a&gt;) at the University of Toronto at Scarborough. The course applies theoretical constructs from development studies to understanding the role of technology, the internet, and knowledge in international development processes, and as implemented in specific development projects.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-03-a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content_-_half-01.png" alt="" class="img-left"&gt;I began the course with two fairly tech-centric classes, because I believe that having an intuition about how the Internet works, is important for subsequent discussions. I've also realized in previous years that even the "digital generation" often has very little understanding of what happens when you for example send a Facebook message from one computer to another.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;So I spent the first week focused on the physical infrastructure of the Internet, the history and development, protocols and structure. We used traceroute to explore connections between local computers, from Scarborough to downtown Toronto, or to servers in the US or China. We used telnet to simulate a connection to a web server, we handwrote a very simple webpage, and I had students in groups "draw" the process of sending an e-mail from Toronto to China (inspired by &lt;a href="http://openmatt.org/2010/10/04/draw-how-the-internet-works/"&gt;John Britton's idea&lt;/a&gt; in his &lt;a href="http://archive.p2pu.org/webcraft/web-200-anatomy-request.html"&gt;P2PU course&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-03-a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content_-_half-03.png" alt="" class="img-right"&gt;The second week, &lt;strong&gt;in which I implemented this script&lt;/strong&gt;, was focused more on the software side, particularly Web 2.0. I wanted the students to understand the history of the Web, the concept of the social web, how websites have gone from simple containers of content to web apps and places of interaction, the interaction between commercial providers and user-generated content, etc. &lt;em&gt;I also thought this would be a great opportunity to try out some interactive production and organization of knowledge in practice.&lt;/em&gt;&lt;/p&gt;

&lt;h1&gt;The script&lt;/h1&gt;

&lt;p&gt;My basic idea was to have the students work in small groups on adding ideas to group Etherpads. My key question was: &lt;strong&gt;What are the distinguishing features of Web 2.0&lt;/strong&gt;. I would have a few different "external stimuli", after which they would add more ideas to their own pad. At the end of the class, we would use tagging to create a &lt;a href="http://en.wikipedia.org/wiki/Folksonomy"&gt;folksonomy&lt;/a&gt;, and then collaboratively reduce the folksonomy to a few key categories. Based on the categories, my script would extract the tagged pieces of text from everyone's pad, and create one new pad per category. My idea was then to assign the students to edit these pads and turn them into wiki articles that students could later refer to when studying, or working on the final project.&lt;/p&gt;

&lt;p&gt;Inspired by Pierre Dillenbourg's work on orchestration graphs, which he talked about at &lt;a href="https://www.youtube.com/watch?v=UIRDxHAnAlg"&gt;ICLS in Sydney&lt;/a&gt; and &lt;a href="http://new.livestream.com/hgselive/events/3105335/videos/55325177"&gt;LASI&lt;/a&gt; in Cambridge, MA, I attempted to visually represent the script.&lt;/p&gt;

&lt;p&gt;Script (click for a bigger version)&lt;/p&gt;

&lt;p&gt;&lt;a href="/blog/images/2014-10-03-a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content_-_whole-04-orig.png"&gt;&lt;img src="/blog/images/2014-10-03-a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content_-_whole-04.png" alt=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The graph is organized across three vertical levels, representing work done by students individually, in small teams (in our case groups of 2-3 students), and done in the whole class. The x-axis represents time, and the stars represents processing by an external script which shuffles information between pads. Below, I'll discuss the script more in detail.&lt;/p&gt;

&lt;h2&gt;Phase 1&lt;/h2&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-03-a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content_-_whole-03.png" alt="" class="img-left"&gt;The students had been assigned two readings as homework, and came to the class prepared. In addition, they mostly had pre-existing ideas about Web 2.0, the social web, etc. Before doing any lecturing at all, I asked them to gather into groups of 2-3 people (based on where they were sitting), and counted off. A script had already generated a number of Etherpads with identical prompts, that were all linked from the week's wiki page.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-03-a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content_-_whole-05.png" alt=""&gt;&lt;/p&gt;

&lt;h2&gt;Phase 2&lt;/h2&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-03-a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content_-_whole-02.png" alt=""&gt;&lt;/p&gt;

&lt;p&gt;The idea was to have a few rounds of providing the students with stimulus that would lead to the generation of more ideas, which could be recorded in the pads. After the students had finished their initial brainstorming, I led the class in a discussion, talked through the articles, and showed two Michael Wesch videos: &lt;a href="http://www.youtube.com/watch?v=NLlGopyXT_g"&gt;The Machine is Us/ing Us&lt;/a&gt; and &lt;a href="https://www.youtube.com/watch?v=-4CV05HyAbM"&gt;Information R/evolution&lt;/a&gt;. I then gave the students some time to write down ideas and thoughts.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-03-a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content_-_half-04.png" alt="" class="img-right"&gt;I wanted to give the students a sense of how the web has developed, and I've found the &lt;a href="https://archive.org/web/"&gt;Internet Archive Wayback Machine&lt;/a&gt; a great tool in the past. This let's the students explore the history of websites they are familiar with. We looked at a few examples together, and then I asked each group to choose a few websites they were familiar with (and encouraged them to also look at websites in other languages, since we have a very multilingual student body), and specifically look at features and functionality that has changed, not just design. We then discussed some of these examples together, and the students added more ideas to their pads.&lt;/p&gt;

&lt;p&gt;The final stimulus was looking in-depth at a specific Web 2.0 site. I had prepared a list of URLs, and the script automatically inserted a random URL at the bottom of each group's pad. The URLs included Wikipedia, Flickr, a newspaper, the university LMS, etc. I always enjoy walking around the room when students are working on these kinds of tasks, because the quick conversations we have while they are in the middle of exploring a topic, and I am able to point them in the right direction, seem really valuable. (And something I have not really captured in the orchestration graph).&lt;/p&gt;

&lt;h2&gt;Phase 3&lt;/h2&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-03-a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content_-_whole-01.png" alt=""&gt;&lt;/p&gt;

&lt;h2&gt;Individual tagging&lt;/h2&gt;

&lt;p&gt;At this point, each student group had a pad filled with notes and ideas about Web 2.0 features and functionality, the shift from Web 1.0 to Web 2.0, etc. They had gradually added to this pad during the class, and I had provided rich external input, and opportunities to discuss in small groups, and in the whole class. If the class had ended here, I would have already considered it a success.&lt;/p&gt;

&lt;p&gt;However, one of my design goals for this iteration was to make content generated in class more accessible and useful to students going forwards. Nobody would probably wade through a bunch of unorganized Etherpads when they were later working on the final project, or preparing for the exam. I wanted to see if my &lt;a href="/blog/2012/06/13/tag-extract-a-tool-to-automatically-restructure-textoutline-using-tags/"&gt;tag-extract&lt;/a&gt; approach, which had been so useful to me in organizing and structuring my own notes and ideas when doing a literature review, could help the students collaboratively reorganize all of their notes into a coherent whole.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-03-a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content_-_half-05.png" alt="" class="img-right"&gt;I first asked the students to add tags to the ideas in their pads. They were all familiar with tagging from Twitter, and from the readings and videos in class, and I connected the exercise to this, without telling them about the next step yet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Below you can see an example of the evolution of the pad from one student group, ending with the students adding tags to their existing ideas.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-09-22-tagging-and-convergence-for-small-groups-collaboration-with-multiple-etherpad-01-static.png" alt="" class="post-thumb" onclick='this.src="/blog/images/2014-09-22-tagging-and-convergence-for-small-groups-collaboration-with-multiple-etherpad-01.gif"'&gt;&lt;/p&gt;

&lt;h2&gt;Collectively organizing the folksonomy&lt;/h2&gt;

&lt;p&gt;After the students had added tags to their pads, the script extracted all the tags, and created a new Etherpad listing them all. Including misspellings and variants, the students had come up with almost 100 tags. The pad with all the tags offered a simple way of categorizing tags, if you have a list of tags that are similar and should be grouped together, such as this list from our class:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;advertisments&lt;/li&gt;
&lt;li&gt;viralmarketing&lt;/li&gt;
&lt;li&gt;consumerism&lt;/li&gt;
&lt;li&gt;marketization&lt;/li&gt;
&lt;li&gt;onlinerevenue&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We just needed to combine them together into one category:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;advertisments: viralmarketing, consumerism, marketization, onlinerevenue
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Because we were using Etherpad for this as well, the whole class could participate, and in just a few minutes, we had organized the tags into a much smaller number of categories. In the example above, this was quite easy, but for others, the ordering might have been a bit arbitrary. Since we did not have access to the actual text tagged at this point, it was also sometimes hard to know what people had meant.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;See the animation below of this process.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-09-22-tagging-and-convergence-for-small-groups-collaboration-with-multiple-etherpad-02-static.png" alt="" class="post-thumb" onclick='this.src="/blog/images/2014-09-22-tagging-and-convergence-for-small-groups-collaboration-with-multiple-etherpad-02.gif"'&gt;&lt;/p&gt;

&lt;h2&gt;Putting it all together&lt;/h2&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-03-a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content_-_half-06.png" alt="" class="img-right"&gt;The final move was to run a script which would go through all the pads, and extract all the text tagged into appropriate category pads (ie. above, text tagged either #advertisments or #viralmarketing or #consumerism would all end up in a pad called advertisments). There was a bit of an anti-climax as that script failed to run at the end of the class, because of a silly error. That was OK, because when I fixed the script after class, and ran it, I realized that the output was often not as useful as I had naively hoped.&lt;/p&gt;

&lt;p&gt;On the right, you see an example of a pad, with the title listing all the tags in the category, and the specific category also appended after each text snippet. (I could also have appended the name of the group pad that the snippet came from, but in this case, that would just make it more cluttered.)&lt;/p&gt;

&lt;h1&gt;Evaluation&lt;/h1&gt;

&lt;p&gt;The first part of the script worked great, but was not that different from the way I often teach. Pushing different URLs to each pad worked very well (and students loved having something suddenly show up beneath the text they were just editing), I've used this particular script again in a later class.&lt;/p&gt;

&lt;p&gt;However, the grand idea of using tags to reorganize and structure the text still needs more work. I think the prompt was the first problem, "What distinguishes Web 2.0" was not perhaps the most productive or clear question I could have asked. I could have made it more problem-oriented, or debate (pro/con), etc. I also never explained to the students the purpose of what we were doing - they are used to taking notes in the Etherpad during group work, but I never told them that it would be tagged, cut up into small pieces, and shared.&lt;/p&gt;

&lt;h2&gt;Now do this, then do that&lt;/h2&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-03-a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content_-_half-07.png" alt="" class="img-right"&gt;Once we tagged, I didn't tell them that we would gather all the tags and organize them, and while we were organizing the tags, I didn't tell them what we were doing it for. I think teachers often do this unconsciously – lead students through activities that they have clearly thought out, without telling the students where they are going, or why they are moving in a certain way. Sometimes this is a necessary part of the script, whereas at other times, students need conscious training in certain epistemic moves, to be able to carry them out effectively.&lt;/p&gt;

&lt;p&gt;I was very impressed by Kate Bielaczyc's presentation at ICLS 2010, where she used a sport metaphore to distinguish between practice and playing a match. To get good at a sport, you don't just play matches, the coach leads you into targetted practice, setting up contrived situations to challenge you. Similarly, she had her students practice the act of knowledge building with contrived situations, before they actually got into authentic knowledge building.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Hypothetical game-configurations are used to reflect on the knowledge building moves made possible by a particular configuration of knowledge objects. The configurations consist of “snapshots” of hypothetical student work in Knowledge Forum, meant to capture game play at a fixed point in time in order to engage the community in asking: given this configuration, what types of knowledge building moves would best contribute to advancing our knowledge? Some configurations focus on single moves, such as presenting a possible initial explanation generated in response to the problem the students are working on.&lt;/p&gt;

&lt;p&gt;Students then generate a knowledge building move meant to advance this initial idea. There are also more complex game configurations that present not only a possible initial explanation in response to a problem, but also provide a series of possible knowledge building moves. In this case, students both evaluate the quality of the provided moves and generate a possible next move that contributes to the progressive improvement of ideas. In all cases, students each work on the same hypothetical game-configurations so that they can then compare and contrast their proposed knowledge building moves in whole-class discussions about issues such as what makes a “good contribution” and what does it mean to advance the community understanding.
(&lt;a href="http://dl.acm.org/citation.cfm?id=1854471&amp;amp;dl=ACM&amp;amp;coll=DL&amp;amp;CFID=579302205&amp;amp;CFTOKEN=60206663"&gt;Bielaczyc &amp;amp; Ow, 2010&lt;/a&gt;)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So perhaps if I did this again, with a better prompt, and students knew that they were taking notes that would be cut apart and reassembled, they would write differently. However, it is also possible that this kind of reorganization is not as useful for collaborative work – part of why it works when working on my own ideas, is that I wrote the notes myself, and I remember why I wrote them. Looking at other people's notes is difficult at best, and especially when they are out of context.&lt;/p&gt;

&lt;h1&gt;Future work&lt;/h1&gt;

&lt;p&gt;I still believe that Etherpad and the wiki, which is also scriptable, has a lot of potential for rapid experimentation with collaborative scripts. We have experimented with a number of other scripts in the two classes that I teach/co-teach, and I might blog about other ones later. I've used a simplified version of the tagging script in two later classes, where I focused on the ability to pick out and aggregate ideas based on pre-determined tags. &lt;em&gt;The fact that Etherpad records the entire history, and let's me go back, is also useful for research, and I plan to spend a bit longer looking at how the individual team pads evolved during the different stages of the script.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-10-03-a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content_-_half-08.png" alt="" class="img-right"&gt;The week after the script above, we talked about theoretical perspectives on development and technology, and as an ice-breaker before the lecture, I asked them to write down what they thought poverty was, and what they perceived the "official" definition to be. After some group discussion, I asked each group to come up with a single sentence for each, and tag them with respectively #ithink and #worldthinks. The script then simply aggregated all of these on a simple webpage, which was a great way of exploring different approaches (ethics, capability approach, materialistic, relative/objective poverty, happiness, etc).&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Writing these &lt;a href="/blog/2014/10/03/supporting-idea-convergence-through-pedagogical-scripts/"&gt;two&lt;/a&gt; blog posts has also been interesting. I know they ended up very long, and kind of mix technical stuff, CSCL theory, my own ideas about how people work with ideas, and my practical experience in the class. But this is how I think – and it's liberating to not be constrained by the academic paper format, especially in the exploratory phase. Writing it down enables me to reflect more deeply on my own design, and execution. Putting the script into Dillenbourg's orchestration graph format was also a great exercise, even though I have barely scratched the surface of his framework.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The drawing of an internet workflow comes from &lt;a href="/blog/images/2014-10-03-a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content_-_half-02.png"&gt;a student in the P2PU course&lt;/a&gt;. The Web 2.0 image comes from &lt;a href="http://www.nptechforgood.com/2010/01/28/web-1-0-web-2-0-and-web-3-0-simplified-for-nonprofits/"&gt;Nonprofit Tech for Good&lt;/a&gt;. The Do as I tell You sign comes from &lt;a href="http://icehousecrafts.com/item_371/If-You-Would-Just-Do-What-I-Tell-You-I-Wouldnt-Have-To-Be-So-Bossy-Sign.htm"&gt;Ice House Crafts&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <id>tag:reganmian.net,2014-10-03:/blog/blog/2014/10/03/supporting-idea-convergence-through-pedagogical-scripts/</id>
    <title type="html">Supporting idea convergence through pedagogical scripts and Etherpad APIs, an introduction</title>
    <published>2014-10-03T20:17:26Z</published>
    <updated>2014-10-03T20:17:26Z</updated>
    <link rel="alternate" href="https://reganmian.net/blog/2014/10/03/supporting-idea-convergence-through-pedagogical-scripts/"/>
    <content type="html">&lt;p&gt;&lt;strong&gt;We can script Etherpad to push discussion prompts out to many small groups, and then pull the information back. Using tags, we can extract information, and as a community organize the emerging folksonomy&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This blog post brings together two long-standing interests of mine. The first is how to script web tools to support small-group collaborative learning, and the other is how to support reorganization of ideas across groups.&lt;/em&gt;&lt;/p&gt;

&lt;h1&gt;Two meanings of the word "scripting"&lt;/h1&gt;

&lt;p&gt;There is an interesting intersection between two quite different meanings of the word scripting in the context of my work. In the CSCL literature, &lt;a href="http://edutechwiki.unige.ch/en/CSCL_script"&gt;scripts&lt;/a&gt; refer to sequences of activities that support groups of students in carrying out a collaborative learning activity. They can be very simple and generic, like the well-known &lt;a href="http://olc.spsd.sk.ca/DE/PD/instr/strats/jigsaw/"&gt;jigsaw script&lt;/a&gt;, or very content-specific. There is active on-going research on scripting, including &lt;a href="http://epub.ub.uni-muenchen.de/14328/"&gt;how external scripts interfer with internal scripts&lt;/a&gt;, and &lt;a href="http://hal.archives-ouvertes.fr/docs/00/19/02/30/PDF/Dillenbourg-Pierre-2002.pdf"&gt;the dangers of over-scripting&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-09-22-tagging-and-convergence-for-small-groups-collaboration-with-multiple-etherpad_-_whole-02.png" alt=""&gt;&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;From the technology side, we talk about writing scripts as a light-weight form of programming, to ask the computer to do something for us. Scripting a program, or a web site, means that we can interact with it automatically, for example asking Google Docs to set up 20 documents for us. There is obviously a connection between developing educational scripts and actually "implementing" them in technology. Often, specialized software has been written to enable specific scripts, but there has been some attempts at enabling more generic implementations of learning designs into software.&lt;/p&gt;

&lt;p&gt;One inspiring example for me has been &lt;a href="http://www.gsic.uva.es/glueps/"&gt;GLUE-PS&lt;/a&gt; (&lt;a href="http://libgen.org/scimag/get.php?doi=10.1016/j.compedu.2013.12.008"&gt;paper&lt;/a&gt;), which can take a learning design from a learning design software, like &lt;a href="http://compendiumld.open.ac.uk/"&gt;CompendiumLD&lt;/a&gt;, and set-up the necessary software, for example pre-generating Google Docs and wiki-pages, splitting students into groups, and assigning different resources to different groups, etc.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-09-22-tagging-and-convergence-for-small-groups-collaboration-with-multiple-etherpad_-_whole-01.png" alt=""&gt;&lt;/p&gt;

&lt;h1&gt;Etherpad for collaborative writing, and more&lt;/h1&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-09-22-tagging-and-convergence-for-small-groups-collaboration-with-multiple-etherpad_-_half-01.png" alt="" class="img-right"&gt;For those who have not heard of, or used, Etherpad, I always introduce it as "like Google Docs". It enables very easy collaborative editing of documents, where everyone logged in immediately see all changes made by other users. It features very light-weight formatting, and is not designed to generate print-ready documents, but is perfect for brainstorming, note taking, etc. It also focuses more on the live-editing situation, with the text written by each user colored differently. It can be quite mesmerizing to see a document being filled in by multiple people at the same time.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-09-22-tagging-and-convergence-for-small-groups-collaboration-with-multiple-etherpad_-_half-02.png" alt="" class="img-left"&gt;We often used Etherpad for collaborative meetings at P2PU (&lt;a href="http://en.flossmanuals.net/etherpad/case-studies/"&gt;case study&lt;/a&gt;), and I also had some interesting experiences using in the P2PU course &lt;a href="https://p2pu.org/en/groups/introduction-to-the-field-of-computer-supported-co/content/full-description/"&gt;"Intro to Computer-Supported Collaborative Learning"&lt;/a&gt;. We began by chatting in the chat feature, but gradually migrated over to use the main editing space, which enabled us to keep multiple separate discussions going simultaneously. I have &lt;a href="/wiki/analysis_of_cscl-intro/#synchronous_meetings"&gt;written&lt;/a&gt; and &lt;a href="https://www.youtube.com/watch?v=kGXRdj9F_4E&amp;amp;list=UUEqYaJY03O0tC9Q1oQuXtKA"&gt;spoken&lt;/a&gt; about this experience, and it's something I'd love to experiment more with.&lt;/p&gt;

&lt;p&gt;For a large course I am currently helping teach, we began by using Etherpads to support note-taking in small group discussions. Setting up 10-15 Google Docs, remembering to set the sharing options correctly, and copying and pasting the URLs to the wiki was a lot of work, and it was much easier to do it with Etherpad. The students all found Etherpad very easy to use, and it was nice to be able to pull up group notes on the projector afterwards while groups were presenting their ideas to the rest of the class.&lt;/p&gt;

&lt;h1&gt;Scripting Etherpad&lt;/h1&gt;

&lt;p&gt;Etherpad has a nice simple API, which let's you create new pads and fill them with content, and read the content of existing pads. (Since there are no permissions or logins, that's basically all the actions that you would need). Since Etherpad conserves the full history of all changes, you can also access older versions of a pad. Here is how little code you need to get started.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A simple helper function to execute Etherpad API calls, and properly format the return message:&lt;/strong&gt;&lt;/p&gt;

&lt;pre&gt;&lt;code class="python"&gt;def run_etherpad(path, **params):
    params['apikey'] = settings.ETHERPAD_API
    data = urlencode(dict(params)).encode('ascii')
    url = ("%s/api/1.2.10/%s" % (settings.ETHERPAD_URL, path))
    r = json.loads(urlopen(url, data).read().decode('utf-8'))
    return(r)
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;And then I can simply execute any of the API commands listed in the &lt;a href="http://etherpad.org/doc/v1.3.0/#index_api_methods"&gt;Etherpad documentation&lt;/a&gt;, for example:&lt;/p&gt;

&lt;pre&gt;&lt;code class="python"&gt;text = run_etherpad("getText", padID=pad)
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;The first thing I did with this, was to write a little script that took a template (typically a few questions for discussion), and created a number of identical pre-populated Etherpads for the discussion groups. It then generated an HTML page with all the links, which I could copy and paste into the course wiki. This already saved a lot of time. Given that we now have a list of all the Etherpads, it's just as easy to pull all the text from the Etherpads back into a script, and do something with it.&lt;/p&gt;

&lt;p&gt;For example, I can generate a quick HTML page consisting of the text from all of the Etherpads together, which makes it much faster to quickly scan through what groups have been doing, rather than opening 15 tabs in the browser. Or I can automatically pull each group's page and post it to the wikipage of that group for easier access in the future (Etherpad is great for live editing, but not very trustworthy as a long-term repository of information).&lt;/p&gt;

&lt;h1&gt;Using tags to extract and organize ideas&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;A little detour about making ideas moveable, supporting organizing and synthesis&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;One of the things we found from teaching this course last year, is that the ideas and information entered tended to be "stuck" on the week and group-specific pages. After several weeks of the course, we would have lot's of great reflections, notes from readings, project ideas, etc., but they were not organized in a way that was easily accessible to the students. When working on their final projects, the students were more likely to use Google, rather than accessing the common knowledge base which we had tried to structure throughout the course.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-09-22-tagging-and-convergence-for-small-groups-collaboration-with-multiple-etherpad_-_half-04.png" alt="" class="img-left"&gt;I have long been inspired by a course that I took in the first year of my MA, which used an environment called Knowledge Forum. I made a &lt;a href="https://vimeo.com/17143638"&gt;screencast&lt;/a&gt; to showcase the course, and how the environment enabled us to go back to our old notes, reorganize them, and see new connections and gaps. At that time, I contrasted it with threaded discussion forums, where a post is usually "captured" in the chronology/thread where it is posted.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-09-22-tagging-and-convergence-for-small-groups-collaboration-with-multiple-etherpad_-_half-03.png" alt="" class="img-right"&gt;Inspired by this course, and by workshop methodology, I began thinking about what I called &lt;strong&gt;"the cycle of divergence an convergence"&lt;/strong&gt; which I began exploring in a &lt;a href="/wiki/grappling_with_ideas/"&gt;talk&lt;/a&gt; and a &lt;a href="/wiki/grappling_with_ideas-the_paper/"&gt;paper&lt;/a&gt;. I also thought that &lt;a href='/wiki/tagging/'&gt;tagging&lt;/a&gt; could be a low-tech way to enable "movable ideas". An example script could be the following: after a class had discussed a topic for a few weeks, each week with a new "input" (stimulus), they could collectively or individually decide on a few emerging overarching themes, and go back through the posts, tagging them accordingly. The discussion forum software would then create "saved searches" for these tags -- dynamic folders containing all the posts tagged with a certain tag. Students could then revisit these posts in a new context, continue the discussion, etc.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-09-22-tagging-and-convergence-for-small-groups-collaboration-with-multiple-etherpad_-_half-05.png" alt="" class="img-left"&gt;This turns out to be very similar to the process of qualitative research, and a methodology supported by tools like NVivo (&lt;a href="https://www.youtube.com/watch?v=0YyVySrV2cM"&gt;good video intro&lt;/a&gt;), where one codes individual pieces of text, and then pulls up a list of all the coded pieces for a given code, to see them in a new context. Something that might be useful for quantitative research, for organizing a literature review... and for supporting community knowledge and deep thinking in a collaborative learning situation?&lt;/p&gt;

&lt;h1&gt;Tag-extract, short recap&lt;/h1&gt;

&lt;p&gt;&lt;img src="/files/uploads/2012/06/Screen-Shot-2012-06-13-at-10.50.59.png" alt="" class="img-left"&gt;Programmers both like to reinvent the wheel, and to scratch their own itch, so faced with a large amount of notes that I needed to structure into a literature review, while keeping the source-information (which paper did a certain idea come from), I experimented with tagging. (Adding some markup to the text, and then processing it with a script, is far easier than developing a new graphical interface for tagging). I had three simple rules, or principles, which turned out to be quite powerful:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;all the text below a source indicator "belongs" to that source, and will always be tagged as coming from that source&lt;/li&gt;
&lt;li&gt;any tag applies to the line it is on&lt;/li&gt;
&lt;li&gt;all text indented below a tagged line, belongs to that line&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;img src="/files/uploads/2012/06/Screen-Shot-2012-06-13-at-10.57.31.png" alt="" class="img-right"&gt;I wrote a &lt;a href="/blog/2012/06/13/tag-extract-a-tool-to-automatically-restructure-textoutline-using-tags/"&gt;detailed blog-post&lt;/a&gt; about this tool, complete with &lt;a href="https://www.youtube.com/watch?v=NEfdPDptD5U"&gt;an extensive screencast&lt;/a&gt;, showing the tool embedded in an academic workflow process, and &lt;a href="https://www.youtube.com/watch?v=NEfdPDptD5U"&gt;a more focused screencast&lt;/a&gt; focusing on the tool itself.&lt;/p&gt;

&lt;p&gt;At the top, you see a text file with notes, and bibdesk-identifiers as the "source". After adding tags, and running the script, I get the output seen on the right, where the notes are reordered according to tag, but with the source still there. When writing about for example metalearning, you then have all the relevant ideas/quotes, together with the article citations. Here's &lt;a href="/wiki/litreview_raw_sorted/"&gt;an example of raw sorted notes&lt;/a&gt; about open courses, and &lt;a href="/wiki/draft_literature_review_open_courses/"&gt;here the resulting draft&lt;/a&gt;.&lt;/p&gt;

&lt;h1&gt;Bringing all the pieces together&lt;/h1&gt;

&lt;p&gt;So we have introduced the idea of pedagogical scripting, as well as implementing scripts in computer code. I've discussed the desire to make ideas more "moveable", and support deeper work on ideas, and talked about the idea of using tags to support this. Finally, I introduced the tool tag-extract, which I developed to work on a literature review. This blog post is already long enough, so I will write a &lt;strong&gt;&lt;a href="/blog/2014/10/03/a-pedagogical-script-for-idea-convergence-through-tagging-etherpad-content/"&gt;separate blog post&lt;/a&gt;&lt;/strong&gt; about the actual design, implementation and evaluation of the pedagogical script using these ideas.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The script flowchart is from an unpublished manuscript by Pierre Dillenbourg. The GLUE-PS image is from the &lt;a href="http://www.gsic.uva.es/wikis/gs2/index.php/LDGResourcesGLUEPS"&gt;GLUE-PS website&lt;/a&gt;, the concept map about concept maps from the &lt;a href="http://cmapskm.ihmc.us/viewer/cmap/1064009710027_1483270340_27090"&gt;Cmap site&lt;/a&gt;, and the NVivo image from the &lt;a href="http://ebabbie.net/resource/NVivo/primer.html"&gt;NVivo primer&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <id>tag:reganmian.net,2014-09-22:/blog/blog/2014/09/22/easy-interoperability-between-ruby-and-python-scripts-with-json/</id>
    <title type="html">Easy interoperability between Ruby and Python scripts with JSON</title>
    <published>2014-09-22T21:57:15Z</published>
    <updated>2014-09-22T21:57:15Z</updated>
    <link rel="alternate" href="https://reganmian.net/blog/2014/09/22/easy-interoperability-between-ruby-and-python-scripts-with-json/"/>
    <content type="html">&lt;p&gt;I recently needed to call a Ruby script from Python to do some data processing. I was generating some Etherpad-scripts in Python, and needed to restructure tags (using &lt;a href="/blog/2012/06/13/tag-extract-a-tool-to-automatically-restructure-textoutline-using-tags/"&gt;tag-extract&lt;/a&gt;) with a Ruby script. The complication was that this script does not just return a simple string or number, but a somewhat complex data structure, that I needed to process further in the Python script.&lt;/p&gt;

&lt;p&gt;Luckily, searching online I came across the idea to use JSON as the interchange format, which worked swimmingly. Given that all the information was in text format, and the data structure was not that complex (just some lists and dictionaries), JSON could cope well with the complexity, and was easier to debug, since it's a text format. If I had had other requirements, like binary data, I would have had to investigate other data formats.&lt;/p&gt;

&lt;!-- more --&gt;

&lt;p&gt;&lt;a href="https://github.com/houshuang/folders2web/blob/master/tag-extract.rb"&gt;The Ruby script&lt;/a&gt; already had several different output options, controlled through command-line switches. I added the option to output through JSON, and the code required to do the actual output was simply:&lt;/p&gt;

&lt;pre&gt;&lt;code class="ruby"&gt;puts JSON.generate(tags)
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;I also reconfigured the script to accept input from standard-in:&lt;/p&gt;

&lt;pre&gt;&lt;code class="ruby"&gt;if ARGV.size == 0
  a = ARGF.read
else
  a = try { File.read(ARGV[1]) }
  unless a
    puts "Could not read input file"
    exit
  end
end
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;This meant that I could simply pipe the information from the Python script through the Ruby script, writing to standard-out and reading from standard-in, not having to create any temporary files at all.&lt;/p&gt;

&lt;p&gt;I created a helper function in Python to let me run a command, putting into standard-out, and reading from standard-in:&lt;/p&gt;

&lt;pre&gt;&lt;code class="python"&gt;def apply_external(cmd_ary, out):
    proc = subprocess.Popen(
        cmd_ary,stdout=subprocess.PIPE,
        stdin=subprocess.PIPE)
    proc.stdin.write(bytes(out, "UTF-8"))
    proc.stdin.close()
    result = proc.stdout.read()
    result = result.decode("UTF-8")
    return(result)
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;and then I simply call it with the text I want transformed, and receive the transformed text structure back.&lt;/p&gt;

&lt;pre&gt;&lt;code class="python"&gt;def tag_extract(text):
    ret = apply_external(['ruby', '/Users/Stian/src/folders2web/tag-extract.rb'], text)
    return json.loads(ret)
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Although the idea was simple, there was a bit of fiddling getting the pipes set up in Python, etc., so I figured that documenting it might be helpful to others.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <id>tag:reganmian.net,2014-04-16:/blog/blog/2014/04/16/creating-anki-cards-for-russian-coursera-mooc-with-stemming-and-frequency-lists/</id>
    <title type="html">Creating Anki cards for Russian Coursera MOOC with stemming and frequency lists</title>
    <published>2014-04-16T18:44:51Z</published>
    <updated>2014-04-16T18:44:51Z</updated>
    <link rel="alternate" href="https://reganmian.net/blog/2014/04/16/creating-anki-cards-for-russian-coursera-mooc-with-stemming-and-frequency-lists/"/>
    <content type="html">&lt;p&gt;&lt;em&gt;In which I automatically generate Anki review cards for vocabulary based on subtitles from a Russian Coursera MOOC&lt;/em&gt;&lt;/p&gt;

&lt;h1&gt;Learning Russian, a 10-year project&lt;/h1&gt;

&lt;p&gt;I have been working on my Russian on and off for many years, I'm at the level where I don't feel the need for textbooks, but my understanding is not quite good enough for authentic media yet. I've experimented with &lt;a href="/blog/2012/02/02/my-one-month-russian-challenge/"&gt;readings novels in parallel&lt;/a&gt; (English and Russian side by side, or Swedish and Russian, like on the picture), and I've been listening to &lt;a href="http://www.tasteofrussian.com/"&gt;a great podcast&lt;/a&gt; for Russian learners.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-04-16-creating-anki-cards-for-russian-coursera-mooc-with-stemming-and-frequency-lists_-_whole-01.png" alt=""&gt;&lt;/p&gt;

&lt;h1&gt;Language learning with MOOCs&lt;/h1&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-04-16-creating-anki-cards-for-russian-coursera-mooc-with-stemming-and-frequency-lists_-_half-03.png" alt="" class="img-left"&gt;
MOOCs can be a great resource for language learning, whether intentionally or not. At the University of Toronto, &lt;a href="http://www.ocw.utoronto.ca/demographic-reports/"&gt;we found&lt;/a&gt; that more than 60% of learners across all MOOCs spoke English as a second language, no doubt some of them are not just viewing the foreign language of the MOOC as a barrier, but are also hoping to improve their English. To my delight, the availability of non-English language MOOCs has been growing steadily. For example, Coursera has courses in Chinese, French, Spanish, Russian, Turkish, German, Hebrew and Arabic. There are also MOOC providers focused on specific linguistic areas, for example China's &lt;a href="http://xuetangx.com/"&gt;XueTangX&lt;/a&gt;, France's &lt;a href="https://www.france-universite-numerique-mooc.fr/"&gt;Université Numérique&lt;/a&gt;, and Germany's &lt;a href="https://iversity.org/"&gt;iversity&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The existence of global open educational resources is nothing new. Five years ago, I wanted to share my enthusiasm about being able to "peek through the windows" of universities around the world, and edited &lt;a href="https://www.youtube.com/watch?v=eRbWXKnxB2c"&gt;a YouTube mashup&lt;/a&gt;. I've also blogged about &lt;a href="/blog/2010/12/07/oer-for-a-multicultural-classroom-student-as-user-and-producer"&gt;OER for a multicultural classroom&lt;/a&gt;, and written about OER from &lt;a href="/blog/the-chinese-national-top-level-courses-project/"&gt;China&lt;/a&gt;, &lt;a href="/blog/2008/12/05/worlds-largest-university-opens-almost-all-its-materials/"&gt;India&lt;/a&gt;, &lt;a href="/blog/2009/03/19/407-indonesian-textbooks-openly-available/"&gt;Indonesia&lt;/a&gt;, etc.&lt;/p&gt;

&lt;p&gt;However, earlier OER videos tended to be full lecture recordings, with poor video and audio quality, often on proprietary and hard to use platforms, with very little facilitation for online learning. Current MOOC platforms have converged around short (5-15 minutes) videos, with high video and audio quality, and decent video players. The courses are coherent, designed for online learning, and often feature supporting readings, sometimes even subtitles, etc.&lt;/p&gt;

&lt;h1&gt;A Russian MOOC&lt;/h1&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-04-16-creating-anki-cards-for-russian-coursera-mooc-with-stemming-and-frequency-lists_-_whole-02.png" alt=""&gt;&lt;/p&gt;

&lt;p&gt;The picture above is from a lecture in the MOOC &lt;a href="https://class.coursera.org/historyofec-001"&gt;"История экономической мысли"&lt;/a&gt; (History of Economic Thought), which is currently in it's second week on Coursera. I signed up for this several months ago, because I wanted to try out my Russian, and I find the topic very interesting. (Extra interesting, of course, is the question of whether this topic will be taught any differently from a country that has a very unique history with regard to economic thought and experiments). The fact that I know something about the topic, and about European history in general, makes it easier to follow the lectures, as do the subtitles, the ability to regulate the speed, and the occasional illustrations in the video. Watching a few of the videos, I found that I could more or less follow along, but there were a number of words that seemed important, and which I could not understand.&lt;/p&gt;

&lt;h1&gt;Stemming and word frequencies&lt;/h1&gt;

&lt;p&gt;I've always been interested in language technology, and last year I spent some time experimenting with different libraries to create word frequency lists from Russian electronic texts. I used a library to "stem words" -- reducing all the conjugated forms of a word to a single word -- for example all these forms of the word &lt;code&gt;big&lt;/code&gt;: &lt;em&gt;большая, больше, большей, большие, большим, большими, большое, большой, большую&lt;/em&gt;. There are several possible libraries for doing this, none are perfect, but most are quite good. In &lt;a href="/blog/2013/10/17/playing-with-word-stemming-and-frequencies-in-russian/"&gt;my previous blog post&lt;/a&gt;, I used the Ruby gem &lt;code&gt;ruby-stemmer&lt;/code&gt;, however when I revisited my code, I noticed that I had switched to using &lt;code&gt;hunspell&lt;/code&gt; after writing the blog post.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Hunspell&lt;/code&gt; is an open-source spellchecker, which is also able to output word stems. In the example below, I simply send a Russian sentence through &lt;code&gt;hunspell&lt;/code&gt;, and get back a list with the original conjugated word on the left, and the stem on the right (I've removed the words that didn't change)&lt;/p&gt;

&lt;pre&gt;&lt;code class="bash"&gt;❯❯❯ echo "Сегодня наше предствавление об экономике очень сильно отличается от того,
    как люди древности эээ, мыслили то, что мы сегодня называем экономикой." |
    hunspell -d ru_RU,ru_RU_google -s

экономике экономика
сильно сильный
отличается отличаться
древности древность
мыслили мыслить
называем называть
экономикой экономика
&lt;/code&gt;&lt;/pre&gt;

&lt;h1&gt;Getting word frequencies from a MOOC subtitle file&lt;/h1&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-04-16-creating-anki-cards-for-russian-coursera-mooc-with-stemming-and-frequency-lists_-_whole-03.png" alt="" class="img-left"&gt;
Not only did I have this old project with stemming and creating frequency lists of Russian words, but I have recently done some experiments with downloading video subtitles from Coursera courses for use in a machine learning project. &lt;a href="https://github.com/houshuang/russian_stemmer/blob/master/parse_subtitles.py"&gt;This script&lt;/a&gt; takes an HTML file with the list of course lectures (you need to be logged in to get to this page, so the easiest is to open it in the browser, and save the page), extracts all the subtitle links, downloads them, and concatenates them to one output file.&lt;/p&gt;

&lt;p&gt;We can then pass this file through the stemmer (all code for this blogpost on &lt;a href="https://github.com/houshuang/russian_stemmer/"&gt;Github&lt;/a&gt;), and generate a list of all word-stems that occur more than five times, sorted by general frequency. (I use a Russian frequency word list I found &lt;a href="http://bokrcorpora.narod.ru/frqlist/frqlist-en.html"&gt;here&lt;/a&gt;). The subtitles for the first two weeks of videos (18 videos) total 27,600 words (50-80 pages of printed text). There are almost 6,400 unique words (of which сельскохозяйственного -- agricultural -- is the longest), but only 3374 stems. By removing all stems that occur less than five times, we're down to a much more manageable 839 stems.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-04-16-creating-anki-cards-for-russian-coursera-mooc-with-stemming-and-frequency-lists_-_half-01.png" alt="" class="img-right"&gt;
I opened these in a text editor, and manually removed words I already knew (I left in many words which I kind of knew, but would like to learn more precisely). Since the words were sorted in order of global frequency, with more "rare" words near the top, I could remove most of the lower half of the file in one swoop (words like "I", "like", "and", etc). Within half an hour, I had now gone from 27,600 words to 184 word stems that I needed to learn or study.&lt;/p&gt;

&lt;p&gt;No problem, just copy the whole list of Russian words into Google Translate, and out comes an equally long list of English corresponding words, which I save to another text file. (These are not 100% perfect, but in my experience, it works for 95% of the words).&lt;/p&gt;

&lt;h1&gt;Bringing it all together - Anki and spaced repetition&lt;/h1&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-04-16-creating-anki-cards-for-russian-coursera-mooc-with-stemming-and-frequency-lists_-_half-02.png" alt="" class="img-right"&gt;&lt;/p&gt;

&lt;p&gt;A common concept on language learning blogs is &lt;a href="http://en.wikipedia.org/wiki/Spaced_repetition"&gt;"spaced repetition"&lt;/a&gt;, the concept of having a computer automatically determine the optimal repetition rate to counteract the natural rate of forgetting. &lt;a href="http://ankisrs.net/"&gt;Anki&lt;/a&gt; is a great free program for practicing flashcards using spaced repetition, with very powerful options to create your own cards, configuring the rate of repetition, etc. I have only ever used it before with public decks, but I thought I would try to use the tools I introduced above to make my own cards.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-04-16-creating-anki-cards-for-russian-coursera-mooc-with-stemming-and-frequency-lists_-_half-04.png" alt="" class="img-left"&gt;
The simplest cards can be represented as two columns in a file, the front and the back. You are presented with the word on the front, try to remember the word on the back, turn the card, and indicate whether you were successful or not. However, Anki supports much more powerful card configurations. You create a new note type, with arbitrary many fields, and then you can create different card designs, incorporating these designs.&lt;/p&gt;

&lt;p&gt;I decided that it would be very helpful to see the words in context as well. And not just in an arbitrary context, but in the specific context from the course. Given that I had already gone from specific words to the word stem, I could go back to using those specific words to search for example sentences, add highlighting of the word in question, and include these in the file. Each card has a Russian word on the front, and when you turn it, you see the English translation (from Google Translate), the Russian stem (for reference), and up to three example sentences with the word highlighted.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-04-16-creating-anki-cards-for-russian-coursera-mooc-with-stemming-and-frequency-lists_-_whole-04.png" alt="" class="img-right"&gt;
I think the result is pretty awesome. It will be fun to practice on these cards, and then go back to watching the videos. When next week's videos are released, I should be able to quickly generate a list of words that are new (were not covered in week 1 or 2).&lt;/p&gt;

&lt;p&gt;There are of course many ways in which this could be improved. I'm thinking of how to keep track of words I know, so that I can easily remove them when getting a word list from new material. Integrating this kind of functionality into the Coursera or other MOOC platforms would be neat, although it will probably never happen. (If they had more open APIs, other people could build it, however). I could for example easily extract the timing of the example sentences, and include that in the flash cards, so you could jump straight to a specific video at the point where the professor says a certain example sentence (however, I don't know how to generate URLs that open a Coursera video at a specific time).&lt;/p&gt;

&lt;p&gt;There's also a lot of other fun stuff one could do with subtitles - it would be great if we could get easy access to subtitles for all the courses, and use it to build a search engine (linked directly to videos etc.)&lt;/p&gt;

&lt;p&gt;But for now, I need to study some flashcards, and watch some Russian videos. Пока!&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(All the code on &lt;a href="https://github.com/houshuang/russian_stemmer"&gt;Github&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <id>tag:reganmian.net,2014-03-11:/blog/blog/2014/03/11/learning-idiomatic-haskell/</id>
    <title type="html">Learning idiomatic Haskell with Exercism.io</title>
    <published>2014-03-11T15:50:24Z</published>
    <updated>2014-03-11T15:50:24Z</updated>
    <link rel="alternate" href="https://reganmian.net/blog/2014/03/11/learning-idiomatic-haskell/"/>
    <content type="html">&lt;p&gt;I got my introduction to functional programming through Clojure, but lately I've been really fascinated by Haskell. I gave up on it a few times, because it can seem quite impenetrable, and very different from anything that I'm used to. But there is something about the elegance, and powerful ideas that keeps me coming back. I've read a bunch of tutorials and papers, followed blog posts, and experimented a bit with ghci and &lt;a href="https://github.com/gibiansky/IHaskell"&gt;IHaskell&lt;/a&gt;, but I never really got off the starting block with writing my own Haskell.&lt;/p&gt;

&lt;h1&gt;How to get practice?&lt;/h1&gt;

&lt;p&gt;The projects I'm currently working on in Python are too complex and urgent for me to try to implement them in Haskell at this point, and because Haskell is so different, I found myself stymied doing even simple things, even though I'd just finished reading a complex CS paper about monads. I needed some structured tasks that were not too hard, and that gave me good feedback on how to improve.&lt;/p&gt;

&lt;p&gt;&lt;a href="http://exercism.io/"&gt;&lt;img src="/blog/images/2014-03-11-learning-idiomatic-haskell_-_whole-02.png" alt=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Listening to the &lt;a href="http://www.functionalgeekery.com/"&gt;Functional Geekery podcast&lt;/a&gt;, I heard about &lt;a href="http://exercism.io/"&gt;exercism.io&lt;/a&gt;, which provides exercises for many popular languages. There are several similar websites in existence, perhaps the most well known is &lt;a href="https://projecteuler.net/"&gt;Project Euler&lt;/a&gt;, however that focuses too much on CS/math-type problems, and is perhaps better for learning algorithms than a particular language. The way exercism.io works, is that you install it as a command line application. The first time you run it, it downloads the first exercise for a number of languages.&lt;/p&gt;

&lt;p&gt;The exercises themselves consist of test cases, and it's up to you to write a module that will fulfill the test case. Once you are done, you use the command line client to quickly submit the answer, which unlocks everyone else's answers, and downloads the next exercise. The fact that you download the test cases and README files, and work on your local computer, is also great. You can use the environment you're used to (Emacs, VIM, EclipseFP), with all the tooling, documentation, autorunning tests, etc., that you are used to.&lt;/p&gt;

&lt;h1&gt;Trying before being told&lt;/h1&gt;

&lt;p&gt;This kind of graded approach is brilliant, because it precludes "cheating". It's easy to skip ahead to the solution, say "that makes sense, I would have come up with that", and move on. But it's important to actually force yourself &lt;em&gt;to come up with that&lt;/em&gt;. Reading other people's versions, and the feedback they received, after having struggled with the problem yourself, is also incredibly informative -- and again, I think you get far more out of it, than just browsing through solutions before you've put in the hard work yourself.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;There's a parallel to some of the work I'm doing with audience response systems in lectures. &lt;a href="http://mazur.harvard.edu/research/detailspage.php?rowid=8"&gt;Eric Mazur at Harvard&lt;/a&gt; puts physics problems on the projector, and asks students to first try to solve them themselves. Then they have to try to convince the person next to them about the right answer, which they submit by clicker. Finally, he explains and discusses. Students who have wrestled with the problem for a while already, and been exposed to other people's ways of thinking, are far more receptive and learn more from the final explanation, than if you had just gone directly to lecture.&lt;/em&gt;&lt;/p&gt;

&lt;h1&gt;Idiomatic code&lt;/h1&gt;

&lt;p&gt;The objective is not just to pass the tests (which is in most cases trivial, for example it would not be difficult to write very specific code that just fulfills the test cases, but does not generalize), but to write as beautiful, idiomatic code as possible. One might start with an inelegant approach that fulfills the test cases, and then begin refactoring into something one can be proud of, while making sure the test-cases continue to work.&lt;/p&gt;

&lt;p&gt;This is an interesting challenge, and it's also very useful, because one of the biggest challenges when learning a new language is internalizing the metaphors and idioms of that language. When writing a simple function, there might be many ways of writing it, that will all "work", and yield the same result. Yet once you begin writing larger and more complex programs, programs that need certain kinds of performance (memory, CPU, dealing with very large inputs etc), the difference between different approaches becomes very large.&lt;/p&gt;

&lt;p&gt;The first project I did was quite embarassing. The task asks you to construct a function that takes a string, and generates a string as a response. There are a bunch of test cases, and you have to figure out yourself how to group them (seems like strings ending with ? are questions, but not if they are all upper-case, etc). I decided to use pattern matching, which was a fair approach.&lt;/p&gt;

&lt;p&gt;However, I instinctively used regexp patterns to match the text, instead of looking for built-in functions. I got it to work mostly (and did learn how to use regexp in Haskell, which I hadn't used before), but I had problems with unicode, and spent too long trying different patterns. After I submitted, I saw an incredibly elegant solution, using only built-in predicates like &lt;code&gt;isAlphaNum&lt;/code&gt;, which read and compose much better than my convoluted regexps. A very good lesson, and one I internalized much quicker than if I read a blog post about it, without having struggled with it myself first.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-03-11-learning-idiomatic-haskell_-_whole-01.png" alt=""&gt;&lt;/p&gt;

&lt;h1&gt;Future growth of exercism.io&lt;/h1&gt;

&lt;p&gt;The platform is quite young, and I'm wondering how it will scale. For the first few exercises in Haskell, there were only about 15 responses (since you cannot proceed until you have submitted the first few responses, I'm guessing the more advanced exercises will have even fewer answers). If there were hundred answers for a given problem, I might want to see as many different approaches as possible, not 20 examples that are basically the same (I'm also interested in as much diverse feedback as possible).&lt;/p&gt;

&lt;p&gt;Perhaps some kind of voting feature is necessary. There is the possibility of providing feedback to individual submissions, and while I really appreciate people's feedback, I almost feel like it isn't necessary once I've seen other people's contributions -- I already have a good idea of what I could do better. (There are only &lt;em&gt;so&lt;/em&gt; many different ways of doing these fairly small and bounded problems).&lt;/p&gt;

&lt;p&gt;Anyway, I was very excited to discover this platform. After the first task, which I spent about two hours on, I did two more. With these, I was less embarassed when seeing other people's code, but in both cases, there were more elegant solutions than the one I suggested. One of the recurring themes also seems to be really understanding the core library, and all the functions that are already implemented, instead of trying to reinvent the wheel.&lt;/p&gt;

&lt;p&gt;Go learn some code!&lt;/p&gt;
</content>
  </entry>
  <entry>
    <id>tag:reganmian.net,2014-03-10:/blog/blog/2014/03/10/parsing-massive-clicklogs/</id>
    <title type="html">Parsing massive clicklogs, an approach to parallel Python</title>
    <published>2014-03-10T15:09:17Z</published>
    <updated>2014-03-10T15:09:17Z</updated>
    <link rel="alternate" href="https://reganmian.net/blog/2014/03/10/parsing-massive-clicklogs/"/>
    <content type="html">&lt;p&gt;I am currently working on analyzing MOOC clicklog data for a research project. The clicklogs themselves are huge text files (from 3GB to 20GB in size), where each line represents one "event" (a mouse click, pausing a video, submitting a quiz, etc). These events are represented as a (sometimes nested) JSON structure, which doesn't really have a clear schema.&lt;/p&gt;

&lt;h1&gt;Introduction&lt;/h1&gt;

&lt;p&gt;Our goal was to run frequent-sequence analyses over these clicklogs, but to do that, we needed to process them in two rounds. Initially, we walk throug the log-file, and convert each line (JSON blob) into a row in a &lt;a href="http://pandas.pydata.org/"&gt;Pandas&lt;/a&gt; table, which we store in a &lt;a href="http://www.hdfgroup.org/HDF5/"&gt;HDF5&lt;/a&gt; file using &lt;a href="http://www.pytables.org/moin"&gt;pytables&lt;/a&gt; (&lt;a href="http://pandas.pydata.org/pandas-docs/dev/cookbook.html#cookbook-hdf"&gt;see also&lt;/a&gt;). We convert the JSON key-values to columns, extract information from the URL (for example &lt;code&gt;/view/quiz?quiz_id=3&lt;/code&gt; results in the column action receiving the value &lt;code&gt;/view/quiz&lt;/code&gt;, and the column &lt;code&gt;quiz_id&lt;/code&gt;, the value 3. We also do a bit of cleaning up of values, throw out some of the columns that are not useful, etc.&lt;/p&gt;

&lt;h1&gt;Speeding it up&lt;/h1&gt;

&lt;p&gt;We use Python for the processing, and even with a JSON parser written in C, this process is really slow. An obvious way of speeding it up would be to parallelize the code, taking advantage of the eight-cores on the server, rather than only maxing out a single core. I did not have much experience with parallel programming in general, or in Python, so this was a learning experience for me. I by no means consider myself an expert, and I might have missed some obvious things, but I still thought what we came up with might be useful to others.&lt;/p&gt;

&lt;p&gt;There are many approaches to parallelizing code, and choosing the right one requires thinking about your algorithms -- which functions need access to shared state (and can you design it so that you minimize shared state), are you CPU-bound, or IO-bound, what is the grain size that it makes sense to split at, etc.&lt;/p&gt;

&lt;h1&gt;Shared state&lt;/h1&gt;

&lt;p&gt;I was using HDF5 as my storage format, and pytables can append to an existing table. Theoretically, one would think that the lines in the log file are completely independent from each other, just get them converted and into the HDF5, and once you've got them all entered (appended), you can index and sort by timestamp, username, etc. However, Pandas Data.Frames have one big limitation compared to R data.frames -- no categorical column type. In R, a data.frame might have a column that stores the answer to a likert-type question (&lt;code&gt;Not at all&lt;/code&gt; to &lt;code&gt;Totally agree&lt;/code&gt;, for example), and even though there are 10,000 rows, those five text answers are only stored once, with each cell only containing a pointer to the correct alternative. This makes the table take much less space.&lt;/p&gt;

&lt;p&gt;Pandas was developed for financial analysis, and still seems fairly focused on numbers. (It also builds upon NumPy, which also doesn't have a categorical format). In our case, this was a big problem, since the difference between storing "/view/lecture" (13 bytes) and 1 (1 byte) for a table with 14 million rows is huge! (Especially when you consider that Pandas does much of its work by loading the entire table into memory).&lt;/p&gt;

&lt;p&gt;I decided to hack my own little categorical format, by creating a memoize function, which would assign a unique number to each new value it was called with, but always return the same number for the same value. Quick example:&lt;/p&gt;

&lt;pre&gt;&lt;code class="python"&gt;memoize("hello") #       =&amp;gt; 1
memoize("how are you") # =&amp;gt; 2
memoize("hello") #       =&amp;gt; 1
memoize("how are you") # =&amp;gt; 2
memoize("i'm fine") #    =&amp;gt; 3
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Initially the function simply kept state in a Pandas Series, and once I had finished parsing the entire clicklog, I wrote these series to the same HDF5 file (HDF5 files can contain multiple tables, files etc). This worked great, and it meant that instead of this:&lt;/p&gt;

&lt;pre&gt;&lt;code class="json"&gt;{"key": "user.video.lecture.action","value": "{\"currentTime\":257.066689,
\"playbackRate\":1,\"paused\":false,\"error\":null,\"networkState\":1,
\"readyState\":4,\"eventTimestamp\":13752243249891,\"initTimestamp\":
137524554345511,\"type\":\"ratechange\",\"prevTime\":227.066689}","username":
"9930a0523403499b7347707c92bccbcbff1a","timestamp": 13752234234277,"page_url":
"https://class.mooc-platform.org/course-001/lecture/view?lecture_id=31","client":
"spark","session":"30423409-1323423042304","language": "en-US,en;q=0.5","from":
"https://class.mooc-platform.org/course-001/lecture/view?lecture_id=31",
"user_ip": "123.255.122.167","user_agent": "Mozilla/5.0 (Windows NT 6.1; rv:22.0)
Gecko/20100101 Firefox/22.0", "12": ["{\"height\":678,\"width\":1207}"], "13": [0]}
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;...I ended up with what was externally represented as a list of 18 floats (I wanted to store them as integers instead, but Pandas needs floats to be able to represent NaNs, which was important, given that many of the values were sparse).&lt;/p&gt;

&lt;h1&gt;The problem with shared state&lt;/h1&gt;

&lt;p&gt;This worked well before I tried to do any kind of speedup (apart from the fact that the script took 12 hours to run on the largest file). However, I wanted to be able to iterate a little bit quicker, and explore ways of speeding up the code. Because I needed to keep the memoized data consistent, I couldn't just fork off a bunch of processes that each would process their own data -- they would have all ended up with their own individual memo tables, where &lt;code&gt;1&lt;/code&gt; corresponds to &lt;code&gt;hello&lt;/code&gt; in one process, and &lt;code&gt;what's up&lt;/code&gt; in another. To address this, I moved the memoization code out of the function that processed individual lines, and applied it to each Pandas column in a final step before appending the processed chunks to the HDF5 file.&lt;/p&gt;

&lt;p&gt;The function to process each line should now be a pure function -- that is, a function that always returns the same output given the same input, not reliant on any external state, etc. -- and thus ripe for parallization. I tried using &lt;code&gt;pool.map&lt;/code&gt; from the &lt;a href="http://docs.python.org/3.3/library/multiprocessing.html"&gt;multiprocessing&lt;/a&gt; library to let the computer automatically figure out how many external processes to run, etc. Here I ran into a problem of grain size... It takes a bit of time to fire up processes, transfer state, coordinate across threads etc, so if you assign each process too little work, the gain from having work done in parallel is more than lost. This was the result -- my naive attempt led to a slowdown. (Some libraries try to tackle this automatically and figure out the correct grain size, but pool.map didn't seem to be able to).&lt;/p&gt;

&lt;p&gt;My second attempt was to do the chunking myself -- I read in 100,000 lines, chunked it into chunks of 20,000 lines, and then mapped over this list of lists. This worked better, but I only saw a small speedup (maybe 50%). Partly because of the extra time spent chunking lines, and partly because a big part of the program was still spent converting the list of dicts to a Pandas DataFrame, memoizing the columns, and writing to file (I could see this very clearly by monitoring &lt;code&gt;top&lt;/code&gt; -- first you see five processes maxing out their respective cores, and then suddenly they drop away, and there is only one process on one core working).&lt;/p&gt;

&lt;h1&gt;Map-reduce&lt;/h1&gt;

&lt;p&gt;A very common paradigm for cluster computing (with multiple computers) is called &lt;a href="https://en.wikipedia.org/wiki/MapReduce"&gt;map-reduce&lt;/a&gt;. Basically, you split the data into chunks, run parallel operations on each chunk, and then somehow "reduce" the result back together. An example could be getting number of page views per URL for an extremely large clicklog, have each serve count up clicks per URL for a piece of the log file, and then have a script that combines the numbers in the end.&lt;/p&gt;

&lt;p&gt;Something similar can be done on a single computer, and for the second phase of the analysis, this is what I ended up with. The second phase, after we got the nice HDF5 file, is to do some processing for each user. Instead of the fact that a user watched video 39, we want to know if he/she watched the video for the first time, or if it was an immediate review, to be able to compare patterns across students, and even across courses.&lt;/p&gt;

&lt;p&gt;This means that we cannot treat each &lt;code&gt;student event&lt;/code&gt; individually -- we need to know which videos the students has already seen, etc. However, we can treat each &lt;code&gt;student&lt;/code&gt; individually -- how one student's actions are tagged, are not at all reliant on other students' actions.&lt;/p&gt;

&lt;h1&gt;Split, process, collect&lt;/h1&gt;

&lt;p&gt;One very nice thing about HDF5 is that once you've indexed the table (which is quite fast), you can read in data based on simple queries. Thus, once we've opened the HDF file with &lt;code&gt;store = pd.HDFStore("hdfstore.h5")&lt;/code&gt;, we don't have to read the entire table into memory with &lt;code&gt;db = store['db']&lt;/code&gt;, which might take 6GB of memory (with a big table), but we can instead run&lt;/p&gt;

&lt;pre&gt;&lt;code class="python"&gt;user = store.select('db', pd.Term('username = 3434')
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;(in our case, we also select on timestamp, because we only want events that happened while the course was active). Add to this that pytables is thread-safe for reading (ie. multiple processes can read from the same file, without problems), but not for writing (it's not like a mysql database that can have multiple processes writing at the same time).&lt;/p&gt;

&lt;p&gt;So instead of having one script that tries to parallelize subtasks, what I ended up doing was to run eight Python scripts concurrently. I simply indicate as a command parameter which usernumber to start with, and how far to jump each time. Thus, the first Python process starts with user 1, then does user 9, user 17 etc, until there are no more users. The second process starts with user number 2, then 10, etc. This way all users get processes, without the need for any coordination. Initially I started the processes from a shell script, but then I put it into Python:&lt;/p&gt;

&lt;pre&gt;&lt;code class="python"&gt;print("Spawning %d processing scripts" % numprocesses)
procs = []
for proc in range(1, numprocesses+1):
    range_start = proc
    range_jump = numprocesses+1
    procs.append(Popen([python_exec, "videos_reduce.py", hdffile, str(range_start), \
        str(range_jump), tmpdir, cutoff]))
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;I can then wait for all processes to finish, and then do any cleanup work that is necessary.&lt;/p&gt;

&lt;pre&gt;&lt;code class="python"&gt;def all_procs_finished(procs):
    for proc in procs:
        if proc.poll() == None:
            return False
        return True

while not all_procs_finished(procs):
    time.sleep(5)
&lt;/code&gt;&lt;/pre&gt;

&lt;h1&gt;Reduce&lt;/h1&gt;

&lt;p&gt;Initially, I wanted to capture the output in a new HDF5 file, so I had the individual scripts dump their output pickled into temporary files (named using &lt;a href="http://docs.python.org/3.3/library/uuid.html"&gt;uuid&lt;/a&gt;) in a temporary directory, which would then be read by the main script, and appended to the HDF5 file (thus ensuring that only one thread wrote to the HDF5, to avoid data corruption). Later, we decided we wanted to output a text file for processing in R, which made it even easier. The individual thread simply spit out a bunch of text file chunks into a temporary directory, and then we combine them with &lt;code&gt;cat * &amp;gt; outfile&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This approach worked swimmingly. Since we are just running the same script eight times in parallel, it's easy to debug by running the script only once (I would run it with &lt;code&gt;1 10000&lt;/code&gt; to start with user 1 and jump 10000 ahead, to 10001, 20001 etc until the file had been exhausted). Because all the hard work was done in the individual threads, I got very close to 100% utilization on all eight cores for the duration of the run (which is a beauty to see in &lt;code&gt;htop&lt;/code&gt;), and even when running a separate process that collected the output files and wrote them to HDF5, that process used so little CPU as to be negligible.&lt;/p&gt;

&lt;p&gt;&lt;img src="/blog/images/2014-03-10-parsing-massive-clicklogs-01.png" alt=""&gt;&lt;/p&gt;

&lt;h1&gt;But what about that shared state?&lt;/h1&gt;

&lt;p&gt;At this point, I needed to go back to the initial phase to redo a bunch of the parsing, because I had received new files. Inspired by my experience with the second phase, I thought about how I could improve the first-phase scripts. The tricky problem was how to coordinate state. I did a bunch of googling to look at "coordinated / parallel memoization libraries for Python", and somehow came upon people using &lt;a href="http://redis.io/"&gt;Redis&lt;/a&gt; for that purpose.&lt;/p&gt;

&lt;p&gt;I had heard very vaguely about Redis, but never known much about it. It turned out to be perfect for the job through. It's basically a very fast distributed key-value store (although it also has other value types, like lists and hashes), with an API that only takes a few hours to learn. I installed a Redis server, and rewrote my memoization function to use Redis instead of the Pandas Series, with only very minimal changes.&lt;/p&gt;

&lt;p&gt;Here is a shortened version of the memoization class:&lt;/p&gt;

&lt;pre&gt;&lt;code class="python"&gt;class Memo(object):
    def __init__(self, name, prefix):
        self.r = redis.StrictRedis(host='localhost', decode_responses=True)
        self.name = name
        self.prefix = prefix
        self.counter = '%s:%s:cnt' % (prefix, name)
        self.store = '%s:%s:store' % (prefix, name)
        if self.r.zcard(self.store) == 0:
            self.add_object('None')

        def add_object(self, obj):
            c = self.r.incr(self.counter)
            self.r.zadd(self.store, c, obj)
            return c

        @lru_cache(maxsize=256)
        def get(self, s):
            if s is None or pd.isnull(s):
                return -1

                existing_score = self.r.zscore(self.store, s)
                if existing_score:
                    return existing_score
                else:
                    c = self.add_object(s)
                    return c
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;You create a new Memo class with a name, and then just call it like described above. &lt;code&gt;lru_cache&lt;/code&gt;, from &lt;a href="http://docs.python.org/3.3/library/functools.html"&gt;functools&lt;/a&gt;, is really neat - it remembers the last n (maxsize) arguments to the function, and their return values, and if you call the function with an argument in the list, it gives you the response, without having to run the function. This is safe, given that once a value is assigned in Redis, it will never change. Having this meant a big speedup (especially since it's very common for adjacent lines to have similar values, for example a chunk of lines all from the same user, etc), however when I tried to set &lt;code&gt;maxsize=None&lt;/code&gt;, ie. to remember all the values, it became much slower. Apparently Redis is much faster at looking up a hash-value among 10,000 than Python, despite the overhead of calling out.&lt;/p&gt;

&lt;p&gt;So the final version of the script does several thing. It begins by running the shell command &lt;code&gt;split -n&lt;/code&gt;, which splits the log file into chunks based on number of lines (currently 40,000 lines seems to be a good size). It does not wait until cut is finished, but begins immediately launching the 8 processes, which begin reading the next available file in the temp dir, and processing it.&lt;/p&gt;

&lt;p&gt;The eight workers pick up a temp file "chew on it", and spit out the result as a pickled array of dicts in another temp dir. The master script, having launched all the processes, begins to cycle, waiting for files to appear, reads them in, and appends them to the HDF5 file. The worker processes continue until they've gotten the signal that cut has finished (through Redis), and there are no more files. The master script continues until all worker processes have finished, and there are no more "chewed" files. It then repacks the HDF5 file (to add indices), and adds the memoized data from Redis.&lt;/p&gt;

&lt;p&gt;Using this approach, I am now able to process a 20GB clicklog file in 12 minutes! (Resulting in a 3.5 GB HDF5 file with something like 14 million events).&lt;/p&gt;
</content>
  </entry>
</feed>
