{"id":562,"date":"2010-10-06T22:21:06","date_gmt":"2010-10-07T03:21:06","guid":{"rendered":"http:\/\/yuguangzhang.com\/blog\/?p=562"},"modified":"2010-10-06T22:21:06","modified_gmt":"2010-10-07T03:21:06","slug":"tokenizing-prerequisites","status":"publish","type":"post","link":"http:\/\/yuguangzhang.com\/blog\/tokenizing-prerequisites\/","title":{"rendered":"Tokenizing Prerequisites"},"content":{"rendered":"<p>I&#8217;m taking a bite by bite approach to interpreting the course prerequisites from <a href=\"http:\/\/yuguangzhang.com\/blog\/uw-course-calendar-scraper\/\">the scraper<\/a> to make its storage in the database possible. Meanwhile, I already finished the part of it that <a href=\"http:\/\/yuguangzhang.com\/blog\/python-iteration-recursion\/\">retrieves the data recursively<\/a>. The approach I decide to take is the same as building a compiler.<\/p>\n<p>I previously favored a quick approach such as the interpreter pattern.<\/p>\n<p><a href=\"http:\/\/yuguangzhang.com\/blog\/wp-content\/uploads\/2010\/10\/design.png\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-medium wp-image-564\" title=\"design\" src=\"http:\/\/yuguangzhang.com\/blog\/wp-content\/uploads\/2010\/10\/design-300x165.png\" alt=\"design\" width=\"300\" height=\"165\" srcset=\"http:\/\/yuguangzhang.com\/blog\/wp-content\/uploads\/2010\/10\/design-300x165.png 300w, http:\/\/yuguangzhang.com\/blog\/wp-content\/uploads\/2010\/10\/design-1024x565.png 1024w, http:\/\/yuguangzhang.com\/blog\/wp-content\/uploads\/2010\/10\/design.png 1064w\" sizes=\"auto, (max-width: 300px) 100vw, 300px\" \/><\/a><\/p>\n<p>However, commas in sentences have different meanings depending on the context.<\/p>\n<pre>(BIOL 140, BIOL 208 or 330) and (BIOL 308 or 330)<\/pre>\n<pre>(CS 240 or SE 240), CS 246, ECE 222<\/pre>\n<p>The interpreter pattern assumes a symbol only has one meaning. Thanks again to a course I took, I was already familiar with the problem and the solution. Recognizing and defining the problem was the hard part. A quick overview of the steps involved from start to finish:<\/p>\n<p><a href=\"http:\/\/yuguangzhang.com\/blog\/wp-content\/uploads\/2010\/10\/coursetree.png\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter\" title=\"Coursetree overview\" src=\"http:\/\/yuguangzhang.com\/blog\/wp-content\/uploads\/2010\/10\/coursetree_thumb.png\" alt=\"\" width=\"500\" height=\"250\" \/><\/a><\/p>\n<p>First step in the compiling process is to tokenize input. \u00a0That takes care of blobs such as &#8220;Level at least 3A; Not open to General Mathematics students.&#8221;. Totally useless in terms of prerequisites. So it should not show up in the token string.<\/p>\n<p>Building the lexer was as simple as specifying the token list, writing regular expressions for them, and setting ignored characters.<br \/>\n[cc lang=&#8221;python&#8221;]<br \/>\nimport ply.lex as lex<br \/>\nt_DEPT = r'[A-Z]{2,5}&#8217;<br \/>\nt_NUM = r&#8217;\\d{3}&#8217;<br \/>\nt_OR = r&#8217;or&#8217;<br \/>\nt_COMMA = r&#8217;,&#8217;<br \/>\nt_SEMI = r&#8217;;&#8217;<br \/>\ntokens = (<br \/>\n\t&#8216;DEPT&#8217;,<br \/>\n\t&#8216;NUM&#8217;,<br \/>\n\t&#8216;OR&#8217;,<br \/>\n\t&#8216;COMMA&#8217;,<br \/>\n\t&#8216;SEMI&#8217;,<br \/>\n\t&#8216;AND&#8217;,<br \/>\n\t&#8216;LPARENS&#8217;,<br \/>\n\t&#8216;RPARENS&#8217;,<br \/>\n)<br \/>\nt_AND = r&#8217;and&#8217;<br \/>\nt_LPARENS = r&#8217;\\(&#8216;<br \/>\nt_RPARENS = r&#8217;\\)&#8217;<br \/>\nt_ignore  = &#8216; \\t&#8217;<br \/>\ndef t_error(t):<br \/>\n\t    print &#8220;Illegal character &#8216;%s'&#8221; % t.value[0]<br \/>\n\t    t.lexer.skip(1)<br \/>\nlexer = lex.lex()<br \/>\ndata = &#8216; (CS 240 or SE 240), CS 246, ECE 222&#8217;<br \/>\nlexer.input(data)<br \/>\nwhile True:<br \/>\n    tok = lexer.token()<br \/>\n    if not tok: break      # No more input<br \/>\n    print tok<br \/>\n[\/cc]<br \/>\nRunning it gives the output:<\/p>\n<pre>LexToken(LPARENS,'(',1,1)\nLexToken(DEPT,'CS',1,2)\nLexToken(NUM,'240',1,5)\nLexToken(OR,'or',1,9)\nLexToken(DEPT,'SE',1,12)\nLexToken(NUM,'240',1,15)\nLexToken(RPARENS,')',1,18)\nLexToken(COMMA,',',1,19)\nLexToken(DEPT,'CS',1,21)\nLexToken(NUM,'246',1,24)\nLexToken(COMMA,',',1,27)\nLexToken(DEPT,'ECE',1,29)\nLexToken(NUM,'222',1,33)<\/pre>\n","protected":false},"excerpt":{"rendered":"<p>I&#8217;m taking a bite by bite approach to interpreting the course prerequisites from the scraper to make its storage in the database possible. Meanwhile, I already finished the part of it that retrieves the data recursively. The approach I decide to take is the same as building a compiler. I previously favored a quick approach [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_import_markdown_pro_load_document_selector":0,"_import_markdown_pro_submit_text_textarea":"","footnotes":""},"categories":[29],"tags":[53,22],"class_list":["post-562","post","type-post","status-publish","format-standard","hentry","category-coursetree","tag-ply","tag-python"],"aioseo_notices":[],"_links":{"self":[{"href":"http:\/\/yuguangzhang.com\/blog\/wp-json\/wp\/v2\/posts\/562","targetHints":{"allow":["GET"]}}],"collection":[{"href":"http:\/\/yuguangzhang.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"http:\/\/yuguangzhang.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"http:\/\/yuguangzhang.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"http:\/\/yuguangzhang.com\/blog\/wp-json\/wp\/v2\/comments?post=562"}],"version-history":[{"count":0,"href":"http:\/\/yuguangzhang.com\/blog\/wp-json\/wp\/v2\/posts\/562\/revisions"}],"wp:attachment":[{"href":"http:\/\/yuguangzhang.com\/blog\/wp-json\/wp\/v2\/media?parent=562"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"http:\/\/yuguangzhang.com\/blog\/wp-json\/wp\/v2\/categories?post=562"},{"taxonomy":"post_tag","embeddable":true,"href":"http:\/\/yuguangzhang.com\/blog\/wp-json\/wp\/v2\/tags?post=562"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}