<turbo-stream action="append" target="posts_list"><template>    <div class="postbit" id="137808" data-post-id="137808">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="mischov" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/mischov/120/2489_2.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  mischov
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<h3><a name="p-137808-release-v0120httpsgithubcommischovmeeseeksreleasestagv0120-1" class="anchor" href="#p-137808-release-v0120httpsgithubcommischovmeeseeksreleasestagv0120-1" aria-label="Heading link" rel="nofollow"></a>Release <a href="https://github.com/mischov/meeseeks/releases/tag/v0.12.0" rel="noopener nofollow ugc">v0.12.0</a></h3>
<p>This release fixes <code>Meeseeks.html/1</code> so that when encoding attribute values double quotes are always used and <code>&amp;</code> and <code>"</code> are escaped as character entities and when encoding text <code>&lt;</code>, <code>&gt;</code>, and <code>&amp;</code> are escaped as character entities.</p>
<p><strong>This should be considered a breaking change</strong> since the output for <code>Meeseeks.html/1</code> may be slightly different, but means for instance that round tripping</p>
<pre data-code-wrap="elixir"><code class="lang-elixir">"&lt;span&gt;&amp;lt;script&amp;gt;Hello&amp;lt;/script&amp;gt;&lt;/span&gt;"
</code></pre>
<p>through <code>Meeseeks.parse</code> then <code>Meeseeks.html</code> produces</p>
<pre data-code-wrap="elixir"><code class="lang-elixir">"&lt;span&gt;&amp;lt;script&amp;gt;Hello&amp;lt;/script&amp;gt;&lt;/span&gt;"
</code></pre>
<p>instead of</p>
<pre data-code-wrap="elixir"><code class="lang-elixir">"&lt;span&gt;&lt;script&gt;Hello&lt;/script&gt;&lt;/span&gt;"
</code></pre>
<p>A big thanks to <a class="mention" href="/u/ericlathrop" rel="nofollow">@ericlathrop</a> for bringing the issues to my attention and proposing a solution.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="137808" data-batch-url="/posts/batch_likers">
                        5
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/meeseeks-a-library-for-extracting-data-from-html-and-xml-with-css-or-xpath-selectors/4315/62">Post #61</a>
	                </div>
	            </div>
              <div id="likers-container-137808" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="137808"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #61"></div>
  </section>
</div>
    <div class="postbit" id="142694" data-post-id="142694">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="mischov" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/mischov/120/2489_2.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  mischov
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<h3><a name="p-142694-release-v0130httpsgithubcommischovmeeseeksreleasestagv0130-1" class="anchor" href="#p-142694-release-v0130httpsgithubcommischovmeeseeksreleasestagv0130-1" aria-label="Heading link" rel="nofollow"></a>Release <a href="https://github.com/mischov/meeseeks/releases/tag/v0.13.0" rel="noopener nofollow ugc">v0.13.0</a></h3>
<p>This release adds support for Erlang/OTP 22 (and Elixir 1.9, though that was working fine before), and removes support for Elixir 1.4, Elixir 1.5, and OTP 19. I had only planned on removing support for Elixir 1.4 and OTP 19, but a recent change in Rustler made 1.6 its minimum version. Sorry for any inconvenience that this might cause.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="142694" data-batch-url="/posts/batch_likers">
                        4
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/meeseeks-a-library-for-extracting-data-from-html-and-xml-with-css-or-xpath-selectors/4315/63">Post #62</a>
	                </div>
	            </div>
              <div id="likers-container-142694" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="142694"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #62"></div>
  </section>
</div>
    <div class="postbit" id="142830" data-post-id="142830">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="mischov" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/mischov/120/2489_2.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  mischov
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<h3><a name="p-142830-release-v0131httpsgithubcommischovmeeseeksreleasestagv0131-1" class="anchor" href="#p-142830-release-v0131httpsgithubcommischovmeeseeksreleasestagv0131-1" aria-label="Heading link" rel="nofollow"></a>Release <a href="https://github.com/mischov/meeseeks/releases/tag/v0.13.1" rel="noopener nofollow ugc">v0.13.1</a></h3>
<p>Since the minimum supported version of Erlang/OTP is now 20 it was possible to switch the NIF from working asynchronously to using a dirty scheduler, which simplified the NIF’s implementation and may speed things up (I’ll get an updated version of the Meeseeks vs. Floki bench out soon).</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="142830" data-batch-url="/posts/batch_likers">
                        3
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/meeseeks-a-library-for-extracting-data-from-html-and-xml-with-css-or-xpath-selectors/4315/64">Post #63</a>
	                </div>
	            </div>
              <div id="likers-container-142830" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="142830"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #63"></div>
  </section>
</div>
    <div class="postbit" id="142860" data-post-id="142860">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="OvermindDL1" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/OvermindDL1/120/2677_2.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  OvermindDL1
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<aside class="quote no-group" data-username="mischov" data-post="64" data-topic="4315">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/mischov/48/2489_2.png" class="avatar"> mischov:</div>
<blockquote>
<p>it was possible to switch the NIF from working asynchronously to using a dirty scheduler</p>
</blockquote>
</aside>
<p>Ooo very cool!</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="142860" data-batch-url="/posts/batch_likers">
                        1
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/meeseeks-a-library-for-extracting-data-from-html-and-xml-with-css-or-xpath-selectors/4315/65">Post #64</a>
	                </div>
	            </div>
              <div id="likers-container-142860" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="142860"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #64"></div>
  </section>
</div>
    <div class="postbit" id="145230" data-post-id="145230">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="mischov" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/mischov/120/2489_2.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  mischov
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<h3><a name="p-145230-release-v0140httpsgithubcommischovmeeseeksreleasestagv0140-1" class="anchor" href="#p-145230-release-v0140httpsgithubcommischovmeeseeksreleasestagv0140-1" aria-label="Heading link" rel="nofollow"></a>Release <a href="https://github.com/mischov/meeseeks/releases/tag/v0.14.0" rel="noopener nofollow ugc">v0.14.0</a></h3>
<p>This release changes how extractors are implemented and makes various other extractor-related improvements.</p>
<h4><a name="p-145230-reimplementation-2" class="anchor" href="#p-145230-reimplementation-2" aria-label="Heading link" rel="nofollow"></a>Reimplementation</h4>
<p>Previously extractors were implemented as callbacks in the private <code>Document.Node</code> behaviour. This didn’t play to the strengths of behaviours because document nodes are a closed set of structs and the extensibility of having extractors as callbacks didn’t provide any concrete benefits. It did, however, set an example that obscured the fact that you can easily write your own extractor functions, and split up extractor implementations over the various node types making them harder to understand.</p>
<p>I refactored this, removing the <code>Document.Node</code> behaviour and moving that functionality to modules under <code>Meeseeks.Extractor</code>.</p>
<h4><a name="p-145230-fixes-and-improvements-3" class="anchor" href="#p-145230-fixes-and-improvements-3" aria-label="Heading link" rel="nofollow"></a>Fixes and Improvements</h4>
<p>I also made a number of improvements to extractors, such as using iodata in string-building extractors instead of string concatenation, improving the performance of whitespace collapsing, and making whitespace collapsing optional in those extractors that do it (<code>data</code>, <code>own_text</code>, <code>text</code>).</p>
<p>The release also contains a couple extractor-related fixes, one from <a href="https://github.com/pclewis" rel="noopener nofollow ugc">pclewis</a> that removes the unnecessary spaces that comments were being wrapped in when encoding them to HTML, and another that limits the adding of a space between sibling nodes during text extraction to when the preceding sibling did not end in whitespace[1].</p>
<p>A big “thank you” to Philip for his fix!</p>
<h4><a name="p-145230-compatibility-4" class="anchor" href="#p-145230-compatibility-4" aria-label="Heading link" rel="nofollow"></a>Compatibility</h4>
<p>By this point you may be concerned about how this release will break your existing use of Meeseeks. Good news- as long as you use the public API the changes to how extractors are implemented are completely backwards compatible. The two fixes mentioned above do, however, mean that using <code>data</code>, <code>html</code>, <code>own_text</code>, or <code>text</code> may yield slightly different results than before.</p>
<p>If you are for some reason using the private  <code>Document.Node</code>  callbacks directly on document nodes, sorry, that behaviour is gone and the callbacks are no longer implemented on nodes. If you are using the private  <code>Document.Node</code>  helper functions that called into those callbacks you will be happy to learn those still work fine.</p>
<h4><a name="p-145230-is-it-fast-tho-5" class="anchor" href="#p-145230-is-it-fast-tho-5" aria-label="Heading link" rel="nofollow"></a>Is it fast tho?</h4>
<p>I want to get an updated version of the Meeseeks vs Floki bench out, but there are some technical difficulties there because <code>html5ever_elixir</code> won’t compile on Erlang/OTP 22 so I have nothing to compare against. I am considering comparing against the <code>:mochiweb_html</code> parser for a round since that’s what a lot of people are being forced to use anyway, but I’m still a bit reluctant because that’s more of an apples to oranges comparison.</p>
<hr>
<p>[1] Efficiently finding out if a binary ends in whitespace is hard. I am really glad José had to figure it out for <code>String.trim</code> and I was able to adapt his solution.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="145230" data-batch-url="/posts/batch_likers">
                        4
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/meeseeks-a-library-for-extracting-data-from-html-and-xml-with-css-or-xpath-selectors/4315/66">Post #65</a>
	                </div>
	            </div>
              <div id="likers-container-145230" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="145230"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #65"></div>
  </section>
</div>
    <div class="postbit" id="163331" data-post-id="163331">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="mischov" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/mischov/120/2489_2.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  mischov
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<h3><a name="p-163331-release-v0150httpsgithubcommischovmeeseeksreleasestagv0150-1" class="anchor" href="#p-163331-release-v0150httpsgithubcommischovmeeseeksreleasestagv0150-1" aria-label="Heading link" rel="nofollow"></a>Release <a href="https://github.com/mischov/meeseeks/releases/tag/v0.15.0" rel="noopener nofollow ugc">v0.15.0</a></h3>
<p>This release adds support for Elixir 1.10 and makes a couple correctness related improvements.</p>
<h4><a name="p-163331-safer-tuple-tree-parser-2" class="anchor" href="#p-163331-safer-tuple-tree-parser-2" aria-label="Heading link" rel="nofollow"></a>Safer tuple tree parser</h4>
<p>The first of these improvements is that the tuple tree parser is now much more strict about not parsing invalid input. Thanks to <a href="https://github.com/pclewis" rel="noopener nofollow ugc">pclewis</a> for pointing out an issue that led to this work.</p>
<h4><a name="p-163331-no-xpath-attribute-steps-outside-of-predicates-3" class="anchor" href="#p-163331-no-xpath-attribute-steps-outside-of-predicates-3" aria-label="Heading link" rel="nofollow"></a>No XPath attribute steps outside of predicates</h4>
<p>The second of these improvements is that XPath attribute steps outside of predicates are now prohibited (rather than just broken).</p>
<p>For example, <code>xpath("\\p[@class]")</code> which returns elements with class attributes is allowed, but <code>xpath("\\p\@class")</code> which would return the class attributes themselves is prohibited. If you do need to extract a selected element’s attribute use the <code>attr</code> extractor.</p>
<pre data-code-wrap="elixir"><code class="lang-elixir">Meeseeks.all(doc, xpath("//p[@class]")) |&gt; Enum.map(&amp;Meeseeks.attr(&amp;1, "class"))
</code></pre>
<p>Thanks to <a class="mention" href="/u/oldhammade" rel="nofollow">@OldhamMade</a> for reporting the issue that prompted this work.</p>
<h4><a name="p-163331-other-changes-4" class="anchor" href="#p-163331-other-changes-4" aria-label="Heading link" rel="nofollow"></a>Other changes</h4>
<p>There are also some minor improvements to the project documentation, and <a href="https://github.com/mischov/meeseeks/blob/v0.15.0/CONTRIBUTING.md" rel="noopener nofollow ugc">contribution guidelines</a> have been added.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="163331" data-batch-url="/posts/batch_likers">
                        4
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/meeseeks-a-library-for-extracting-data-from-html-and-xml-with-css-or-xpath-selectors/4315/67">Post #66</a>
	                </div>
	            </div>
              <div id="likers-container-163331" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="163331"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #66"></div>
  </section>
</div>
    <div class="postbit" id="173268" data-post-id="173268">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="mischov" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/mischov/120/2489_2.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  mischov
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<p>Since <code>html5ever_elixir</code> recently released a new version that works on Erlang/OTP 22 I have finally been able to release a <a href="https://github.com/mischov/meeseeks_floki_bench" rel="noopener nofollow ugc">Meeseeks vs. Floki Benchmark</a> update.</p>
<p>Here’s an excerpt from the Trending JS scenario, but go ahead and check out the whole benchmark for a more complete description.</p>
<pre data-code-wrap="bash"><code class="lang-bash">Name                     ips        average  deviation         median         99th %
Meeseeks CSS           23.22       43.07 ms     ±2.73%       42.79 ms       47.22 ms
Meeseeks XPath         19.47       51.35 ms     ±4.03%       50.77 ms       60.82 ms
Floki CSS              14.01       71.39 ms     ±3.85%       71.31 ms       83.36 ms

Comparison: 
Meeseeks CSS           23.22
Meeseeks XPath         19.47 - 1.19x slower +8.28 ms
Floki CSS              14.01 - 1.66x slower +28.32 ms

Memory usage statistics:

Name              Memory usage
Meeseeks CSS           3.66 MB
Meeseeks XPath         6.57 MB - 1.80x memory usage +2.91 MB
Floki CSS             22.23 MB - 6.08x memory usage +18.57 MB
</code></pre> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="173268" data-batch-url="/posts/batch_likers">
                        3
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/meeseeks-a-library-for-extracting-data-from-html-and-xml-with-css-or-xpath-selectors/4315/68">Post #67</a>
	                </div>
	            </div>
              <div id="likers-container-173268" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="173268"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #67"></div>
  </section>
</div>
    <div class="postbit" id="173390" data-post-id="173390">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="dimitarvp" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/dimitarvp/120/38664_2.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  dimitarvp
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<p>One problem I lately had with Meeseeks was that it complained about an invalid encoding while Floki just parsed the document in UTF-8. I can try and find a few such defective HTML files if you like. But in such situations I appreciate my tool’s ability to cope with the problem instead of erroring out.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="173390" data-batch-url="/posts/batch_likers">
                        1
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/meeseeks-a-library-for-extracting-data-from-html-and-xml-with-css-or-xpath-selectors/4315/69">Post #68</a>
	                </div>
	            </div>
              <div id="likers-container-173390" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="173390"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #68"></div>
  </section>
</div>
    <div class="postbit" id="173402" data-post-id="173402">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="mischov" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/mischov/120/2489_2.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  mischov
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<p>Please do create an issue if you believe you’ve found an error, yes.</p>
<p>If it’s an issue with not parsing a content type <code>text/html; charset=ISO-8859-1</code> or some other charset as UTF-8, yes, that’s a known issue for Meeseeks.</p>
<aside class="onebox githubissue" data-onebox-src="https://github.com/mischov/meeseeks/issues/50">
  <header class="source">

      <a href="https://github.com/mischov/meeseeks/issues/50" target="_blank" rel="noopener nofollow ugc">github.com/mischov/meeseeks</a>
  </header>

  <article class="onebox-body">
    <div class="github-row">
  <div class="github-icon-container" title="Issue" data-github-private-repo="false">
	  <svg width="60" height="60" class="github-icon" viewBox="0 0 14 16" aria-hidden="true"><path fill-rule="evenodd" d="M7 2.3c3.14 0 5.7 2.56 5.7 5.7s-2.56 5.7-5.7 5.7A5.71 5.71 0 0 1 1.3 8c0-3.14 2.56-5.7 5.7-5.7zM7 1C3.14 1 0 4.14 0 8s3.14 7 7 7 7-3.14 7-7-3.14-7-7-7zm1 3H6v5h2V4zm0 6H6v2h2v-2z"></path></svg>
  </div>

  <div class="github-info-container">
    <h4>
      <a href="https://github.com/mischov/meeseeks/issues/50" target="_blank" rel="noopener nofollow ugc"> Panicked with Utf8Error when trying to parse Google result </a>
    </h4>

    <div class="github-info">
      <div class="date">
        opened <span class="discourse-local-date" data-format="ll" data-date="2018-12-25" data-time="17:18:20" data-timezone="UTC">05:18PM - 25 Dec 18 UTC</span>
      </div>

        <div class="date">
          closed <span class="discourse-local-date" data-format="ll" data-date="2018-12-27" data-time="15:01:08" data-timezone="UTC">03:01PM - 27 Dec 18 UTC</span>
        </div>

      <div class="user">
        <a href="https://github.com/raooll" target="_blank" rel="noopener nofollow ugc">
          <img alt="" src="https://avatars.githubusercontent.com/u/3482371?v=4" class="onebox-avatar-inline" width="20" height="20">
          raooll
        </a>
      </div>
    </div>

    <div class="labels">
    </div>
  </div>
</div>

  <div class="github-row">
    <p class="github-body-container">iex(27)&gt; html = HTTPoison.get!("https://www.google.com/search?q=wild%20love&amp;star<span class="show-more-container"><a href="" rel="noopener nofollow" class="show-more">…</a></span><span class="excerpt hidden">t=0&amp;gws_rd=cr&amp;gbv=1").body 

iex(22)&gt; s = Meeseeks.parse(html, :xml)
thread '&lt;unnamed&gt;' panicked at 'called `Result::unwrap()` on an `Err` value: Utf8Error { valid_up_to: 35017, error_len: Some(1) }', libcore/result.rs:1009:5
                                                                                                                                                            {:error,
 "called `Result::unwrap()` on an `Err` value: Utf8Error { valid_up_to: 35017, error_len: Some(1) }"}

Getting the above error while parsing  a google search page.</span></p>
  </div>

  </article>

  <div class="onebox-metadata">
    
    
  </div>

  <div style="clear: both"></div>
</aside>

<p>And for Floki</p>
<aside class="onebox githubissue" data-onebox-src="https://github.com/philss/floki/issues/221">
  <header class="source">

      <a href="https://github.com/philss/floki/issues/221" target="_blank" rel="noopener nofollow ugc">github.com/philss/floki</a>
  </header>

  <article class="onebox-body">
    <div class="github-row">
  <div class="github-icon-container" title="Issue" data-github-private-repo="false">
	  <svg width="60" height="60" class="github-icon" viewBox="0 0 14 16" aria-hidden="true"><path fill-rule="evenodd" d="M7 2.3c3.14 0 5.7 2.56 5.7 5.7s-2.56 5.7-5.7 5.7A5.71 5.71 0 0 1 1.3 8c0-3.14 2.56-5.7 5.7-5.7zM7 1C3.14 1 0 4.14 0 8s3.14 7 7 7 7-3.14 7-7-3.14-7-7-7zm1 3H6v5h2V4zm0 6H6v2h2v-2z"></path></svg>
  </div>

  <div class="github-info-container">
    <h4>
      <a href="https://github.com/philss/floki/issues/221" target="_blank" rel="noopener nofollow ugc">Some attributes are binary</a>
    </h4>

    <div class="github-info">
      <div class="date">
        opened <span class="discourse-local-date" data-format="ll" data-date="2019-09-17" data-time="19:35:47" data-timezone="UTC">07:35PM - 17 Sep 19 UTC</span>
      </div>

        <div class="date">
          closed <span class="discourse-local-date" data-format="ll" data-date="2019-09-21" data-time="17:05:29" data-timezone="UTC">05:05PM - 21 Sep 19 UTC</span>
        </div>

      <div class="user">
        <a href="https://github.com/bitboxer" target="_blank" rel="noopener nofollow ugc">
          <img alt="" src="https://avatars.githubusercontent.com/u/56195?v=4" class="onebox-avatar-inline" width="20" height="20">
          bitboxer
        </a>
      </div>
    </div>

    <div class="labels">
    </div>
  </div>
</div>

  <div class="github-row">
    <p class="github-body-container">I am writing a little crawler to fetch data from other pages and have a weird pr<span class="show-more-container"><a href="" rel="noopener nofollow" class="show-more">…</a></span><span class="excerpt hidden">oblem with [this page](https://www.zvab.com/servlet/BookDetailsPL?bi=30131047830&amp;searchurl=an%3Dk%25E4stner%26hl%3Don%26sortby%3D20%26tn%3D4%2Bb%25E4nde&amp;cm_sp=snippet-_-srp1-_-title2):

This part:

```html
&lt;meta itemprop="name" content="Kästner für Erwachsene / 4 Bände: ..." /&gt;
```

Is represented as binary instead of a string inside of floki:

```
{"meta",
                      [
                        {"itemprop", "name"},
                        {"content",
                         &lt;&lt;75, 228, 115, 116, 110, 101, 114, 32, ...&gt;&gt;}
                      ], []},
```

My current guess is an encoding issue there. Is there a away to fix this inside of floki or should I try to get them to fix it?</span></p>
  </div>

  </article>

  <div class="onebox-metadata">
    
    
  </div>

  <div style="clear: both"></div>
</aside>

<p>It’s just that the <code>mochiweb_html</code> parser will try to treat the other encoding as UTF-8 and give you back gibberish or incorrect results at times while the <code>html5ever</code> parser that Meeseeks uses by default will just not work.</p>
<p>Erroring is the correct response, imo. If there’s a problem it’s better to know about it early and handle it correctly (as suggested in the Floki issue, by converting from whatever charset to UTF-8, then parsing) than to incorrectly parse and potentially return wrong answers or gibberish and not provide any context as to why.</p>
<p>At one point html5ever attempted to provide a mechanism for parsing from other encodings (<code>from_bytes</code>) but that was removed. In practice it’s difficult- I believe to do content sniffing correctly you need to provide information from the HTTP request as well, so it’s not just a matter of parsing HTML any more and consequently not necessarily appropriate for a HTML parser to handle.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="173402" data-batch-url="/posts/batch_likers">
                        0
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/meeseeks-a-library-for-extracting-data-from-html-and-xml-with-css-or-xpath-selectors/4315/70">Post #69</a>
	                </div>
	            </div>
              <div id="likers-container-173402" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="173402"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #69"></div>
  </section>
</div>
    <div class="postbit" id="173403" data-post-id="173403">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="dimitarvp" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/dimitarvp/120/38664_2.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  dimitarvp
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<aside class="quote no-group" data-username="mischov" data-post="70" data-topic="4315">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/mischov/48/2489_2.png" class="avatar"> mischov:</div>
<blockquote>
<p>Erroring is the correct response, imo. If there’s a problem it’s better to know about it early and handle it correctly (as suggested in the Floki issue, by converting from whatever charset to UTF-8, then parsing) than to incorrectly parse and potentially return wrong answers or gibberish and not provide any context as to why.</p>
</blockquote>
</aside>
<p>I agree and I usually write my own libraries like this. But HTML and XML are a very notable exception. I remember back in 2007 a guy was writing an RSS parser and aggregator client and he actually had to normalise invalid XML in his program so it becomes a valid XML that’s parseable. It’s the sad reality of the web, and HTML is much, much worse.</p>
<p>What I would do if I was in the place of <code>html5ever</code> would be to add a setting that allows several encodings to be “tried” before giving up with an error.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="173403" data-batch-url="/posts/batch_likers">
                        0
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/meeseeks-a-library-for-extracting-data-from-html-and-xml-with-css-or-xpath-selectors/4315/71">Post #70</a>
	                </div>
	            </div>
              <div id="likers-container-173403" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="173403"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #70"></div>
  </section>
</div>
</template></turbo-stream><turbo-stream action="replace" target="load-more-container"><template><div id="load-more-container" class="load-more-container">
    <a class="load-more-button" data-turbo-stream="true" href="/topics/4315/load_more?page=8">Load more posts (16 remaining)</a>
</div></template></turbo-stream>