<?xml version="1.0" encoding="UTF-8"?>
<rdf:RDF xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" xmlns="http://purl.org/rss/1.0/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#"><channel rdf:about="https://www.bibsonomy.org/user/stroeh/shingle"><title>BibSonomy bookmarks for /user/stroeh/shingle</title><link>https://www.bibsonomy.org/user/stroeh/shingle</link><description>BibSonomy RSS Feed for /user/stroeh/shingle</description><items><rdf:Seq><rdf:li rdf:resource="http://nlp.stanford.edu/IR-book/html/htmledition/near-duplicates-and-shingling-1.html"/><rdf:li rdf:resource="http://www.std.org/~msm/common/clustering.html"/><rdf:li rdf:resource="http://de.wikipedia.org/wiki/N-Gramm"/><rdf:li rdf:resource="http://en.wikipedia.org/wiki/W-shingling"/></rdf:Seq></items></channel><item rdf:about="http://nlp.stanford.edu/IR-book/html/htmledition/near-duplicates-and-shingling-1.html"><title>Near-duplicates and shingling</title><description>can now generate all pairs $i,j$ for which $x_i^\pi$ is present in both their sketches. From these we can compute, for each pair $i,j$ with non-zero sketch overlap, a count of the number of $x_i^\pi$ values they have in common. By applying a preset threshold, we know which pairs $i,j$ have heavily overlapping sketches. For instance, if the threshold were 80%, we would need the count to be at least 160 for any $i,j$. As we identify such pairs, we run the union-find to group documents into near-duplicate ``syntactic clusters&#039;&#039;. This is essentially a variant of the single-link clustering algorithm introduced in Section 17.2 (page [*]). </description><link>http://nlp.stanford.edu/IR-book/html/htmledition/near-duplicates-and-shingling-1.html</link><dc:creator>stroeh</dc:creator><dc:date>2011-03-09T13:51:35+01:00</dc:date><dc:subject>duplicate near shingle shingling </dc:subject><content:encoded>&lt;span itemprop=&#034;description&#034;&gt;can now generate all pairs $i,j$ for which $x_i^\pi$ is present in both their sketches. From these we can compute, for each pair $i,j$ with non-zero sketch overlap, a count of the number of $x_i^\pi$ values they have in common. By applying a preset threshold, we know which pairs $i,j$ have heavily overlapping sketches. For instance, if the threshold were 80%, we would need the count to be at least 160 for any $i,j$. As we identify such pairs, we run the union-find to group documents into near-duplicate ``syntactic clusters&amp;#039;&amp;#039;. This is essentially a variant of the single-link clustering algorithm introduced in Section 17.2 (page [*]). &lt;/span&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/near"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/shingle"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/shingling"/></rdf:Bag></taxo:topics></item><item rdf:about="http://www.std.org/~msm/common/clustering.html"><title>Syntactic Clustering of the Web</title><description></description><link>http://www.std.org/~msm/common/clustering.html</link><dc:creator>stroeh</dc:creator><dc:date>2011-03-09T13:46:59+01:00</dc:date><dc:subject>resemblance shingle ähnlichkeitsmaß </dc:subject><content:encoded>&lt;a itemprop=&#034;url&#034; data-versiondate=&#034;2011-03-09T13:46:59+01:00&#034; href=&#034;http://www.std.org/~msm/common/clustering.html&#034; rel=&#034;nofollow&#034; class=&#034;description-link&#034;&gt;http://www.std.org/~msm/common/clustering.html&lt;/a&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/resemblance"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/shingle"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/ähnlichkeitsmaß"/></rdf:Bag></taxo:topics></item><item rdf:about="http://de.wikipedia.org/wiki/N-Gramm"><title>N-Gramm – Wikipedia</title><description></description><link>http://de.wikipedia.org/wiki/N-Gramm</link><dc:creator>stroeh</dc:creator><dc:date>2011-03-09T13:07:34+01:00</dc:date><dc:subject>bigramm n-gramm shingle trigramm </dc:subject><content:encoded>&lt;a itemprop=&#034;url&#034; data-versiondate=&#034;2011-03-09T13:07:34+01:00&#034; href=&#034;http://de.wikipedia.org/wiki/N-Gramm&#034; rel=&#034;nofollow&#034; class=&#034;description-link&#034;&gt;http://de.wikipedia.org/wiki/N-Gramm&lt;/a&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/bigramm"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/n-gramm"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/shingle"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/trigramm"/></rdf:Bag></taxo:topics></item><item rdf:about="http://en.wikipedia.org/wiki/W-shingling"><title>w-shingling - Wikipedia, the free encyclopedia</title><description></description><link>http://en.wikipedia.org/wiki/W-shingling</link><dc:creator>stroeh</dc:creator><dc:date>2011-03-09T12:57:28+01:00</dc:date><dc:subject>detection duplicate shingle shingling w-shingling ähnlichkeitsmaß </dc:subject><content:encoded>&lt;a itemprop=&#034;url&#034; data-versiondate=&#034;2011-03-09T12:57:28+01:00&#034; href=&#034;http://en.wikipedia.org/wiki/W-shingling&#034; rel=&#034;nofollow&#034; class=&#034;description-link&#034;&gt;http://en.wikipedia.org/wiki/W-shingling&lt;/a&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/detection"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/shingle"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/shingling"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/w-shingling"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/ähnlichkeitsmaß"/></rdf:Bag></taxo:topics></item></rdf:RDF>