<?xml version="1.0" encoding="UTF-8"?>
<rdf:RDF xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" xmlns="http://purl.org/rss/1.0/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#"><channel rdf:about="https://www.bibsonomy.org/concept/tag/duplicate"><title>BibSonomy bookmarks for /concept/tag/duplicate</title><link>https://www.bibsonomy.org/concept/tag/duplicate</link><description>BibSonomy RSS Feed for /concept/tag/duplicate</description><items><rdf:Seq><rdf:li rdf:resource="https://github.com/spring-projects/spring-framework/issues/20611"/><rdf:li rdf:resource="http://www.chip.de/downloads/AllDup_22241569.html"/><rdf:li rdf:resource="https://github.com/iipc/openwayback/wiki/How-OpenWayback-handles-revisit-records-in-WARC-files"/><rdf:li rdf:resource="https://github.com/adrianlopezroche/fdupes"/><rdf:li rdf:resource="http://stackoverflow.com/questions/1746213/how-to-delete-duplicate-entries"/><rdf:li rdf:resource="http://www.seo-summary.de/doppelte-inhalte-duplicate-content-verhindern/"/><rdf:li rdf:resource="http://www.seerinteractive.com/blog/technical-foul-canonicals-and-duplicate-content"/><rdf:li rdf:resource="https://webarchive.jira.com/wiki/display/Heritrix/Duplication+Reduction+Processors"/><rdf:li rdf:resource="http://alexlurthu.wordpress.com/2008/04/04/skip-duplicate-entries-in-a-slave/"/><rdf:li rdf:resource="http://www.alldup.de/en_alldup.htm"/><rdf:li rdf:resource="http://www.hardcoded.net/dupeguru/"/><rdf:li rdf:resource="http://www.gossamer-threads.com/lists/lucene/java-dev/53351"/><rdf:li rdf:resource="http://wiki.apache.org/solr/Deduplication"/><rdf:li rdf:resource="http://www.gbv.de/wikis/cls/Bibliographic_Hash_Key"/><rdf:li rdf:resource="http://sourceforge.net/projects/doubles/"/><rdf:li rdf:resource="http://www.f2ko.de/programs.php?lang=de&amp;pid=dfe"/><rdf:li rdf:resource="http://www.pixelbeat.org/fslint/"/><rdf:li rdf:resource="http://duplicatefilessearcher.net/de/index_de.php"/><rdf:li rdf:resource="http://nlp.stanford.edu/IR-book/html/htmledition/near-duplicates-and-shingling-1.html"/><rdf:li rdf:resource="http://en.wikipedia.org/wiki/W-shingling"/></rdf:Seq></items></channel><item rdf:about="https://github.com/spring-projects/spring-framework/issues/20611"><title>Remove duplicate commons logging classes from spring-jcl [SPR-16062] · Issue #20611 · spring-projects/spring-framework · GitHub</title><description></description><link>https://github.com/spring-projects/spring-framework/issues/20611</link><dc:creator>jil</dc:creator><dc:date>2022-03-22T17:46:19+01:00</dc:date><dc:subject>package spring commons class jcl classes logging duplicate </dc:subject><content:encoded>&lt;a itemprop=&#034;url&#034; data-versiondate=&#034;2022-03-22T17:46:19+01:00&#034; href=&#034;https://github.com/spring-projects/spring-framework/issues/20611&#034; rel=&#034;nofollow&#034; class=&#034;description-link&#034;&gt;https://github.com/spring-projects/spring-framework/issues/20611&lt;/a&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/package"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/spring"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/commons"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/class"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/jcl"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/classes"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/logging"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/></rdf:Bag></taxo:topics></item><item rdf:about="http://www.chip.de/downloads/AllDup_22241569.html"><title>AllDup - Download - CHIP Online</title><description></description><link>http://www.chip.de/downloads/AllDup_22241569.html</link><dc:creator>thtbln</dc:creator><dc:date>2017-07-16T12:23:03+02:00</dc:date><dc:subject>tool file freeware duplicate </dc:subject><content:encoded>&lt;a itemprop=&#034;url&#034; data-versiondate=&#034;2017-07-16T12:23:03+02:00&#034; href=&#034;http://www.chip.de/downloads/AllDup_22241569.html&#034; rel=&#034;nofollow&#034; class=&#034;description-link&#034;&gt;http://www.chip.de/downloads/AllDup_22241569.html&lt;/a&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/tool"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/file"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/freeware"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/></rdf:Bag></taxo:topics></item><item rdf:about="https://github.com/iipc/openwayback/wiki/How-OpenWayback-handles-revisit-records-in-WARC-files"><title>How OpenWayback handles revisit records in WARC files</title><description></description><link>https://github.com/iipc/openwayback/wiki/How-OpenWayback-handles-revisit-records-in-WARC-files</link><dc:creator>jaeschke</dc:creator><dc:date>2017-01-18T16:12:58+01:00</dc:date><dc:subject>web openwayback wayback revisit archive duplicate warc </dc:subject><content:encoded>&lt;a itemprop=&#034;url&#034; data-versiondate=&#034;2017-01-18T16:12:58+01:00&#034; href=&#034;https://github.com/iipc/openwayback/wiki/How-OpenWayback-handles-revisit-records-in-WARC-files&#034; rel=&#034;nofollow&#034; class=&#034;description-link&#034;&gt;https://github.com/iipc/openwayback/wiki/How-OpenWayback-handles-revisit-records-in-WARC-files&lt;/a&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/web"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/openwayback"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/wayback"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/revisit"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/archive"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/warc"/></rdf:Bag></taxo:topics></item><item rdf:about="https://github.com/adrianlopezroche/fdupes"><title>adrianlopezroche/fdupes · GitHub</title><description></description><link>https://github.com/adrianlopezroche/fdupes</link><dc:creator>jil</dc:creator><dc:date>2015-10-23T19:37:27+02:00</dc:date><dc:subject>linux tool detection file duplicate </dc:subject><content:encoded>&lt;a itemprop=&#034;url&#034; data-versiondate=&#034;2015-10-23T19:37:27+02:00&#034; href=&#034;https://github.com/adrianlopezroche/fdupes&#034; rel=&#034;nofollow&#034; class=&#034;description-link&#034;&gt;https://github.com/adrianlopezroche/fdupes&lt;/a&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/linux"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/tool"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/detection"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/file"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/></rdf:Bag></taxo:topics></item><item rdf:about="http://stackoverflow.com/questions/1746213/how-to-delete-duplicate-entries"><title>sql - How to delete duplicate entries? - Stack Overflow</title><description></description><link>http://stackoverflow.com/questions/1746213/how-to-delete-duplicate-entries</link><dc:creator>jil</dc:creator><dc:date>2014-10-14T11:33:48+02:00</dc:date><dc:subject>duplicate delete sql entries </dc:subject><content:encoded>&lt;a itemprop=&#034;url&#034; data-versiondate=&#034;2014-10-14T11:33:48+02:00&#034; href=&#034;http://stackoverflow.com/questions/1746213/how-to-delete-duplicate-entries&#034; rel=&#034;nofollow&#034; class=&#034;description-link&#034;&gt;http://stackoverflow.com/questions/1746213/how-to-delete-duplicate-entries&lt;/a&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/delete"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/sql"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/entries"/></rdf:Bag></taxo:topics></item><item rdf:about="http://www.seo-summary.de/doppelte-inhalte-duplicate-content-verhindern/"><title>Duplicate Content | Doppelte Inhalte finden &amp; vermeiden!</title><description></description><link>http://www.seo-summary.de/doppelte-inhalte-duplicate-content-verhindern/</link><dc:creator>esistimfluss</dc:creator><dc:date>2014-08-18T10:43:23+02:00</dc:date><dc:subject>technical redirect 2014 duplicate htaccess google rules seo content </dc:subject><content:encoded>&lt;a itemprop=&#034;url&#034; data-versiondate=&#034;2014-08-18T10:43:23+02:00&#034; href=&#034;http://www.seo-summary.de/doppelte-inhalte-duplicate-content-verhindern/&#034; rel=&#034;nofollow&#034; class=&#034;description-link&#034;&gt;http://www.seo-summary.de/doppelte-inhalte-duplicate-content-verhindern/&lt;/a&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/technical"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/redirect"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/2014"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/htaccess"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/google"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/rules"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/seo"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/content"/></rdf:Bag></taxo:topics></item><item rdf:about="http://www.seerinteractive.com/blog/technical-foul-canonicals-and-duplicate-content"><title>Technical Foul! Canonicals and Duplicate Content | SEER Interactive</title><description>Nobody knows technical fouls better than Ron Artest, period. You can probably make an argument for Rasheed or Malone but no one took it to another level than</description><link>http://www.seerinteractive.com/blog/technical-foul-canonicals-and-duplicate-content</link><dc:creator>esistimfluss</dc:creator><dc:date>2014-07-18T12:08:08+02:00</dc:date><dc:subject>t24 2014 duplicate google seo content </dc:subject><content:encoded>&lt;span itemprop=&#034;description&#034;&gt;Nobody knows technical fouls better than Ron Artest, period. You can probably make an argument for Rasheed or Malone but no one took it to another level than&lt;/span&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/t24"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/2014"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/google"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/seo"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/content"/></rdf:Bag></taxo:topics></item><item rdf:about="https://webarchive.jira.com/wiki/display/Heritrix/Duplication+Reduction+Processors"><title>Duplication Reduction Processors - Heritrix - IA Webteam Confluence</title><description></description><link>https://webarchive.jira.com/wiki/display/Heritrix/Duplication+Reduction+Processors</link><dc:creator>jaeschke</dc:creator><dc:date>2013-07-10T11:32:55+02:00</dc:date><dc:subject>heritrix crawling duplicate recrawl </dc:subject><content:encoded>&lt;a itemprop=&#034;url&#034; data-versiondate=&#034;2013-07-10T11:32:55+02:00&#034; href=&#034;https://webarchive.jira.com/wiki/display/Heritrix/Duplication+Reduction+Processors&#034; rel=&#034;nofollow&#034; class=&#034;description-link&#034;&gt;https://webarchive.jira.com/wiki/display/Heritrix/Duplication+Reduction+Processors&lt;/a&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/heritrix"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/crawling"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/recrawl"/></rdf:Bag></taxo:topics></item><item rdf:about="http://alexlurthu.wordpress.com/2008/04/04/skip-duplicate-entries-in-a-slave/"><title>Skip duplicate entries in a slave « Random Thoughts</title><description>The following one liner helps to sync a slave that is facing duplicate entry errors. while [ 1 ]; do …Continue reading »</description><link>http://alexlurthu.wordpress.com/2008/04/04/skip-duplicate-entries-in-a-slave/</link><dc:creator>nosebrain</dc:creator><dc:date>2012-12-10T17:56:53+01:00</dc:date><dc:subject>slave entry replication mysql duplicate master </dc:subject><content:encoded>&lt;span itemprop=&#034;description&#034;&gt;The following one liner helps to sync a slave that is facing duplicate entry errors. while [ 1 ]; do …Continue reading »&lt;/span&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/slave"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/entry"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/replication"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/mysql"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/master"/></rdf:Bag></taxo:topics></item><item rdf:about="http://www.alldup.de/en_alldup.htm"><title>find and delete hard links, best duplicate file finder, duplicate file finder, picture duplicate finder, find duplicate mp3, freeware duplicate, find duplicate pictures, duplicate photo, search for duplicate files, delete duplicate files, duplicate picture</title><description>AllDup is a freeware tool for searching and removing file duplicates on your computer. The fast search algorithm find duplicates of any file type, e.g., text, pictures, music or movies. The powerful search engine enables you to find duplicates with a combination of the following criteria: File Name, File Extension, File Size, File Content, Last Modified Date, Create Date, File Attributes and Hard Links
Features

Freeware for private and commercial use
Fast search algorithm
Find duplicates with a combination of the following criteria: file content, file name, file extension, file dates and file attributes!
Search is performed in multiple specified folders, drives, media storages, CD/DVDs, network drives...
Search through an unlimited number of files and folders
Search for duplicates of music and video files
Search for duplicates of digital photo files
Search for duplicates of executable and any other files
Search for hard links
Entire folders or individual files can be excluded from the search by masks or size conditions
The built-in file viewer allows you to preview many different file formats and analyze the content of the file before deciding what to do with it
Ignore the ID3 tags of MP3 files
Convenient search result list
Many flexible options help you to select unnecessary duplicates automatically
The unnecessary duplicates can be deleted permanently or copied/moved to a folder of your choice
Save and restore the search result for continue working later
Export the search result to TXT or CSV file
Turn duplicate files into hard links (NTFS file systems only) or shortcuts
Detailed log file about all actions
System Requirements

Microsoft Windows 7 (all versions)
Microsoft Windows Server 2008 (all versions)
Microsoft Windows Vista (all versions)
Microsoft Windows Server 2003 (all versions)
Microsoft Windows XP (all versions) *
Microsoft Windows 2000 (all versions) *

AllDup also runs on the 64-bit versions of these operating systems. 

* The portable version of AllDup doesn&#039;t support Windows XP without any Service Pack installed and Microsoft Windows 2000.</description><link>http://www.alldup.de/en_alldup.htm</link><dc:creator>gresch</dc:creator><dc:date>2011-07-27T12:08:28+02:00</dc:date><dc:subject>software duplicate tools windows filesystem </dc:subject><content:encoded>&lt;span itemprop=&#034;description&#034;&gt;AllDup is a freeware tool for searching and removing file duplicates on your computer. The fast search algorithm find duplicates of any file type, e.g., text, pictures, music or movies. The powerful search engine enables you to find duplicates with a combination of the following criteria: File Name, File Extension, File Size, File Content, Last Modified Date, Create Date, File Attributes and Hard Links
Features

Freeware for private and commercial use
Fast search algorithm
Find duplicates with a combination of the following criteria: file content, file name, file extension, file dates and file attributes!
Search is performed in multiple specified folders, drives, media storages, CD/DVDs, network drives...
Search through an unlimited number of files and folders
Search for duplicates of music and video files
Search for duplicates of digital photo files
Search for duplicates of executable and any other files
Search for hard links
Entire folders or individual files can be excluded from the search by masks or size conditions
The built-in file viewer allows you to preview many different file formats and analyze the content of the file before deciding what to do with it
Ignore the ID3 tags of MP3 files
Convenient search result list
Many flexible options help you to select unnecessary duplicates automatically
The unnecessary duplicates can be deleted permanently or copied/moved to a folder of your choice
Save and restore the search result for continue working later
Export the search result to TXT or CSV file
Turn duplicate files into hard links (NTFS file systems only) or shortcuts
Detailed log file about all actions
System Requirements

Microsoft Windows 7 (all versions)
Microsoft Windows Server 2008 (all versions)
Microsoft Windows Vista (all versions)
Microsoft Windows Server 2003 (all versions)
Microsoft Windows XP (all versions) *
Microsoft Windows 2000 (all versions) *

AllDup also runs on the 64-bit versions of these operating systems. 

* The portable version of AllDup doesn&amp;#039;t support Windows XP without any Service Pack installed and Microsoft Windows 2000.&lt;/span&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/software"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/tools"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/windows"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/filesystem"/></rdf:Bag></taxo:topics></item><item rdf:about="http://www.hardcoded.net/dupeguru/"><title>dupeGuru - Duplicate file scanner</title><description>dupeGuru is a tool to find duplicate files on your computer. It can scan either filenames or contents. The filename scan features a fuzzy matching algorithm that can find duplicate filenames even when they are not exactly the same. dupeGuru runs on Windows, Mac OS X and Linux.

dupeGuru is efficient. Find your duplicate files in minutes, thanks to its quick fuzzy matching algorithm. dupeGuru not only finds filenames that are the same, but it also finds similar filenames.

dupeGuru is customizable. You can tweak its matching engine to find exactly the kind of duplicates you want to find. The Preference page of the help file lists all the scanning engine settings you can change.

dupeGuru is safe. Its engine has been especially designed with safety in mind. Its reference directory system as well as its grouping system prevent you from deleting files you didn&#039;t mean to delete.

Do whatever you want with your duplicates. Not only can you delete duplicates files dupeGuru finds, but you can also move or copy them elsewhere. There are also multiple ways to filter and sort your results to easily weed out false duplicates (for low threshold scans).

Supported languages: English, French.

Requirements

Mac OS X: 10.5 and up (Leopard, Snow Leopard or Lion). PowerPC or Intel. (Last version to support Tiger: v2.8.2)
Windows: 2k/XP/Vista/Win7.
Linux: Ubuntu 10.04</description><link>http://www.hardcoded.net/dupeguru/</link><dc:creator>gresch</dc:creator><dc:date>2011-07-27T12:07:40+02:00</dc:date><dc:subject>linux dupl software duplicate tools windows macos </dc:subject><content:encoded>&lt;span itemprop=&#034;description&#034;&gt;dupeGuru is a tool to find duplicate files on your computer. It can scan either filenames or contents. The filename scan features a fuzzy matching algorithm that can find duplicate filenames even when they are not exactly the same. dupeGuru runs on Windows, Mac OS X and Linux.

dupeGuru is efficient. Find your duplicate files in minutes, thanks to its quick fuzzy matching algorithm. dupeGuru not only finds filenames that are the same, but it also finds similar filenames.

dupeGuru is customizable. You can tweak its matching engine to find exactly the kind of duplicates you want to find. The Preference page of the help file lists all the scanning engine settings you can change.

dupeGuru is safe. Its engine has been especially designed with safety in mind. Its reference directory system as well as its grouping system prevent you from deleting files you didn&amp;#039;t mean to delete.

Do whatever you want with your duplicates. Not only can you delete duplicates files dupeGuru finds, but you can also move or copy them elsewhere. There are also multiple ways to filter and sort your results to easily weed out false duplicates (for low threshold scans).

Supported languages: English, French.

Requirements

Mac OS X: 10.5 and up (Leopard, Snow Leopard or Lion). PowerPC or Intel. (Last version to support Tiger: v2.8.2)
Windows: 2k/XP/Vista/Win7.
Linux: Ubuntu 10.04&lt;/span&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/linux"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/dupl"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/software"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/tools"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/windows"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/macos"/></rdf:Bag></taxo:topics></item><item rdf:about="http://www.gossamer-threads.com/lists/lucene/java-dev/53351"><title>DuplicatesFilter - one for contrib? | Lucene | Java-Dev</title><description></description><link>http://www.gossamer-threads.com/lists/lucene/java-dev/53351</link><dc:creator>folke</dc:creator><dc:date>2011-07-25T18:00:12+02:00</dc:date><dc:subject>detection dedupe remove duplicate lucene duplicatesfilter </dc:subject><content:encoded>&lt;a itemprop=&#034;url&#034; data-versiondate=&#034;2011-07-25T18:00:12+02:00&#034; href=&#034;http://www.gossamer-threads.com/lists/lucene/java-dev/53351&#034; rel=&#034;nofollow&#034; class=&#034;description-link&#034;&gt;http://www.gossamer-threads.com/lists/lucene/java-dev/53351&lt;/a&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/detection"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/dedupe"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/remove"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/lucene"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicatesfilter"/></rdf:Bag></taxo:topics></item><item rdf:about="http://wiki.apache.org/solr/Deduplication"><title>Deduplication - Solr Wiki</title><description></description><link>http://wiki.apache.org/solr/Deduplication</link><dc:creator>stroeh</dc:creator><dc:date>2011-06-29T12:37:29+02:00</dc:date><dc:subject>searchengine solr ir deduplication duplicate </dc:subject><content:encoded>&lt;a itemprop=&#034;url&#034; data-versiondate=&#034;2011-06-29T12:37:29+02:00&#034; href=&#034;http://wiki.apache.org/solr/Deduplication&#034; rel=&#034;nofollow&#034; class=&#034;description-link&#034;&gt;http://wiki.apache.org/solr/Deduplication&lt;/a&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/searchengine"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/solr"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/ir"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/deduplication"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/></rdf:Bag></taxo:topics></item><item rdf:about="http://www.gbv.de/wikis/cls/Bibliographic_Hash_Key"><title>Bibliographic Hash Key – Verbund-Wiki GBV</title><description></description><link>http://www.gbv.de/wikis/cls/Bibliographic_Hash_Key</link><dc:creator>stroeh</dc:creator><dc:date>2011-06-27T08:55:59+02:00</dc:date><dc:subject>hashkey bibkey bibliographic duplicate hash </dc:subject><content:encoded>&lt;a itemprop=&#034;url&#034; data-versiondate=&#034;2011-06-27T08:55:59+02:00&#034; href=&#034;http://www.gbv.de/wikis/cls/Bibliographic_Hash_Key&#034; rel=&#034;nofollow&#034; class=&#034;description-link&#034;&gt;http://www.gbv.de/wikis/cls/Bibliographic_Hash_Key&lt;/a&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/hashkey"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/bibkey"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/bibliographic"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/hash"/></rdf:Bag></taxo:topics></item><item rdf:about="http://sourceforge.net/projects/doubles/"><title>Duplicate Files Finder | Download Duplicate Files Finder software for free at SourceForge.net</title><description>Duplicate Files Finder is a cross-platform application for finding and removing duplicate files by deleting, creating hardlinks or creating symbolic links. A special algorithm minimizes the amount of data read from disk, so the program is very fast.</description><link>http://sourceforge.net/projects/doubles/</link><dc:creator>gresch</dc:creator><dc:date>2011-05-27T09:34:41+02:00</dc:date><dc:subject>doubletten java software duplicate tools </dc:subject><content:encoded>&lt;span itemprop=&#034;description&#034;&gt;Duplicate Files Finder is a cross-platform application for finding and removing duplicate files by deleting, creating hardlinks or creating symbolic links. A special algorithm minimizes the amount of data read from disk, so the program is very fast.&lt;/span&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/doubletten"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/java"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/software"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/tools"/></rdf:Bag></taxo:topics></item><item rdf:about="http://www.f2ko.de/programs.php?lang=de&amp;pid=dfe"><title>Duplicate File Eraser</title><description>Duplicate File Eraser findet und entfernt Duplikate.

Duplikate sind mehrfach vorkommende identische Dateien und
belegen oftmals unnötigen Speicherplatz,
der mit diesem Programm wieder freigegeben werden kann.

Die Duplikate werden mittels MD5, CRC32 oder SHA1 ermittelt,
und können anschliessend gelöscht werden.</description><link>http://www.f2ko.de/programs.php?lang=de&amp;pid=dfe</link><dc:creator>gresch</dc:creator><dc:date>2011-05-27T09:33:54+02:00</dc:date><dc:subject>doubletten linux software duplicate tools </dc:subject><content:encoded>&lt;span itemprop=&#034;description&#034;&gt;Duplicate File Eraser findet und entfernt Duplikate.

Duplikate sind mehrfach vorkommende identische Dateien und
belegen oftmals unnötigen Speicherplatz,
der mit diesem Programm wieder freigegeben werden kann.

Die Duplikate werden mittels MD5, CRC32 oder SHA1 ermittelt,
und können anschliessend gelöscht werden.&lt;/span&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/doubletten"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/linux"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/software"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/tools"/></rdf:Bag></taxo:topics></item><item rdf:about="http://www.pixelbeat.org/fslint/"><title>FSlint - Duplicate file finder for linux</title><description>FSlint is a utility to find and clean various forms of lint on a filesystem.
I.E. unwanted or problematic cruft in your files or file names.
For example, one form of lint it finds is duplicate files.
It has both GUI and command line modes.</description><link>http://www.pixelbeat.org/fslint/</link><dc:creator>gresch</dc:creator><dc:date>2011-05-27T09:31:13+02:00</dc:date><dc:subject>doubletten linux software duplicate tools filesystem </dc:subject><content:encoded>&lt;span itemprop=&#034;description&#034;&gt;FSlint is a utility to find and clean various forms of lint on a filesystem.
I.E. unwanted or problematic cruft in your files or file names.
For example, one form of lint it finds is duplicate files.
It has both GUI and command line modes.&lt;/span&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/doubletten"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/linux"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/software"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/tools"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/filesystem"/></rdf:Bag></taxo:topics></item><item rdf:about="http://duplicatefilessearcher.net/de/index_de.php"><title>Duplicate Files Searcher - Das Tool zur Suche nach doppelten Dateien, findet und löscht Duplikate von Dateien im Computer.</title><description>Was ist DFS – DuplicateFiles Seracher - Das Tool zur Suche nach doppelten Dateien?
    
  DFS – Duplicate Files Seracher - Das Tool zur Suche nach doppelten Dateien  findet und löscht Duplikate von Dateien im Computer. Es kann aber auch benutzt werden, um MD5 oder SHA Hashes zu berechnen. Es sind zwei Versionen verfügbar
die open source – Version.  Die Version open source befindet sich in der Datei dfs.jar
die voll version (unentgeltlich, aber nicht open source). Die Vollversion ist zugänglich als Datei mit dem Namen dfsfull.jar.</description><link>http://duplicatefilessearcher.net/de/index_de.php</link><dc:creator>gresch</dc:creator><dc:date>2011-05-27T09:28:28+02:00</dc:date><dc:subject>doubletten java software duplicate tools </dc:subject><content:encoded>&lt;span itemprop=&#034;description&#034;&gt;Was ist DFS – DuplicateFiles Seracher - Das Tool zur Suche nach doppelten Dateien?
    
  DFS – Duplicate Files Seracher - Das Tool zur Suche nach doppelten Dateien  findet und löscht Duplikate von Dateien im Computer. Es kann aber auch benutzt werden, um MD5 oder SHA Hashes zu berechnen. Es sind zwei Versionen verfügbar
die open source – Version.  Die Version open source befindet sich in der Datei dfs.jar
die voll version (unentgeltlich, aber nicht open source). Die Vollversion ist zugänglich als Datei mit dem Namen dfsfull.jar.&lt;/span&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/doubletten"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/java"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/software"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/tools"/></rdf:Bag></taxo:topics></item><item rdf:about="http://nlp.stanford.edu/IR-book/html/htmledition/near-duplicates-and-shingling-1.html"><title>Near-duplicates and shingling</title><description>can now generate all pairs $i,j$ for which $x_i^\pi$ is present in both their sketches. From these we can compute, for each pair $i,j$ with non-zero sketch overlap, a count of the number of $x_i^\pi$ values they have in common. By applying a preset threshold, we know which pairs $i,j$ have heavily overlapping sketches. For instance, if the threshold were 80%, we would need the count to be at least 160 for any $i,j$. As we identify such pairs, we run the union-find to group documents into near-duplicate ``syntactic clusters&#039;&#039;. This is essentially a variant of the single-link clustering algorithm introduced in Section 17.2 (page [*]). </description><link>http://nlp.stanford.edu/IR-book/html/htmledition/near-duplicates-and-shingling-1.html</link><dc:creator>stroeh</dc:creator><dc:date>2011-03-09T13:51:35+01:00</dc:date><dc:subject>near shingle duplicate shingling </dc:subject><content:encoded>&lt;span itemprop=&#034;description&#034;&gt;can now generate all pairs $i,j$ for which $x_i^\pi$ is present in both their sketches. From these we can compute, for each pair $i,j$ with non-zero sketch overlap, a count of the number of $x_i^\pi$ values they have in common. By applying a preset threshold, we know which pairs $i,j$ have heavily overlapping sketches. For instance, if the threshold were 80%, we would need the count to be at least 160 for any $i,j$. As we identify such pairs, we run the union-find to group documents into near-duplicate ``syntactic clusters&amp;#039;&amp;#039;. This is essentially a variant of the single-link clustering algorithm introduced in Section 17.2 (page [*]). &lt;/span&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/near"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/shingle"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/shingling"/></rdf:Bag></taxo:topics></item><item rdf:about="http://en.wikipedia.org/wiki/W-shingling"><title>w-shingling - Wikipedia, the free encyclopedia</title><description></description><link>http://en.wikipedia.org/wiki/W-shingling</link><dc:creator>stroeh</dc:creator><dc:date>2011-03-09T12:57:28+01:00</dc:date><dc:subject>detection ähnlichkeitsmaß shingle duplicate shingling w-shingling </dc:subject><content:encoded>&lt;a itemprop=&#034;url&#034; data-versiondate=&#034;2011-03-09T12:57:28+01:00&#034; href=&#034;http://en.wikipedia.org/wiki/W-shingling&#034; rel=&#034;nofollow&#034; class=&#034;description-link&#034;&gt;http://en.wikipedia.org/wiki/W-shingling&lt;/a&gt;</content:encoded><taxo:topics><rdf:Bag><rdf:li rdf:resource="https://www.bibsonomy.org/tag/detection"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/ähnlichkeitsmaß"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/shingle"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/duplicate"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/shingling"/><rdf:li rdf:resource="https://www.bibsonomy.org/tag/w-shingling"/></rdf:Bag></taxo:topics></item></rdf:RDF>