Commit Graph

175 Commits (2fd8a0ead6747091a633310183c0498515bb6f28)

Author SHA1 Message Date
orbiter 2f63bd0261 enhanced Host Balancer strategy: fair round robin 11 years ago
Michael Peter Christen 8b32dd5f9e special strategy for balancer: do not remove targets with zero wait time 11 years ago
Michael Peter Christen 9c6228d948 fix for deadlocks in crawler 11 years ago
Michael Peter Christen 10cf8215bd added crawl depth for failed documents 11 years ago
Michael Peter Christen 06afb568e2 new Strategies in Balancer: 11 years ago
Michael Peter Christen da86f150ab - added a new Crawler Balancer: HostBalancer and HostQueues: 11 years ago
Michael Peter Christen 075b6f9278 refactoring of the crawl balancer: the balancer is turned into an 11 years ago
Michael Peter Christen 6bd8c6f195 fix for wrong status codes of error pages 11 years ago
Michael Peter Christen 9e503b3376 also delete the robots.txt file from the cache when a new crawl is 11 years ago
Michael Peter Christen 1c21b3256d fix for robots.txt handling: delete old entry before starting a new 11 years ago
Michael Peter Christen 926d28dd3f fixed a bug which prevented crawl starts after a network switch 11 years ago
Michael Peter Christen d4b5c457e4 NPE fix 11 years ago
Michael Peter Christen 8b44fcf0f4 added missing @Override annotation 11 years ago
Michael Peter Christen 85a427ec54 support for multiple sitemaps in robots.txt 11 years ago
Michael Peter Christen b08375da33 fix for bad/missing values of size_i 11 years ago
reger dd5bf0b71b cleanup old reference to HTTPDemon.setAlternativeResolver 11 years ago
Michael Peter Christen e485fbd0ce - let crawl loader jobs die after 10 seconds without new jobs 11 years ago
Michael Peter Christen bcd9dd9e1d enhanced concurrent loading by using a fixed set of concurrent loader 11 years ago
Michael Peter Christen 6ed9c0164e attaching names to all Threads to get a better view in profiling tools 11 years ago
Michael Peter Christen fdaeac374a - enhanced postprocessing speed and memory footprint (by using HashMaps 11 years ago
orbiter da5d4128bf prevent npe 11 years ago
orbiter a878c7982c prevent npe 11 years ago
orbiter ced1a96f9c fixed error cache 11 years ago
Michael Peter Christen 69391e5d9e changed strategy to test existence of documents in Solr: using the 11 years ago
Michael Peter Christen 8b14e92ba4 added button in host browser to re-load 404/failed documents 11 years ago
Michael Peter Christen 6ada0daae9 making latency_factor and maximum number of same hosts in loader queue 11 years ago
Michael Peter Christen 0168f80c28 new crawling factors can now be changed during runtime 11 years ago
Michael Peter Christen 77531850b5 reverted crawling strategy from latest commit. 11 years ago
Michael Peter Christen c0da966dfa enhanced crawler speed 11 years ago
Michael Peter Christen 0d235a565b cleanup crawl loader jobs 11 years ago
Michael Peter Christen 1ea17bd9f3 - removed old metadata database and all migration code 11 years ago
Michael Peter Christen 022c6d3ce1 do YaCy p2p connections using a timeout-request which covers the http 11 years ago
reger 28eae57e8b spend CrawlQueues a fremem routine 11 years ago
reger 6932aa4d7a use configured admin-username for api calls 11 years ago
orbiter 3cb6c7861f fixed shutdown authenticaton problem 11 years ago
orbiter f3ac923a7e ftp client shall be able to open non-anonymous ftp servers if login 11 years ago
Michael Peter Christen 82c0525e71 wrong logger fix 11 years ago
Michael Peter Christen 552ef9f18e fix for bad ErrorCache.exists test (bug from latest commit) 11 years ago
Michael Peter Christen 303f5694ba avoid usage of existsByQuery. If a document can be loaded by the ID 11 years ago
Michael Peter Christen 0db8e34625 enhanced webgraph processing 11 years ago
Michael Peter Christen 1a4a69c226 set more logger to 'final static' 11 years ago
Michael Peter Christen 87a956e881 calculating and showing the number of files and the average size of a 11 years ago
Michael Peter Christen 234a974955 load image only if their parser flag is activated 12 years ago
Michael Peter Christen 030d0776ff Enhanced crawl start for very, very large crawl lists (i.e. > 5000) 12 years ago
Michael Peter Christen 4948c39e48 added concurrency for mass crawl check 12 years ago
Michael Peter Christen 1b4fa2947d - fixed a problem which ocurred when a document was not recognized with 12 years ago
orbiter 20bbde8665 fix for mustmatch regex computation: result had correct semantic, but 12 years ago
Michael Peter Christen 74d0256e93 enhanced postprocessing: fixed bugs, enable proper postprocessing also 12 years ago
Michael Peter Christen 101a6e6e14 Patch the citation index for links with canonical tags. 12 years ago
reger fd119deb00 fix NPE on modified since check ( Response.requestHeader allowed to be null) 12 years ago