Commit Graph

11156 Commits (ebd0be2cea2bf06a5bafa3e521d7b797f2df78c1)
 

Author SHA1 Message Date
orbiter c48d2a2a02 npe fix 11 years ago
reger 121d25be38 recover sax fatal error on OAI-PMH import of xml with entity error 11 years ago
reger 81dc2aa536 add current css to HTMLResponseWriter to fix metadata view 11 years ago
orbiter 2fd8a0ead6 Merge branch 'master' of git@gitorious.org:yacy/rc1.git 11 years ago
orbiter 8e5ce7cd51 fixed a situation where finished crawls had not been detected. 11 years ago
orbiter c6f0bd05f8 better removal of stored urls when doing a crawl start 11 years ago
orbiter 2f63bd0261 enhanced Host Balancer strategy: fair round robin 11 years ago
orbiter 0c88a32c36 do not apply lazy value instantiation for numeric or boolean values 11 years ago
orbiter 8e04030596 in case of short memory, do not cut down robinson peers to 1, just 11 years ago
reger 86f6975edc exclude html tags in in/outboundlinks_anchortext_txt parsed text 11 years ago
orbiter 469e0a62f1 added new button to terminate all crawls 11 years ago
orbiter ccb1864d55 catch IllegalArgumentException for wrong process types (that is needed 11 years ago
orbiter 4ee4ba1576 fix for NPE in IndexCreateParserErrors_p.html caused by bad handling of 11 years ago
orbiter 12ba890205 removed warnings 11 years ago
reger d51f9cc863 add custom Jetty errorhandler 11 years ago
reger c193a02023 defer creation of new ArrayList after possible early return 11 years ago
reger 727dfb5875 refactore URIMetadataNode to further unify interaction with index 11 years ago
reger 79e7947442 - remove empty http0_9 status text array 11 years ago
reger 2dabe2009d - remove unused manual http KeepAlive config 11 years ago
Michael Peter Christen 5746aae3db add canonical links to the same crawldepth, not the next crawldepth 11 years ago
Michael Peter Christen 74ab5ef9fa increased runtime for postprocessing query job 11 years ago
Michael Peter Christen 8b32dd5f9e special strategy for balancer: do not remove targets with zero wait time 11 years ago
Michael Peter Christen 9c6228d948 fix for deadlocks in crawler 11 years ago
Michael Peter Christen 7a2f3e2353 increased resource.disk.used.max.steadystate and 11 years ago
Michael Peter Christen 10cf8215bd added crawl depth for failed documents 11 years ago
Michael Peter Christen 7fefebaeca Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 11 years ago
Michael Peter Christen c2f62e783f - better subgraph handling, less overhead for crawls without the 11 years ago
Michael Peter Christen 06afb568e2 new Strategies in Balancer: 11 years ago
Michael Peter Christen 1aea01fe5b fix for Table in case that requested file does not exist and paths also 11 years ago
reger 710054bb37 implement gzip input handling directly in defaultservlet 11 years ago
Michael Peter Christen b4b0d14c04 fix for display bug 11 years ago
Michael Peter Christen 9a5ab4e2c1 removed clickdepth_i field and related postprocessing. This information 11 years ago
Michael Peter Christen da86f150ab - added a new Crawler Balancer: HostBalancer and HostQueues: 11 years ago
Michael Peter Christen 075b6f9278 refactoring of the crawl balancer: the balancer is turned into an 11 years ago
Michael Peter Christen 8470dfe3f8 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 11 years ago
reger 46016fa153 autoupdate fails to download latest release (1.71) due to default release blacklist 11 years ago
Michael Peter Christen 8aeef73d49 fix for virtual root nodes 11 years ago
Michael Peter Christen 7c7fbb9818 find depth-matches also for edge targets 11 years ago
Michael Peter Christen dd12dd392f introduction of a data structure for HyperlinkEdges which should use 11 years ago
Michael Peter Christen 6ea8bb7348 using MultiProtocolURL for edge data which is faster (hash computation 11 years ago
Michael Peter Christen b21c208b4d enhanced hashcode computation for MultiProtocolURL 11 years ago
Michael Peter Christen ce1d1b2fa0 fix for maximum tag length in parser 11 years ago
Michael Peter Christen 17e0956312 refactoring of SystemLoad calls (only one backend tool) 11 years ago
Michael Peter Christen a37d067692 refactoring 11 years ago
orbiter 95780eed32 Merge branch 'master' of git@gitorious.org:yacy/rc1.git 11 years ago
Michael Peter Christen 67beef657f strong redesign of html parser: object recursion is now made using a 11 years ago
Michael Peter Christen 6bd8c6f195 fix for wrong status codes of error pages 11 years ago
Michael Peter Christen 9e503b3376 also delete the robots.txt file from the cache when a new crawl is 11 years ago
orbiter 67501c9dda Merge branch 'master' of git@gitorious.org:yacy/rc1.git 11 years ago
Michael Peter Christen 1c21b3256d fix for robots.txt handling: delete old entry before starting a new 11 years ago