Commit Graph

287 Commits (c5df34989eef3c81c58016a473d5319d2d4814f1)

Author SHA1 Message Date
Michael Peter Christen f5ca5cea44 - added field options to all solr queries. This can be used to restrict 13 years ago
cominch d2a94cc55e refactor package 13 years ago
Michael Peter Christen 8b1c9cba3d fixed a problem with non-terminating crawls 13 years ago
Michael Peter Christen 71ed8e5e07 bugfixes for crawler 13 years ago
Michael Peter Christen 158732af37 automatically delete entries from the crawl profile list if crawl is 13 years ago
Michael Peter Christen d481abd087 added the visualization of error-urls to host browser 13 years ago
Michael Peter Christen 791e1dcfdf when a new crawl is started, delete all entries about error-urls for 13 years ago
orbiter 354ef8000d - added 'deleteold' option to crawler which causes that documents are 13 years ago
Michael Peter Christen 75dd706e1b update to HostBrowser: 13 years ago
Michael Peter Christen 0716a24737 added more / all new crawl profile fields into crawl profile editor 13 years ago
Michael Peter Christen 4a14122ba7 in case that a crawl profile has a collection assigned, use the 13 years ago
Michael Peter Christen 0fe8be7981 enhaced data structures for balancer and latency computation which 13 years ago
Michael Peter Christen ac9540dfb6 removed options for stopwords which are not used 13 years ago
Michael Peter Christen b2ffd49817 less latency 13 years ago
Michael Peter Christen 0833937c1c better balancing and duetime-cumputation also for no-delay intranet 13 years ago
Michael Peter Christen c326aa8f67 disabled writing new entries to crawl stacks to prevent that a domain 13 years ago
Michael Peter Christen c25d7bcb80 - added concurrency for robots.txt loading 13 years ago
Michael Peter Christen a87811bc38 more auto-commit calls when a search interface is opened, but not when a 13 years ago
Michael Peter Christen 2d9e577ad0 replaced the custom robots.txt loader by the standard http loader 13 years ago
Michael Peter Christen a33e2742cb - removed unnecessary synchronized and deadlock in crawler 13 years ago
orbiter 8952153ecf update to Balancer algorithm: 13 years ago
Michael Peter Christen 85ca07b90e when a new crawl is started, an equal crawl, if still running, is 13 years ago
Michael Peter Christen ae6feb5610 showing the web structure graph as animation in the crawl monitor 13 years ago
Michael Peter Christen ccc3760a47 Refactoring and redesign of data architecture to make URIMetadataRow 13 years ago
Michael Peter Christen e5b3c172ff removed hack which translated Solr documents to virtual RWI entries 13 years ago
Michael Peter Christen 43f3345c90 - removed dependencies from URIMetadataRow and made direct access to 13 years ago
Michael Peter Christen 21fe8339b4 - enhanced generation of url objects 13 years ago
Michael Peter Christen 5f0ab25382 removed the option to prevent removal of & parts inside of the 13 years ago
Michael Peter Christen 53789555b9 fix for crawl start filter 13 years ago
Michael Peter Christen a06930662c replaced some more .getBytes() with UTF8/ASCII.getBytes() 13 years ago
Michael Peter Christen 4b5e0c1500 added an url rewriter which can be used to remove session ids from urls 13 years ago
Michael Peter Christen 76d218fbef fixes to crawl profiles 13 years ago
Michael Peter Christen 1533bfd63b refactoring 13 years ago
Michael Peter Christen 872f83ebe0 refactoring 13 years ago
Michael Peter Christen 8219a445f3 refactoring 13 years ago
Michael Peter Christen f879a344e7 fix for no depth limit default value 13 years ago
Michael Peter Christen 00c1c777fa refactoring 13 years ago