Commit Graph

7245 Commits (05d58e4df00fdecac3f785878e54c3fefd6bc136)

Author SHA1 Message Date
orbiter f15c832587 Merge branch 'master' of git@gitorious.org:yacy/rc1.git 11 years ago
Marc Nause c97da1a0d8 First draft of a blacklist API. 11 years ago
Michael Peter Christen d4f65833a1 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 11 years ago
Michael Peter Christen c1c1be8f02 fix for slow crawling and better logging in balancer 11 years ago
Michael Peter Christen 3acf416335 npe fix 11 years ago
reger 2eb7682772 add html5 audio/video <source> tag to html content scraper 11 years ago
reger 0b6db04e40 fix contentscraper img height/width parsing 11 years ago
reger ffc5b75c73 optimize and fix lat / lon assignment 11 years ago
reger 9313447de2 reimplement tighter lat/lon calc in URIMetadataNode 11 years ago
reger d812f80784 add exit proxy link to UrlProxy 11 years ago
reger 78d08998db throw MalformedURLException on unknown protocol 11 years ago
reger bb8181b2be fix: resolve url without path but searchpart 11 years ago
orbiter a3542f29b4 npe fix 11 years ago
orbiter c48d2a2a02 npe fix 11 years ago
reger 121d25be38 recover sax fatal error on OAI-PMH import of xml with entity error 11 years ago
reger 81dc2aa536 add current css to HTMLResponseWriter to fix metadata view 11 years ago
orbiter 2fd8a0ead6 Merge branch 'master' of git@gitorious.org:yacy/rc1.git 11 years ago
orbiter 8e5ce7cd51 fixed a situation where finished crawls had not been detected. 11 years ago
orbiter 2f63bd0261 enhanced Host Balancer strategy: fair round robin 11 years ago
orbiter 0c88a32c36 do not apply lazy value instantiation for numeric or boolean values 11 years ago
orbiter 8e04030596 in case of short memory, do not cut down robinson peers to 1, just 11 years ago
reger 86f6975edc exclude html tags in in/outboundlinks_anchortext_txt parsed text 11 years ago
orbiter ccb1864d55 catch IllegalArgumentException for wrong process types (that is needed 11 years ago
orbiter 4ee4ba1576 fix for NPE in IndexCreateParserErrors_p.html caused by bad handling of 11 years ago
orbiter 12ba890205 removed warnings 11 years ago
reger d51f9cc863 add custom Jetty errorhandler 11 years ago
reger c193a02023 defer creation of new ArrayList after possible early return 11 years ago
reger 727dfb5875 refactore URIMetadataNode to further unify interaction with index 11 years ago
reger 79e7947442 - remove empty http0_9 status text array 11 years ago
reger 2dabe2009d - remove unused manual http KeepAlive config 11 years ago
Michael Peter Christen 5746aae3db add canonical links to the same crawldepth, not the next crawldepth 11 years ago
Michael Peter Christen 74ab5ef9fa increased runtime for postprocessing query job 11 years ago
Michael Peter Christen 8b32dd5f9e special strategy for balancer: do not remove targets with zero wait time 11 years ago
Michael Peter Christen 9c6228d948 fix for deadlocks in crawler 11 years ago
Michael Peter Christen 10cf8215bd added crawl depth for failed documents 11 years ago
Michael Peter Christen 7fefebaeca Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 11 years ago
Michael Peter Christen c2f62e783f - better subgraph handling, less overhead for crawls without the 11 years ago
Michael Peter Christen 06afb568e2 new Strategies in Balancer: 11 years ago
Michael Peter Christen 1aea01fe5b fix for Table in case that requested file does not exist and paths also 11 years ago
reger 710054bb37 implement gzip input handling directly in defaultservlet 11 years ago
Michael Peter Christen 9a5ab4e2c1 removed clickdepth_i field and related postprocessing. This information 11 years ago
Michael Peter Christen da86f150ab - added a new Crawler Balancer: HostBalancer and HostQueues: 11 years ago
Michael Peter Christen 075b6f9278 refactoring of the crawl balancer: the balancer is turned into an 11 years ago
Michael Peter Christen 8470dfe3f8 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 11 years ago
reger 46016fa153 autoupdate fails to download latest release (1.71) due to default release blacklist 11 years ago
Michael Peter Christen 8aeef73d49 fix for virtual root nodes 11 years ago
Michael Peter Christen 7c7fbb9818 find depth-matches also for edge targets 11 years ago
Michael Peter Christen dd12dd392f introduction of a data structure for HyperlinkEdges which should use 11 years ago
Michael Peter Christen 6ea8bb7348 using MultiProtocolURL for edge data which is faster (hash computation 11 years ago
Michael Peter Christen b21c208b4d enhanced hashcode computation for MultiProtocolURL 11 years ago
Michael Peter Christen ce1d1b2fa0 fix for maximum tag length in parser 11 years ago
Michael Peter Christen 17e0956312 refactoring of SystemLoad calls (only one backend tool) 11 years ago
Michael Peter Christen a37d067692 refactoring 11 years ago
orbiter 95780eed32 Merge branch 'master' of git@gitorious.org:yacy/rc1.git 11 years ago
Michael Peter Christen 67beef657f strong redesign of html parser: object recursion is now made using a 11 years ago
Michael Peter Christen 6bd8c6f195 fix for wrong status codes of error pages 11 years ago
Michael Peter Christen 9e503b3376 also delete the robots.txt file from the cache when a new crawl is 11 years ago
orbiter 67501c9dda Merge branch 'master' of git@gitorious.org:yacy/rc1.git 11 years ago
Michael Peter Christen 1c21b3256d fix for robots.txt handling: delete old entry before starting a new 11 years ago
orbiter c250fac9f4 linkstructure refactoring to get more options for clickdepth analysis 11 years ago
Michael Peter Christen 8068e68474 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 11 years ago
Michael Peter Christen bd886054cb new structure and enhancements for link graph computation: 11 years ago
reger f326a67561 fix: typo in default charset in metadata2solr 11 years ago
Michael Peter Christen df138084c0 do solr optimization independently from memory and load constraints: 11 years ago
Michael Peter Christen ebd44a7080 replaced solr 4.6.1 with solr 4.7.1 and added index migration to 11 years ago
Michael Peter Christen 734778c0c8 fixed a time-out problem in the default servlet which is also a logging 11 years ago
Michael Peter Christen 466d90ad42 fixed a problem with resource observer; probably coming from uncatched 11 years ago
Michael Peter Christen e8ddd415a8 enhanced the new link structure graph 11 years ago
Michael Peter Christen 926d28dd3f fixed a bug which prevented crawl starts after a network switch 11 years ago
Michael Peter Christen 3ce8eff21b another fix for inbound/outbound detection 11 years ago
Michael Peter Christen d4b5c457e4 NPE fix 11 years ago
Michael Peter Christen 36a66b0704 fix for parsing of numeric value in case that boolean values are given 11 years ago
orbiter 41730c8048 better logging in template engine: shows filename of servlets where 11 years ago
orbiter 3c1274057d fixed thread dump in case of wrong seeds 11 years ago
orbiter 18f9c40302 moved Edge class out of linkstructure servlet as this does not work on 11 years ago
orbiter de95e5e524 reduced search activity corona strength in network image 11 years ago
reger da413af664 move baseurl after parsing orig source in urlproxyservlet 11 years ago
reger af6ad20728 fix: remove obsolete ref to yacy.home 11 years ago
Michael Peter Christen 74ab094587 fix for solr query size; too many documents had been retrieved in case 11 years ago
Michael Peter Christen c64c10ef00 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 11 years ago
Michael Peter Christen 48fbfa60c1 bugfix to inbound/outbound identification 11 years ago
reger 227c42bc96 eleminate obsolete URIMetaDataRow class 11 years ago
Michael Peter Christen cca851a417 introduced new solr field crawldepth_i which records the crawl depth of 11 years ago
orbiter b1ba764d81 fix for first start options and added german translation for popup texts 11 years ago
orbiter 429a874222 - added COLS field in GSA response (non-gsa standard by customer 11 years ago
Michael Peter Christen 1b9ec9a1c5 - added popover to p2p/stealth mode button to explain the peer mode and 11 years ago
Michael Peter Christen 62a36fa584 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 11 years ago
reger c9f92abddc fix: application link count 11 years ago
Michael Peter Christen a267c46e1a Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 11 years ago
Michael Peter Christen 5b83887da8 npe fix 11 years ago
Michael Peter Christen 63c9fcf3e0 free configuration of postprocessing clickdepth maximum depth and time 11 years ago
Michael Peter Christen 39b641d6cd added tutorial mode - some menu items will only appear if you 'qualify' 11 years ago
sixcooler f06775850f fix receiving DHT / parse pultipart 11 years ago
reger 49e76a1c55 make use of detected charset in htmlParser if none is given. 11 years ago
reger e11504309f adding a hint to javascript browser short cut on Url-Proxy page (AugmentedBrowsing_p.html) 11 years ago
reger b12200cafe alternative UrlProxyServlet (for /proxy.html) using different url rewrite rules 11 years ago
reger 2953ebe701 fix: port in local target adress 11 years ago
Michael Peter Christen fda591695c fixed visibility of custom icon 11 years ago
Michael Peter Christen a9b9950d7f Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 11 years ago
Michael Peter Christen b488f33975 added close to fix possible resource leak warning 11 years ago
Michael Peter Christen 56710ecb26 prevent opening of new files as that could be a cause for the latest 11 years ago
Michael Peter Christen 8b44fcf0f4 added missing @Override annotation 11 years ago
reger d7055904a6 fix: proxyservlet path header setting 11 years ago
Michael Peter Christen e515dd460d added linkscount_i and linksnofollowcount_i to the default solr schema 11 years ago
Michael Peter Christen 1a764135be one more Thread Dump fix for new bootstrap css style 11 years ago
Michael Peter Christen bb21d825f9 fix for thread dump line spacing 11 years ago
Michael Peter Christen cbdfef7ce1 changed protocol facet to show also all other counts if one facet is 11 years ago
reger b9056ef2db remove unused private header entries (HeaderFramework) 11 years ago
sixcooler 6d16fa993d make transparent proxy handle https-connections: 11 years ago
Michael Peter Christen 61ad194065 fix for source and target clickdepth in webgraph index 11 years ago
Marc Nause 809b4e1fd9 Team added support for URLs with unicode characters in host part to 11 years ago
reger b126b9ba17 add some InputFileStream close at end of reads 11 years ago
reger ca7444dbdf limit filetype nav to known extension also on image/media search 11 years ago
reger 651d057e93 surrogate import translate dc:language 3-char codes 11 years ago
orbiter 22618e3ba2 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 11 years ago
orbiter 01989f6af9 restrict write buffer size to a limit 11 years ago
Michael Peter Christen d1091e79f8 - added stealth button to navigation menu 11 years ago
reger c297de5145 remove check for unused virtual path /currentyacypeer/ 11 years ago
orbiter 3c8d6e1eee added adminAccount switch to ConfigAccounts_p servlet to switch on 11 years ago
orbiter 7d24bcb98d added flag to require that all web pages, even such without a "_p" 11 years ago
Michael Peter Christen 7a6658abec removed synchronization in embedded solr connection (that was probably 11 years ago
Michael Peter Christen a7d4379ef9 fixed shutdown of solr cores in case that more than one local core is to 11 years ago
Michael Peter Christen 453bfd0f17 removed unused variables and warnings 11 years ago
Michael Peter Christen 05655d98df Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 11 years ago
reger 9f02d2c47b fix: remove link to triplestore in Vocabulary_p (triplestore does not longer exist) 11 years ago
reger 81a846ec33 fix: set YaCy CONNECTION_PROP_HOST Header in ProxyServlet to host incl. port 11 years ago
reger 251be9ecfa remove unused ProxySettings ref. from loader 11 years ago
reger 82dc815af9 cleanup: remove unrelated and unused code 11 years ago
Michael Peter Christen 85a427ec54 support for multiple sitemaps in robots.txt 11 years ago
reger a373fb717d remove more unused from legacy server.http 11 years ago
reger 749d020aeb remove redundant url string manipulation in HTTPDProxyHandler 11 years ago
reger 612294cf84 use servletPath in ProxyServlet instead of fixed name 11 years ago
reger 1d01672bd3 fix DCEntry.getIdentifier 11 years ago
Michael Peter Christen b08375da33 fix for bad/missing values of size_i 11 years ago
reger 6306d28a6a OAI import get multivalued keywords (dc:subject) 11 years ago
reger 0a8c8102de allow YaCy to start w/o ssl if JKS init fails 11 years ago
sixcooler 0b2101c59c Speed up the ProxyHandler: 11 years ago
reger 516f8c2489 fix: to allow unix scripts (bin/*.sh) to allways submit http admin apicalls 11 years ago
Michael Peter Christen ea3aa30593 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 11 years ago
reger dd5bf0b71b cleanup old reference to HTTPDemon.setAlternativeResolver 11 years ago
Michael Peter Christen 51800007c4 - added concurrency to postprocessing of webgraph document 11 years ago
Michael Peter Christen 5f4a6892c1 enhanced RowSet re-sort limit for small sets 11 years ago
reger 351c2be68d fix: make sure adminAccount changes made via ConfigAccounts_p are effective immediately 11 years ago
reger 5c9dcc269d improve OAI-PMH import identifier recognition 11 years ago
Michael Peter Christen 0e7d249a69 fixed another shutdown problem (only occurs if webgraph core is enabled) 11 years ago
Michael Peter Christen e485fbd0ce - let crawl loader jobs die after 10 seconds without new jobs 11 years ago
Michael Peter Christen bcd9dd9e1d enhanced concurrent loading by using a fixed set of concurrent loader 11 years ago
orbiter 051328271c bugfix-bugfix 11 years ago
orbiter eedcbcd906 bugfix to proxy handler: recognize the own yacyh-host 11 years ago
orbiter d68e5ad0c4 NPE fix for Thread name (just commited yesterday, sorry) 11 years ago
reger 6878c90f99 fix: IPv6 INTRANET_PATTERNS for local ip (see http://bugs.yacy.net/view.php?id=378) 11 years ago
reger a2e5ea2026 status panel link to set max mem 11 years ago
Michael Peter Christen 6ed9c0164e attaching names to all Threads to get a better view in profiling tools 11 years ago
Michael Peter Christen fdaeac374a - enhanced postprocessing speed and memory footprint (by using HashMaps 11 years ago
reger ba49ff81ed little more verbose proxy 403 error message 11 years ago
Michael Peter Christen d325cb8912 fixes and enhancements for postprocessing 11 years ago
Michael Peter Christen 7c1b968378 another fix for the shutdown exceptions 11 years ago
orbiter 133d41386c (again) full redesign of ConcurrentUpdateSolrConnector to remove 11 years ago
Michael Peter Christen a632b0d2a4 added a forced commit to index deletion to enable synchronized index 11 years ago
Michael Peter Christen 1d069c5861 make sure that postprocessed documents are overwritten 11 years ago
Michael Peter Christen 0d2342575e Merge branch 'master' of ssh://gitorious.org/yacy/rc1 11 years ago
Michael Peter Christen 3cc5c0ffdd a concurrency enhancement which was not used because tests showed worse 11 years ago
Michael Peter Christen e644981697 added one more postprocessing low memory check 11 years ago
reger 5e645f4449 Merge origin/master 11 years ago
reger 3b89176b9f use config value htroot in Jetty init (was hardcoded) 11 years ago
Michael Peter Christen e1bf65c892 added short memory protection during postprocessing 11 years ago
Michael Peter Christen 90b47e83e6 fixed shutdown error when closing solr connectors 11 years ago
Michael Peter Christen 7640834b37 removed double concurrency to put Solr documents into the index. The 11 years ago
Michael Peter Christen 0f6b72f24b do not use luke requests for remote solr servers if the result is 11 years ago
Michael Peter Christen c57026e242 recover from OOM 11 years ago
Michael Peter Christen 907db8b7a6 fix for bad query shortcut hack 11 years ago
Michael Peter Christen a2b66fe2eb Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 11 years ago
Michael Peter Christen 9f6be762a6 - better logging for postprocessing 11 years ago
orbiter da5d4128bf prevent npe 11 years ago
orbiter a878c7982c prevent npe 11 years ago
orbiter e4eb87d924 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 11 years ago
orbiter ced1a96f9c fixed error cache 11 years ago
reger 3ba81bd08a Merge origin/master 11 years ago
reger 4d896383db fix: use timeout = proxy.ClientTimeout in ProxyHandler 11 years ago
orbiter cfb647db6e - introduced a miss cache in ConcurrentUpdateSolrConnector 11 years ago
orbiter a87d8e4a8e changed caching of ConcurrentUpdateSolrConnector: it caches now also the 11 years ago
orbiter f6e441dd77 refactoring 11 years ago
orbiter 76c53faeb2 removed unused code (HostStat) 11 years ago
orbiter d3a88eaecb introducing ConcurrentUpdateSolrServer for remote solr servers. 11 years ago
reger 809e976578 remove unused java imports form yacy.java 11 years ago
reger a9b06f8719 add a -config command line parameter e.g. -config "port=9090" "port.ssl=8043" 11 years ago
reger 0923b09216 fix: allow 4 character admin user name 11 years ago
Michael Peter Christen 254a7ac66c fixed cleaning of index 11 years ago
Michael Peter Christen 28a7b42e6b removed warning "sun.misc.BASE64Encoder is internal proprietary API and 11 years ago
Michael Peter Christen 046f5a03cb one more SolrIndexSearcher bugfix 11 years ago
sixcooler 78c01b3eff fix for 'AlreadyClosedException: this IndexReader is closed' 11 years ago
Michael Peter Christen 1b5e3d523a better control over close-state of remote solr connections 11 years ago
Michael Peter Christen 1a364572a5 fix for 11 years ago
Michael Peter Christen 69391e5d9e changed strategy to test existence of documents in Solr: using the 11 years ago
Michael Peter Christen 790f103f32 delete fail-docs during postprocessing to prevent that they will appear 11 years ago
Michael Peter Christen ff656ce860 explicit call to optimize to add a expungeDeleted flag 11 years ago
Michael Peter Christen 9eb668e951 enhanced the resource observer 11 years ago
Michael Peter Christen fbee98c06f fixed shortcut self-reference bug 11 years ago
Michael Peter Christen e7a29a2851 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 11 years ago
Michael Peter Christen bf97e38b83 removed clearURLIndex, which is a stub remaining from the old metadata 11 years ago