Commit Graph

733 Commits (da380343c2bc9c673900fc28f568474c20cf2fbb)

Author SHA1 Message Date
Michael Peter Christen 1b61bd40ed - Added new solr field url_file_name_tokens_t which stores the file name 12 years ago
orbiter 5f5a97bafc added the anchor text within web pages to the searcheable entities of a 12 years ago
orbiter 705b3338ee list more fields available for search and for ranking boosts 12 years ago
Michael Peter Christen 78e7aadb26 removed unused initialization method 12 years ago
Michael Peter Christen 4fbc4740df removed warnings 12 years ago
Michael Peter Christen 21aa6a0321 migration to Solr 4.5.0 12 years ago
Michael Peter Christen 101a6e6e14 Patch the citation index for links with canonical tags. 12 years ago
Michael Peter Christen b28d43decc added two more fields source_cr_host_norm_i,target_cr_host_norm_i in 12 years ago
Michael Peter Christen a52f3a597e fix for canonical-from-http-header feature 12 years ago
Michael Peter Christen 2dd7c5be44 added parsing of http-canonical tags (untested, could not find an 12 years ago
Michael Peter Christen 3bf0104199 fix for crawl domain counter limitation (limit was reached too early) 12 years ago
Michael Peter Christen 82bfd9e00a - crawl profiles shall be deleted from active and passive stacks if they 12 years ago
Michael Peter Christen 91a875dff5 self-healing of mistakenly deactivated crawl profiles. This fixes a bug 12 years ago
Michael Peter Christen 095053a9b4 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 12 years ago
sixcooler 0cae420d8e some dns-timing changes: 12 years ago
Michael Peter Christen 4f83d5f18c added the new field harvestkey_s to the collection index and the 12 years ago
orbiter 14442efa6d when profiles are cleaned, there shall be first a callback showing which 12 years ago
orbiter 8ac2e8c8c9 added location navigator which causes that the image to the map search 12 years ago
Michael Peter Christen 96ed0c980e - added hosthash to all documents (also fail documents which is needed 12 years ago
orbiter 828603e4f1 fix for 100%CPU problem in error cache cleaning process 12 years ago
orbiter c64b51134e hack to add all tokens from the url to text_t. This was working for the 12 years ago
orbiter f3be1930cb CPU problem when pusing to the error cache; wrong class, 12 years ago
Michael Peter Christen e40671ddb7 better and consistent deletions for error urls 12 years ago
Michael Peter Christen 2602be8d1e - removed ZURL data structure; removed also the ZURL data file 12 years ago
Michael Peter Christen 31920385f7 set anchor rel attribute of all links to "nofollow" if the html meta 12 years ago
Michael Peter Christen 61c5e40687 - replaced the properties object in AnchorURL with distinct variables 12 years ago
Michael Peter Christen 5e31bad711 - the webgraph shall store all links which appear on a web page and not 12 years ago
Michael Peter Christen 35ab2cef7b added parsing of 'date', 'dc:date', 'dc.date' and 'last-modified' in 12 years ago
Michael Peter Christen 9cc8468b30 added tools to visualize image generation (i.e. during testing) 12 years ago
Michael Peter Christen dbef8ccfcb forced deletion of ZURL entries for a specific host for each host that 12 years ago
Michael Peter Christen e137ff4171 refactoring (im preparation for new removeHost method) 12 years ago
Michael Peter Christen 7a5574cd51 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 12 years ago
Michael Peter Christen 85456f46b2 added two new fields, exact_signature_copycount_i and 12 years ago
orbiter 26366596d9 fix for a problem which ocurres when a site is crawled where the start 12 years ago
Michael Peter Christen a2511b5600 turned images_alt_txt back to images_alt_sxt because it is not necessary 12 years ago
Michael Peter Christen 85b1922244 activated image type navigation for image search 12 years ago
Michael Peter Christen 9e12fdff23 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 12 years ago
Michael Peter Christen ab1201fdfd fixed wrong facet count 12 years ago
Michael Peter Christen 049c3b3f2e added an option to exclude image search results from text search. This 12 years ago
Michael Peter Christen 69f85265e1 added an option to put image links to the crawl queue and handle these 12 years ago
Michael Peter Christen a8c5bfcf58 avoid to create unnecessary objects 12 years ago
Michael Peter Christen 5a0de1b77d moving image description text to image text field 12 years ago
Michael Peter Christen dc179bd61f fix for catchall query goal for image search 12 years ago
reger 392174de8c remove all_words, all_strings lists from QueryGoal 12 years ago
Michael Peter Christen 169ef8963d one more fix for image search 12 years ago
Michael Peter Christen cb85b22725 redesign of the image search process (with much better results, 12 years ago
reger 29967102a2 optimized QueryGoal (reducing mem and computation by removing all_hashes) 12 years ago
orbiter f106345eef link strings should not be tokenized 12 years ago
orbiter deadeb406e image alt tag strings should be tokenized 12 years ago
Michael Peter Christen 1a3e42eca4 index migration to lucene 4.4 12 years ago
Michael Peter Christen a88a62f7aa added a feature to set a collection for a crawl result based on a 12 years ago
Michael Peter Christen 765943a4b7 Redesign of crawler identification and robots steering. A non-p2p user 12 years ago
Michael Peter Christen 47b1c81d08 - refactoring 12 years ago
Michael Peter Christen 697613170d less logging for postprocessing (this was a debugging logging with high 12 years ago
reger a5019bc470 make Vocabulary Navigator tags a hard result entry filter 12 years ago
reger a67a4b7d86 improve tld: query modifier filter pattern (to prevent tld:net accepting www.abcinet.org) 12 years ago
reger 02fe8b43ba Field Re-Indexing: display list of fields in reindex queue 12 years ago
sixcooler 7f501b7c38 clear some caches before reporting low Memory 12 years ago
Michael Peter Christen 2857499467 fix to collection schema; bug appeared for _txt fields with empty String 12 years ago
Michael Peter Christen 58fe986cca Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 12 years ago
Michael Peter Christen cf12835f20 replaced the single-text description solr field with a multi-value 12 years ago
reger f2d99053ed Field Re-Indexing: prevent endless error loop in ReindexSolrBusyThread on Solr exception (by skipping query causing the exception) 12 years ago
orbiter d05e0c5368 wait a bit longer before doing the first peer ping 12 years ago
orbiter b8f57f7703 don't be noisy when doing background tasks that may be allowed to fail 12 years ago
Roland Haeder 0343f0668c Fix for NPE: 12 years ago
Roland Haeder b58ca8622d Some cleanups: 12 years ago
Roland Haeder 7263bb82fb Fix for NPE on shutdown: 12 years ago
orbiter 080d80c9de do not write an empty failreason in case that there is no fail. Because 12 years ago
Michael Peter Christen 61e015268b fix in forced deletion: forced commit needed 12 years ago
Michael Peter Christen c3b2301b2f fix for http://bugs.yacy.net/view.php?id=268 12 years ago
orbiter 3e901dcb06 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 12 years ago
orbiter f50b596e0b do not run dht ditribution if system load is over 2.5 12 years ago
orbiter 056b42f5aa - added information about segment count to status_p.xml 12 years ago
orbiter 6fb2811e68 fixes for problems with remote solr and non-activated webgraph index 12 years ago
sixcooler af740f3058 changed optimization to a segment-size of index-size/5.000.000 12 years ago
orbiter 5364c4dcc9 delayed first peer-ping to send the first ping out after the http got 12 years ago
orbiter e24016e30a added the property federated.service.solr.indexing.timeout to yacy.init 12 years ago
orbiter c124037f19 removed forced non-soft commits to prevent index fragmentation 12 years ago
Michael Peter Christen c15aa758dc removed failreason_t removal patch because that causes too much 12 years ago
Roland Haeder be0ff6018f Removed trailing spaces + some more final 12 years ago
Roland Haeder 841a28ae76 Added 'final' for all exception blocks as this helps the Java compiler 12 years ago
Michael Peter Christen 89c0aa0e74 added collection_sxt to error documents 12 years ago
Michael Peter Christen 0df5195cb0 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 12 years ago
Michael Peter Christen 1fd006cc56 fixes using the embedded connector 12 years ago
orbiter d0dc86cf3d logging of deadlocks (if any) during cleanup process 12 years ago
Michael Peter Christen c6a6f159e8 fix for crawl stack domain counter 12 years ago
Michael Peter Christen 93d1bac140 do a more frequent optimization, reduces IO after optimization 12 years ago
orbiter 290e24564b Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 12 years ago
orbiter 5533fc8e01 fix for bug 260 12 years ago
Michael Peter Christen b79471ee67 grr 12 years ago
Michael Peter Christen a79f288ac1 automatically running optimize on solr if user/search is idle for some 12 years ago
orbiter a9c8046c87 do a light optimization at the end of a crawl postprocessing 12 years ago
orbiter a548354c71 replaced type of solr schema object sku of text_en_splitting_tight by 12 years ago
orbiter 2f1ec8d4a2 npe fix 12 years ago
Michael Peter Christen bcc623a843 refactoring of load_delay: this is a matter of client identification 12 years ago
orbiter 0d0b3a30f5 activate api actions after postprocessing of crawls 12 years ago
orbiter 2be456e7fb added a postprocessing field into api/status_p.xml to show if the 12 years ago
Michael Peter Christen 5878c1d599 - refactoring of log to ConcurrentLog: 12 years ago
Michael Peter Christen a2c8116a8f accept (but ignore) a '+' sign in front of search words 12 years ago
sixcooler d5d8936f9d For indexes that are changing rapidly in NRT situations, fcs (stands for 12 years ago