Commit Graph

66 Commits (d64d45361cbe9e6898c3d02430dfac2d737dbee5)

Author SHA1 Message Date
Michael Peter Christen b28d43decc added two more fields source_cr_host_norm_i,target_cr_host_norm_i in 12 years ago
Michael Peter Christen a52f3a597e fix for canonical-from-http-header feature 12 years ago
Michael Peter Christen 2dd7c5be44 added parsing of http-canonical tags (untested, could not find an 12 years ago
Michael Peter Christen 4f83d5f18c added the new field harvestkey_s to the collection index and the 12 years ago
Michael Peter Christen 96ed0c980e - added hosthash to all documents (also fail documents which is needed 12 years ago
orbiter c64b51134e hack to add all tokens from the url to text_t. This was working for the 12 years ago
Michael Peter Christen 2602be8d1e - removed ZURL data structure; removed also the ZURL data file 12 years ago
Michael Peter Christen 31920385f7 set anchor rel attribute of all links to "nofollow" if the html meta 12 years ago
Michael Peter Christen 61c5e40687 - replaced the properties object in AnchorURL with distinct variables 12 years ago
Michael Peter Christen 5e31bad711 - the webgraph shall store all links which appear on a web page and not 12 years ago
Michael Peter Christen 35ab2cef7b added parsing of 'date', 'dc:date', 'dc.date' and 'last-modified' in 12 years ago
Michael Peter Christen 85456f46b2 added two new fields, exact_signature_copycount_i and 12 years ago
Michael Peter Christen a2511b5600 turned images_alt_txt back to images_alt_sxt because it is not necessary 12 years ago
Michael Peter Christen 5a0de1b77d moving image description text to image text field 12 years ago
orbiter f106345eef link strings should not be tokenized 12 years ago
orbiter deadeb406e image alt tag strings should be tokenized 12 years ago
Michael Peter Christen a88a62f7aa added a feature to set a collection for a crawl result based on a 12 years ago
Michael Peter Christen 47b1c81d08 - refactoring 12 years ago
Michael Peter Christen 697613170d less logging for postprocessing (this was a debugging logging with high 12 years ago
Michael Peter Christen 2857499467 fix to collection schema; bug appeared for _txt fields with empty String 12 years ago
Michael Peter Christen 58fe986cca Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git 12 years ago
Michael Peter Christen cf12835f20 replaced the single-text description solr field with a multi-value 12 years ago
Roland Haeder 0343f0668c Fix for NPE: 12 years ago
orbiter 080d80c9de do not write an empty failreason in case that there is no fail. Because 12 years ago
orbiter 6fb2811e68 fixes for problems with remote solr and non-activated webgraph index 12 years ago
orbiter c124037f19 removed forced non-soft commits to prevent index fragmentation 12 years ago
Roland Haeder 841a28ae76 Added 'final' for all exception blocks as this helps the Java compiler 12 years ago
Michael Peter Christen 89c0aa0e74 added collection_sxt to error documents 12 years ago
orbiter a9c8046c87 do a light optimization at the end of a crawl postprocessing 12 years ago
orbiter a548354c71 replaced type of solr schema object sku of text_en_splitting_tight by 12 years ago
orbiter 2f1ec8d4a2 npe fix 12 years ago
Michael Peter Christen 5878c1d599 - refactoring of log to ConcurrentLog: 12 years ago
Michael Peter Christen 57ffdfad4c added a crawl option to obey html-meta-robots-noindex. This is on by 12 years ago
Michael Peter Christen 5a5d411ec0 new robots_i attribute fields 12 years ago
Michael Peter Christen e6f361f474 adding the canonical tag to crawl queues 12 years ago
Michael Peter Christen 203921006a redesign of citation index storage 12 years ago
Michael Peter Christen 823ae4d6a7 added url_protocol_s to error documents 12 years ago
Michael Peter Christen 9a6fcdf597 npe fix 12 years ago
Michael Peter Christen 16d1d744fa added url_file_name_s in default collection schema for the file name 12 years ago
Michael Peter Christen f9d859f5dc now writing image alt texts and (camelcase-)parsed urls into a text 12 years ago
orbiter 8792e6c6e9 stub for better image indexing 12 years ago
Michael Peter Christen 570511f3c8 removed fields references_internal_id_sxt and 12 years ago
Michael Peter Christen f7e77a21bf Added a citation reference computation for intra-domain link structures. 12 years ago
Michael Peter Christen a1644ca0fd new workflow processor in Segment to enqueue indexing documents to solr 12 years ago
Michael Peter Christen cca19d94d4 re-declared some fields to be of type string rather than text which 12 years ago
Michael Peter Christen ea85674be2 added the date to error documents 12 years ago
Michael Peter Christen 3502b4c697 refactoring (renaming) of yacy-solr api 12 years ago
Michael Peter Christen c091000165 added collection attribute also to the rss feed reader 12 years ago
Michael Peter Christen 50421171c3 added new schema fields: 12 years ago
Michael Peter Christen 7ab5093321 added new solr title_exact_signature_l and 12 years ago