Commit Graph

9973 Commits (c152d996e6404cd67e1d80ad363aca1780149f3d)
 

Author SHA1 Message Date
Michael Peter Christen 3bf0104199 fix for crawl domain counter limitation (limit was reached too early)
12 years ago
Michael Peter Christen 82bfd9e00a - crawl profiles shall be deleted from active and passive stacks if they
12 years ago
Michael Peter Christen 1b3d26dd23 hack to remove most of the warning: deprecated messages (but not all,
12 years ago
Michael Peter Christen a496313248 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git
12 years ago
sixcooler 3c48fc65fd reverted RemoteInstance to deprecated methods of httpClient-4.2
12 years ago
Michael Peter Christen 91a875dff5 self-healing of mistakenly deactivated crawl profiles. This fixes a bug
12 years ago
Michael Peter Christen 095053a9b4 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git
12 years ago
sixcooler 0cae420d8e some dns-timing changes:
12 years ago
sixcooler 15b1bb2513 bump to httpClient-4.3
12 years ago
Michael Peter Christen 4f83d5f18c added the new field harvestkey_s to the collection index and the
12 years ago
orbiter 14442efa6d when profiles are cleaned, there shall be first a callback showing which
12 years ago
orbiter 0013d0d0bb removed superfluous class
12 years ago
orbiter f90d5296cb Added new data structure to be used by the balancer (not used yet).
12 years ago
orbiter 0e8d752462 refactoring
12 years ago
orbiter 8ac2e8c8c9 added location navigator which causes that the image to the map search
12 years ago
orbiter d86d2be5c3 automatically removed Places autotagging if no location library is
12 years ago
orbiter 214a087cdf Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git
12 years ago
Michael Peter Christen 96ed0c980e - added hosthash to all documents (also fail documents which is needed
12 years ago
Michael Peter Christen 179ad281f9 close include byte buffer after usage
12 years ago
reger 6b9a624808 remove double declaration of TLD_any_zone_filter
12 years ago
orbiter d2effd21db fix for npe during location search
12 years ago
orbiter 828603e4f1 fix for 100%CPU problem in error cache cleaning process
12 years ago
orbiter c64b51134e hack to add all tokens from the url to text_t. This was working for the
12 years ago
orbiter 6e8377b8ad do not check all words with synonym library if the library is empty
12 years ago
orbiter 70ba74b23a disabled ipv4 preference to enable ipv6-only networks like freifunk
12 years ago
orbiter f3be1930cb CPU problem when pusing to the error cache; wrong class,
12 years ago
Michael Peter Christen e40671ddb7 better and consistent deletions for error urls
12 years ago
Michael Peter Christen 2602be8d1e - removed ZURL data structure; removed also the ZURL data file
12 years ago
Michael Peter Christen 31920385f7 set anchor rel attribute of all links to "nofollow" if the html meta
12 years ago
Michael Peter Christen 57e00baf26 fix for parsing of image links inside of anchor links (image-links)
12 years ago
Michael Peter Christen 61c5e40687 - replaced the properties object in AnchorURL with distinct variables
12 years ago
Michael Peter Christen 3ea9bb4427 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git
12 years ago
Michael Peter Christen 5e31bad711 - the webgraph shall store all links which appear on a web page and not
12 years ago
reger 603368fc3e remove redundant declaration of USER_AGENT
12 years ago
Michael Peter Christen 1a8c64117f decreased the responseHeaderDB database which is now flushed more
12 years ago
Michael Peter Christen 3e22d05290 added option for daterange properties in GSA interface to use an left-
12 years ago
reger fe87fb638a adjust test/ParserTest to dc_description data type
12 years ago
Michael Peter Christen 35ab2cef7b added parsing of 'date', 'dc:date', 'dc.date' and 'last-modified' in
12 years ago
Michael Peter Christen 9cc8468b30 added tools to visualize image generation (i.e. during testing)
12 years ago
Michael Peter Christen dbef8ccfcb forced deletion of ZURL entries for a specific host for each host that
12 years ago
Michael Peter Christen e137ff4171 refactoring (im preparation for new removeHost method)
12 years ago
Michael Peter Christen 7a5574cd51 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git
12 years ago
Michael Peter Christen 85456f46b2 added two new fields, exact_signature_copycount_i and
12 years ago
orbiter 26366596d9 fix for a problem which ocurres when a site is crawled where the start
12 years ago
Michael Peter Christen a2511b5600 turned images_alt_txt back to images_alt_sxt because it is not necessary
12 years ago
Michael Peter Christen 85b1922244 activated image type navigation for image search
12 years ago
Michael Peter Christen 9e12fdff23 Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git
12 years ago
Michael Peter Christen ab1201fdfd fixed wrong facet count
12 years ago
Michael Peter Christen 049c3b3f2e added an option to exclude image search results from text search. This
12 years ago
Michael Peter Christen 69f85265e1 added an option to put image links to the crawl queue and handle these
12 years ago