yacy_search_server

Commit Graph

Author	SHA1	Message	Date
borg-0300	1fbd72f9e0	rename "index.html" to "ndx" git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@1032 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
borg-0300	cd1107d85e	added support for URLs with '?&' git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@1030 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
borg-0300	5fb2b017cb	small change git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@1029 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
theli	b8ceb1ffde	) Adding better https support for crawler - solving problems with unkown certificates by implementing a dummy trust Manager - adding https support to robots-parser - Seed File can now be downloaded from https resources - adapting plasmaHTCache.java to support https URLs properly ) URL Normalization - sub URLs are now normalized properly during indexing - pointing urlNormalForm function of plasmaParser to htmlFilterContentScraper function - normalizing URLs which were received by a crawlOrder request git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@1024 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
borg-0300	a803a509ae	bugfix: port handling in HTCache grogram flow, cleared up git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@1021 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
theli	9a5ab62928	) Adding yacy specific X-YACY-Index-Control header which can be used by clients to disallow yacy to index the response that belongs to the request where X-YACY-Index-Contro is set to "no-index" ) Bugfix for Seed-List download via Remote Proxy. Now the pragma and cache-control http headers of the request are properly set to "no-cache" See: http://www.yacy-forum.de/viewtopic.php?p=11639#11639 *) Bugfix for http-Proxy yacy has ignored "no-cache"- pragma and cache-control http headers that were send in requests. Now, these request headers are evaluated properly TODO: Missing evaluation of "no-store" request headers git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@971 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
borg-0300	58b670201d	now, changed HTCacheSize needs no restart git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@961 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
theli	40777556c5	) Connection Tracking - adding automatic refresh - accepts new parameter nameLookup which can be used to deactivate yacy-peer name lookup (because we have problems with this on large seed-dbs) ) ViewFile New page that can be used to view - original content - plain text content - parsed content - parsed sentences of a webpage specified by there url hash Mainly for debugging purpose at the moment ) Robots.txt Bugfix for if-modified-since usage TODO: synchronization of downloads to avoid loading the same robots-file multiple times in parallel by different threads ) Shutdown Better abortion of transferRWI and transferURL sessions on server shutdown *) Status Page Adding icon to start/stop crawling via status page git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@950 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
theli	a2fa75e688	) Asynchronous queuing of crawl job URLs (stackCrawl) various checks like the blacklist check or the robots.txt disallow check are now done by a separate thread to unburden the indexer thread(s) TODO: maybe we have to introduce a threadpool here if it turn out that this single thread is a bottleneck because of the time consuming robots.txt downloads ) improved index transfer The index selection and transmission is done in parallel now to improve index transfer performance. TODO: maybe we could speed up performance by unsing multiple transmission threads in parallel instead of only a single one. ) gzip encoded post requests it is now configureable if a gzip encoded post request should be send on intex transfer/distribution ) storage Peer (very experimentell and not optimized yet) Now it's possible to send the result of the yacy indexer thread to a remote peer istead of storing the indexed words locally. This could be done by setting the property "storagePeerHash" in the yacy config file - Please note that if the index transfer fails, the index ist stored locally. - TODO: currently this index transfer is done by the indexer thread. To seedup the indexer a) this transmission should be done in parallel and b) multiple chunks should be bundled and transfered together ) general performance improvements - better memory cleanup after http request processing has finished - replacing some string concatenations with stringBuffers - replacing BufferedInputStreams with serverByteBuffer - replacing vectors with arraylists wherever possible - replacing hashtables with hashmaps wherever possible This was done because function calls to verctor or hashtable functions take 3 time longer than calls to functions of arraylists or hashmaps. TODO: we should take a look on the class serverObject which is inherited from hashmap Do we realy need a synchronization for this class? TODO: replace arraylists with linkedLists if random access to the list elements is not needed ) Robots Parser supports if-modified-since downloads now If the downloaded robots.txt file is older than 7 days the robots parser tries to download the robots.txt with the if-modified-since header to avoid unnecessary downloads if the file was not changed. Additionally the ETag header is used to detect changes. ) Crawler: better handling of unsupported mimeTypes + FileExtension ) Bugfix: plasmaWordIndexEntity was not closed correctly in - query.java - plasmaswitchboard.java *) function minimizeUrlDB added to yacy.java this function tests the current urlHashDB for unused urls ATTENTION: please don't use this function at the moment because it causes the wordIndexDB to flush all words into the word directory! git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@853 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
orbiter	7fc822a59b	changed handling of time-zones git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@801 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
orbiter	495bc8bec6	removed cache-control from low and medium priority caches which reduces memory use and computation overhead git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@774 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
orbiter	71a31f0902	integrated and extended new memory performance menu; found and fixed bug in DHT caching git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@752 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
orbiter	fb52a82008	added new performance page for memory settings git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@751 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
borg-0300	8260128ee9	changed getFreeSize(); git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@675 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
borg-0300	0a57fbcde5	Added new HashSet filesInUse; Added new Function getFreeSize(); git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@672 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
borg-0300	da9c6857fb	) changed a misunderstand, no BUG ;) ) finals and other git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@668 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
borg-0300	81cb8feb15	back to 649 :/ git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@651 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
borg-0300	5194511e8e	*) attempt to find bug See: http://www.yacy-forum.de/viewtopic.php?t=1121 git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@650 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
borg-0300	7626823519	BUGFIX for last 'commit' git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@635 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
borg-0300	971756e8dd	the delete size is smaller See: http://www.yacy-forum.de/viewtopic.php?t=1084 git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@634 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
borg-0300	cc493ef8c1	Added change from Hermes See: http://www.yacy-forum.de/viewtopic.php?t=1050 git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@629 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
borg-0300	c1d7527929	better cache cleanup git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@621 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
theli	4fd5b95b1f	*) Renaming Logger function names to reflect the proper Java Logging API Loglevels - please use logFine instead of logDebug - please use logSevere instead of logFailure and logError See: http://www.yacy-forum.de/viewtopic.php?p=8726#8726 git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@615 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
theli	6adf8a4bde	*) Renaming Logger function names to reflect the proper Java Logging API Loglevels - please use logFine instead of logDebug - please use logFailure instead of logError See: http://www.yacy-forum.de/viewtopic.php?p=8726#8726 git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@614 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
theli	cc1df08069	*) Adding missing synchronized blocks git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@608 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
borg-0300	bf14e6def5	) proxyCache, proxyCacheSize can be changed under 'Proxy Indexing' - path now are absolute ) move path check from plasmaHTCache to plasmaSwitchboard - only one path check when starting *) small other git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@606 6c8d7289-2bf4-0310-a012-ef5d649a1542	19 years ago
theli	0c8a48e2cb	*) converting php Session ID to lower case in funktion isCGI See: http://www.yacy-forum.de/viewtopic.php?p=7671#7671 git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@552 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
theli	4654eae4e2	*) adding php Session ID to argument in funktion isCGI See: http://www.yacy-forum.de/viewtopic.php?p=7671#7671 git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@546 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	cd10370992	several bugfixes and dht selection / logging improvement git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@531 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	3610fe6b3a	see http://www.yacy-forum.de/viewtopic.php?p=7410#7410 git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@530 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	f5259f29e8	word cache behaviour fix and other fixes git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@519 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	91163db52e	fix for more time-related problems in proxy git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@486 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	fb6f238d70	fix for expires-problem git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@485 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	e84a177c49	many bigfixes git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@475 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	36707586c7	filtering of jsessionid git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@447 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	ad90f0ad13	activated RWI distribution to DHT for senior peers (default redundancy 3), necessary now for network growth git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@438 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	3470a72d48	fixed div by zero, set default delays, fixed release number format and display git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@435 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	be1f324fca	performance setting for remote indexing configuration and latest changes for 0.39 git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@424 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	c64970fa47	re-implemented proxy-busy-check and fixed some other things git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@421 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	51962d55bf	added 'PPM', page-per-minute statistics git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@405 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	2f0d7ea8d3	removed htcache stati (superfluous now) git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@396 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	419f8fb398	fixed bugs/missing code regarding new crawl stack git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@384 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	858cd94299	replaced indexing ram-queue by file-based stack-queue git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@381 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	1e7f062350	many bugfixes, memory leak fixes, performance enhancements; new kelondroHashtable; activated snippets git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@313 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	a19541e563	code-enhancements after analysis with AppPerfect git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@307 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	4f9c30ef49	using mime-type instead of file extension for doctype git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@269 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
theli	ee9e110366	) removing old logging configuration properties from yacy.init ) serverLog.java logging functions now also accept exceptions als additional parameters. The Stacktrace of this ecceptions will then be appended to the logging message and can e.g. be viewed on the gui logging page git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@265 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
theli	68f30811fa	) changing reference to logger ) bugfix in function getCachePath git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@249 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	ca3b4ccaf4	added snippet-routines (not yet finished) git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@218 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	ee0758fe4d	bugfixes/empty-dir-deletion/snippet-test-activation git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@212 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
theli	361f05978d	Multiple updates regarding the yacy seedUpload facility, optional content parsers, thread pool configuration ... Please help me testing if everything works correct. ) Migration of yacy seedUpload functionality See: http://www.yacy-forum.de/viewtopic.php?t=256 - new uploaders can now be easily introduced because of a new modulare uploader system - default uploaders are: none, file, ftp - adding optional uploader for scp - each uploader provides its own configuration file that will be included into the settings page using the new template include feature - Each uploader can define its libx dependencies. If not all needed libs are available, the uploader is deactivated automatically. ) Migration of optional parsers See: http://www.yacy-forum.de/viewtopic.php?t=198 - Parsers can now also define there libx dependencies - adding parser for bzip compressed content - adding parser for gzip compressed content - adding parser for zip files - adding parser for tar files - adding parser to detect the mime-type of a file this is needed by the bzip/gzip Parser.java - adding parser for rtf files - removing extra configuration file yacy.parser the list of enabled parsers is now stored in the main config file ) Adding configuration option in the performance dialog to configure See: http://www.yacy-forum.de/viewtopic.php?t=267 - maxActive / maxIdle / minIdle values for httpd-session-threadpool - maxActive / maxIdle / minIdle values for crawler-threadpool ) Changing Crawling Filter behaviour See: http://www.yacy-forum.de/viewtopic.php?p=2631 ) Replacing some hardcoded strings with the proper constants of the httpHeader class ) Adding new libs to libx directory. This libs are - needed by new content parsers - needed by new optional seed uploader - needed by SOAP API (which will be committed later) git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@126 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	b4030e5023	implemented serverSwitchActions - action-hooks git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@105 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
theli	6f4d2e5272	*) fixing replace bug. using stringvar = stringvar.replace(xxx) istead of stringvar.replace() git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@101 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
theli	2aa5fe8f50	*) Import statements reorganized Now it's easier to determine which class really uses which other class git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@82 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	2de90020ed	fixed caching+synchronization+brute-force-denial git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@67 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	7fb645b0ab	enhanced crawling performance, changed memory settings, new performace options git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@51 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	8b31f9e202	enhanced shut-down behaviour & added experimental nio-wrapper for kelondroRA (not active yet) git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@44 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	00f223cfc1	fixed post-parsing (a case when the bluelist is empty) git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@41 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
(no author)	1fec00bc24	*) Bugfix to avoid Nullpointer-Exceptions git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@30 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	b9203bdb50	bug fixes and code cleaning git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@22 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	c0807abd33	new crawl/proxy/cache design + fixes git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@18 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	e7d055b98e	very experimental integration of the new generic parser and optional disabling of bluelist filtering in proxy. Does not yet work properly. To disable the disable-feature, the presence of a non-empty bluelist is necessary git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@17 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago
orbiter	248077d3f0	initial load with yacy 0.36 git-svn-id: https://svn.berlios.de/svnroot/repos/yacy/trunk@1 6c8d7289-2bf4-0310-a012-ef5d649a1542	20 years ago

1 2 3 4

163 Commits (7429687601d13972bedc5ea5d015af36ac9085a0)