yacy_search_server

Commit Graph

Author	SHA1	Message	Date
Michael Peter Christen	4e3e2acc69	Merge branch 'master' of gitorious.org:yacy/rc1-fixed_percent-encoding	10 years ago
Michael Peter Christen	ecb6a59e9e	do not translate gif images into png images for thumbnails. Instead, stream the original to the search result thumb viewer. This has two reasons: - animated gifs cause 100% cpu and deadlocks in the jvm gif parser; a known bug which is obviously not yet fixed - animated gifs now appear in the search result also as animation	10 years ago
Michael Peter Christen	d9603039ff	automatically set the Q flag for smb/ftp start urls (split pdf support)	10 years ago
Michael Peter Christen	8600ea01dd	automatically swith on query option in case intranet protocols (smb/ftp) are used. This supports the new split-pdf option.	10 years ago
arucard21	3e9871291f	Applied URL-decoding prior to HTML-encoding. This removes percent-encoding from text shown in HTML	10 years ago
Ryszard Goń	3144313974	Postprocessing progress bar fix (Make it work as [probably] actually intended)	10 years ago
reger	6a04563578	Init Jetty using setDefaultDescriptor (web.xml) to defaults/web.xml so web.xml in defaults dir is applied first and optional DATA/SETTINGS/web.xml loaded on top. By using this Jetty feature (default web.xml) we assure that changes to the default are applied to existing installations and individual addition/changes are still respected.	10 years ago
reger	51ec9c1f44	fix "null" title in response writer for documents with multivalued title	10 years ago
reger	73ba5d8ef7	adjust fieldtype and description of field httpstatus_redirect_s in CollectionSchema - the field is not used (delete candidate)	10 years ago
reger	1f9389396a	fix NPE related 500 (Bad Request) response of UrlProxy on blacklisted urls, by adding parameter HTTPDeamon and removing unused hostAddress lookup code in sendRespondError	10 years ago
reger	7e4e9f7e32	improve yacysearchitem, prevent allocation of String (modifyURL) if feature not used	10 years ago
reger	61f75d6019	add xmpcore as direct dependency to pom (otherwise it's looked up at pdfbox archive path and not found there)	10 years ago
Michael Peter Christen	8ef56eda90	Merge branch 'master' of git@gitorious.org:yacy/rc1.git	10 years ago
Michael Peter Christen	9fce8bf2a5	crawling of multi-page pdfs with artificial post part on smb or ftp shares is not possible with the disabled setting; this is not temporary disabled until a better solution is on the hand.	10 years ago
reger	682dd94925	fix div by 0 in hello Caused by: java.lang.ArithmeticException: / by zero at hello.respond(hello.java:159)	10 years ago
reger	17808898c6	update to SLF4J 1.7.9	10 years ago
reger	f856edecb6	fix proxy redirect (http status 302) response fixes http://mantis.tokeek.de/view.php?id=517 The url given in bug report uses a gzip input stream which causes the HTTPClient.writeto() throw an IOException due to incomplete input stream. This in turn prevents the 302 reponse to the client browser. By limiting to serve target content just on httpstatus=200 will proxy the header reponse and client browsers redirect settings can be honored.	10 years ago
Michael Peter Christen	cc090bcb01	enhanced initialization of autotagging	10 years ago
Michael Peter Christen	003ec43bee	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	10 years ago
Michael Peter Christen	bef689d0a2	NPE fix	10 years ago
reger	1de33c6a53	add hint to Heuristics Config on "Greedy Learning Mode" in portal config, to point to a option to make this setting permanent.	10 years ago
reger	5332c9df21	update to commons-fileupload-1.3.1.jar (includes a security fix)	10 years ago
Michael Peter Christen	a0576ec737	fix for pdf sub-page result preparation	10 years ago
Michael Peter Christen	6ad43c4a8b	removed debug code	10 years ago
Michael Peter Christen	407cfff010	fix to wkhtmltopdf usage	10 years ago
Michael Peter Christen	5d321d3dc5	fixes to wkhtmltopdf call	10 years ago
Michael Peter Christen	eb78388a98	changed prefer strategy for http unique in such a way that http is preferred over https. While this is a bad idea from the standpoint of security it is more common applicable for environments where http and https mix and for some domains https is not available. Then the double-check is possible even if no postprocessing is performed.	10 years ago
Michael Peter Christen	84e2cccab4	fix to prevent assertion error in ranking servlet if no vocabularies are present that could be evaluated	10 years ago
Michael Peter Christen	9e588944fa	prevent NPE during initialization of very large vocabularies	10 years ago
Michael Peter Christen	aaf7d4775a	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	10 years ago
Michael Peter Christen	8c3e5b7b6d	added experimental pdf splitting which enables YaCy to split pdfs during parsing into individual pages and add them all using different URLs. These constructed urls are generated from the source url with an appended page=<pagenumber> attribute to the url get/post properties. This will distinguish the different page entries. The search result list will then replace the post parameter with a url anchor # mark which causes that the original url is presented in the search result. These URLs can be opened directly on the correct page using pdf.js which is now built-in into firefox. That means: if you find a search hit on page 5 and click on the search result, firefox will open the pdf viewer and shows page 5.	10 years ago
Michael Peter Christen	85773ebd4f	removed debug lines	10 years ago
Michael Peter Christen	d14114697c	the miss cache does not seem to work, it sometimes contains urlhashes from documents which actually are inside the index. This can be reproduced using the crawl result table at http://localhost:8090/CrawlResults.html?process=5 The cache is temporary disabled to remove the bad behaviour, however a later reactivation of that feater may be possible.	10 years ago
reger	deb75a1dbe	fix refactored size() -> filesize() in YMarkMetadata	10 years ago
reger	198102304b	refactor size() -> filesize() of URIMetadataNode (harmonize with ResultEntry and to not get confused with Collection.size())	10 years ago
reger	c6f634a4f2	remove redundant caching of urlhash in URIMetadataNode (is already cached in underlaying DigestURL .url) upd pom keyword for maven-antrun-plugin	10 years ago
Michael Peter Christen	445fafeb7c	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	10 years ago
Michael Peter Christen	0d69089c61	fix for division by zero	10 years ago
reger	ac61a39828	use peeraddress for link in remote crawl list to make link work without enabled proxy upd pom for Jetty (missing in last commit)	10 years ago
reger	fe5d4e6c7b	update to Jetty 9.2.6	10 years ago
Michael Peter Christen	5516819354	preventing the use of no-cache and expires in case that images are generated dynamically which will stay static in the future. This applies mainly to the search result favicon in front of search hits. These icons will now be generated once, but then caches in the browser. There is also a YaCy-internal cache for these icons which had prevented the re-generation of the icons in YaCy, but this cache is now superfluous since the browser should not call the servlet ViewImage again.	10 years ago
Michael Peter Christen	d3e71ed070	fixes for searches when initialization of large autotagging libraries have not been finished	10 years ago
Michael Peter Christen	28683530cd	fixes to usage of no-cache: use and recognize also the no-store directive	10 years ago
Michael Peter Christen	c9c700b510	reduction of http requests to YaCy using the correct cache-control, expires and last-modified headers in http response.	10 years ago
reger	eca578a5fa	update to PDFBox 1.8.8	10 years ago
reger	13cca2b114	fix missing AppPath upd Maven plugin versionid	10 years ago
Michael Peter Christen	d7e2f08a89	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	10 years ago
reger	0f7d4c42e9	include xmpcore.jar in classpath used by metadata-extractor	10 years ago
malykhin.dmitry	bd39e009ac	Update russian translation	10 years ago
Michael Peter Christen	65125439fe	added query modifier 'on'. This makes it possible to search for date occurrences within the (web) page documents (not the document last-modified!). This works only if the solr field dates_in_content_sxt is enabled. A search request may then have the form "term on:<date>", like gift on:24.12.2014 gift on:2014/12/24 * on:2014/12/31 For the date format you may use any kind of human-readable date representation(!yes!) - the on:<date> parser tries to identify language and also knows event names, like: bunny on:eastern .. as long as the date term has no spaces inside (use a dot). Further enhancement will be made to accept also strings encapsulated with quotes.	10 years ago

1 2 3 4 5 ...

11570 Commits (5a060c9f26c70ad79d4b0f6a8714cfc41bd32496) All Branches Search

11570 Commits (5a060c9f26c70ad79d4b0f6a8714cfc41bd32496)

All Branches