yacy_search_server

Commit Graph

Author	SHA1	Message	Date
Michael Peter Christen	c35170a305	more logging	11 years ago
Michael Peter Christen	e8be07ec78	grr	11 years ago
Michael Peter Christen	6f81bb756c	wrap wkhtmltopdf with xvfb if necessary	11 years ago
Michael Peter Christen	0119f8665d	more logging when failing to create pdf snapshot	11 years ago
Michael Peter Christen	416fe886e3	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	11 years ago
Michael Peter Christen	60f27bdf49	added the property timeoutrequests to configuration to disable TimeoutRequests. The purpose is to test if YaCy runs better on VMs where there is a limitation of concurrent processes; see /proc/user_beancounters in row numproc; this value is limited and should be low. Try to set timeoutrequests to keep this low. (works only after restart)	11 years ago
Michael Peter Christen	97f6089a41	YaCy can now create web page snapshots as pdf documents which can later be transcoded into jpg for image previews. To create such pdfs you must do: Add wkhtmltopdf and imagemagick to your OS, which you can do: On a Mac download wkhtmltox-0.12.1_osx-cocoa-x86-64.pkg from http://wkhtmltopdf.org/downloads.html and downloadh ttp://cactuslab.com/imagemagick/assets/ImageMagick-6.8.9-9.pkg.zip In Debian do "apt-get install wkhtmltopdf imagemagick" Then check in /Settings_p.html?page=ProxyAccess: "Transparent Proxy" and "Always Fresh" - this is used by wkhtmltopdf to fetch web pages using the YaCy proxy. Using "Always Fresh" it is possible to get all pages from the proxy cache. Finally, you will see a new option when starting an expert web crawl. You can set a maximum depth for crawling which should cause a pdf generation. The resulting pdfs are then available in DATA/HTCACHE/SNAPSHOTS/<host>.<port>/<depth>/<shard>/<urlhash>.<date>.pdf	11 years ago
Michael Peter Christen	41d00350e4	moved network configuration to Use Case submenu; this is necessary because the definiton of portal peers within the YaCy freeworld network is otherwise splitted into two different main menus.	11 years ago
reger	ff80700aff	replace depreciated Solr DateField.formatExternal with recommended TrieDateField.formatExternal	11 years ago
Michael Peter Christen	9ea120dbe5	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	11 years ago
reger	aa7122f079	update to guava.18.0.jar and jsch.0.1.51.jar	11 years ago
reger	0c97cc2440	skip unused call parameter for hashSentence()	11 years ago
reger	221f86dd5e	position api icon (ViewFile.html)	11 years ago
reger	4c14a8b44d	update to poi-3.10.1.jar	11 years ago
reger	ea633a794c	including small junit test case for WordTokenizer	11 years ago
reger	5790c7242e	skip to tokenize punktuation as word in WordTokenizer remove unused variables in condenser related to Tokenizer	11 years ago
reger	f07392ff17	add. use host port parameter in YaCyApp	11 years ago
Michael Peter Christen	09d2867050	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	11 years ago
Michael Peter Christen	ad0da5f246	added new web page snapshot infrastructure which will lead to the ability to have web page previews in the search results. (This is a stub, no function available with this yet...)	11 years ago
reger	aa0faeabc5	adjust translation text of error msg on empty query (ru: needs correction)	11 years ago
reger	c475be2937	fix (enable) error msg on empty query	11 years ago
reger	ef5c5b4489	update to Jetty 9.2.4	11 years ago
reger	f709132961	remove obsolete alternate link fix api link	11 years ago
Michael Peter Christen	5f5c7d69d1	added image screenshot generator	11 years ago
Michael Peter Christen	3c71e1c872	show vocabularies in search result (in case of debugging)	11 years ago
Michael Peter Christen	1d45d9405a	security bugfix	11 years ago
Michael Peter Christen	ff728b4aa5	ignore url errors during search	11 years ago
Michael Peter Christen	c94c24638f	disabled postprocessing by default. If you read this: please disable postprocessing in your peer as well: open /IndexSchema_p.html, then deselect field process_sxt	11 years ago
Michael Peter Christen	2fce2e2697	larger boost fields for ranking	11 years ago
Michael Peter Christen	6c03ff8355	bold words in snippets should not be coloured black in the base style because there are styles with dark backgrounds which make the bold word invisible	11 years ago
Michael Peter Christen	8317914ce3	changed vocabulary navigator object type to TreeMap to get a specific order into the vocabularies. This is now lexicographic which is not so much random as a hashed order	11 years ago
Michael Peter Christen	d5c1b07768	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	11 years ago
Michael Peter Christen	c0f9f6ac66	added option to change the navbar-default, i.e. usable for dark skins	11 years ago
Michael Peter Christen	10794e8efd	trying facet.method fc instead of fcs to handle large facets	11 years ago
Michael Peter Christen	041b605cfe	Merge branch 'master' of git@gitorious.org:yacy/rc1.git	11 years ago
Michael Peter Christen	f1f74e8626	toString fix	11 years ago
Michael Peter Christen	30276a2b48	prevent that a local Solr search and a local RWI search are running concurrently. When a RWI search result is flushed into the result set, id does Solr Queries (which replaced the old-style Metadata Queries) and they are possibly running concurrently to a previously startet Solr search. Both methods may block each other with IO. To enhance the speed, they are now serialized. Because the Solr search results may result in better results using the more advanced and configurable Ranking methods, this result is preverred over the RWI search result. However, remote RWI search results are still feeded concurrently into the search result as well.	11 years ago
Michael Peter Christen	84763126e0	added option to make the YaCy proxy act as the cache is never stale. If set to 'Always Fresh' the cache is always used if the entry in the cache exist. This is a good way to archive web content and access it without going online again in case the documents exist. To do so, open /Settings_p.html?page=ProxyAccess and check the "Always Fresh" checkbox. This is set do false which behave as set before. If you set this to true, then you have your web archive in DATA/HTCACHE. Copy this to carry around your private copy of the internet!	11 years ago
reger	1e7ee72240	fix path lookup to ./defaults/yacy.badwords (fix of commit `ee277b9b3e`)	11 years ago
reger	7d863d6254	fix empty text facet entry (noticed on Author facet)	11 years ago
Michael Peter Christen	a39419f2ef	more stacks shall be considered for on-demand loading, not only deep-depth stacks to prevent "too many open files" problem	11 years ago
Michael Peter Christen	5bb52f79be	reduce number of calls to queue.size() because that may be a bottleneck during crawling	11 years ago
Michael Peter Christen	4920ab7b76	optimize usage of size() cache	11 years ago
reger	ee277b9b3e	allow for local yacy.stopwords and yacy.badwords list (in DATA/SETTINGS/) if file in DATA/SETTINGS it is loaded otherwise file in ./defaults is loaded (if locale ./defaults/stopwords.xx doesn't exist take solr/lang/stopwords_xx.txt as default) move yacy.stopwords, yacy.stopwords.de and yacy.badwords.example out of root directory to ./defaults directory	11 years ago
reger	de56266bcb	remove redundant toLower for topwords	11 years ago
Michael Peter Christen	a34f837592	better delete all files in path when removing host crawl stack	11 years ago
Michael Peter Christen	10b1db430a	if we have many hosts, use on-demand earlier	11 years ago
Michael Peter Christen	1324927e66	prevent division by zero	11 years ago
Michael Peter Christen	2beb6abeb6	disabled crazy sleep loop	11 years ago
Michael Peter Christen	092d97d7ac	when importing vocabulary csv files, accept also files without semicolon and truncate quotes from literals	11 years ago

1 2 3 4 5 ...

11411 Commits (c35170a30549f13a11579a637568416fe0bda948) All Branches Search

11411 Commits (c35170a30549f13a11579a637568416fe0bda948)

All Branches