yacy_search_server

Commit Graph

Author	SHA1	Message	Date
Michael Peter Christen	500cfa9457	enhanced logging	10 years ago
Michael Peter Christen	c14bc8d9b7	revert of fq transformation (recent fix)	10 years ago
Michael Peter Christen	203df5a750	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	10 years ago
reger	fa08ca207e	! finish running crawls before applying ! Allow crawl urls up to 2048 character fix for http://mantis.tokeek.de/view.php?id=575	10 years ago
reger	ee77f24e52	use some more declared HeaderFramework constants	10 years ago
Michael Peter Christen	11a848da5a	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	10 years ago
Michael Peter Christen	b94bd7f20a	a collection of search query enhancements: - fixed superfluous space in query field list - fixed filter query logic - removed look-ahead query which caused that each new search page submitted two solr queries - fixed random solr result orders in case that the solr score was equal: this was then re-ordered by YaCy using the document hash which came from the solr object and that appeared to be random. Now the hash of the url is used and the score is additionally modified by the url length to prevent that this particular case appears at all.	10 years ago
reger	dbe2594c38	replace deprecated myPublicLocalIP() in AbstractRemoteHandler	10 years ago
reger	6d3534e725	remove unused Transmission hit counter	10 years ago
reger	cb67eb7baf	use more absolute path for config file opening as suggested in pull request 5 (https://github.com/yacy/yacy_search_server/pull/5)	10 years ago
Michael Peter Christen	1ccbf739b1	added bayes filter from Philipp Nolte, originally taken from https://github.com/ptnplanet/Java-Naive-Bayes-Classifier and modified inside the loklak.org project. After optimization in loklak it was inserted into the net.yacy.cora.bayes package. It shall be used to create custom search navigation filters. The original copyright notice was copied from the README.md from https://github.com/ptnplanet/Java-Naive-Bayes-Classifier/blob/master/README.md The original package domain was de.daslaboratorium.machinelearning.classifier	10 years ago
Michael Peter Christen	1bced1ae60	using latest enhanced (un/)gzip methods from loklak for yacy	10 years ago
Michael Peter Christen	3e6657288d	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	10 years ago
Michael Peter Christen	de8cfbe1d7	added export option to export the fulltext of the search index text only	10 years ago
reger	2fb6ebe88a	move java environment parameter setting disabling SNI (Server Name Indicator) support for https connections from code to startup script allowing admin to ~easy/transparent alter the YaCy default FALSE setting. Background: some user report problem with connecting/crawling some sites via https which require SNI support (by default switched off in YaCy). On the other hand systems not demanding SNI support are sometimes not properly configured and due to a bug/feature in java 1.7 connection is aborted. The later is more often the case, so the default is still fine. With the java start parameter expert user can no alter the startparameter to -Djsse.enableSNIExtension=true (java default) if they crawl more hosts requiring SNI support. The alternative to let YaCy try both during https handshake (deep inside the httpclient) is not pursut at this time.	10 years ago
Michael Peter Christen	fbeae20b3a	try a healing of the cache if the index file is corrupted	10 years ago
Michael Peter Christen	03ea723889	added log lines for query performance profiling	10 years ago
Michael Peter Christen	0e87a99ab8	more fixes for special windows paths	10 years ago
Michael Peter Christen	e5b6424eed	patch for bad windows file paths	10 years ago
Michael Peter Christen	0aa6fcf259	remove old vocabularies and synonyms before adding new	10 years ago
Michael Peter Christen	289018b559	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	10 years ago
Michael Peter Christen	7b412e8c07	added msg (text emails) format; should be handled by html parser.	10 years ago
reger	f91298d3b6	fix one implicit Integer/Long type conversion -> causes Java 1.8 compile error	10 years ago
reger	821262a179	add CommonPattern for multiple spaces to eliminate empty split words on following spaces	10 years ago
Michael Peter Christen	90f75c8c3d	added enrichment of synonyms and vocabularies for imported documents during surrogate reading: those attributes from the dump are removed during the import process and replaced by new detected attributes according to the setting of the YaCy peer. This may cause that all such attributes are removed if the importing peer has no synonyms and/or no vocabularies defined.	10 years ago
Michael Peter Christen	7829480b82	refactoring: separated condenser and tokenizer	10 years ago
Michael Peter Christen	593de05922	enhanced surrogate import process speed (dramatically!)	10 years ago
Michael Peter Christen	3c4c69adea	fix for - bad regex computation for crawl start from file (limitation on domain did not work) - servlet error when starting crawl from a large list of urls	10 years ago
Michael Peter Christen	1fec7fb3c1	suppress access to solr when doing search suggestions in case that the index has more than two million documents. This protects the index from beeing flooded with search requests that cannot be resolved before the real search query has to be computet.	10 years ago
Michael Peter Christen	694b22f165	migration to Solr 5.2: huge benefits - this is a lot faster! This is a very complex migration: many classes had been renamed or removed, dependencies changed and the solr index type is now aligned to be a solr cloud repository. Together with the Solr 5.2 library update, one other dependent library had been updated as well: httpclient 4.4->4.4.1 Older indexes are migrated from 4_10 to 5_2. However, the new index structure is more efficient and we recommend to re-index everything. Please use the index export before you do the update to a large surrogate xml file. After the update, start with an empty index and then initialize this with your dump.	10 years ago
sixcooler	e427efbe54	Next Try for a fix for upload-connection staying in blocked state. This was caused by reading via GZIP from close-wait connection an caused high cpu- and system-loads. Instat of implementing handling of the RedListener now I found a timelimeted 'get' "realy" solving this problem.	10 years ago
reger	0fab445b19	Resourceobserver log warning - deleting releases files - only on actual deletes instead of entering routine	10 years ago
sixcooler	ef6a64b2a4	Fix for upload-connection staying in blocked state. This was caused by reading via GZIP from close-wait connection an caused high cpu- and system-loads. Solved by implementing handling of the RedListener.	10 years ago
reger	c973f94936	add log entry on release file delete by ResourceObserver	10 years ago
reger	121972752c	implement deleteOldDownloads in RexourceObserver on low diskspace - direct assign sb.observer (skip redundant InitThread)	10 years ago
Michael Peter Christen	9c12555be5	added link to Snapshots in search results if the snapshot exists and option is set in ConfigSearchPage_p (this is a stub: we also need a visualization of pdf files!)	10 years ago
reger	72f6a0b0b2	enhance recrawl job - allow to modify the query to select documents to process (after job has started) - allow to include failed urls (httpstatus <> 200)	10 years ago
reger	7478338a40	remove augmented parsing activation from frontend experimental implementation not used and based on error prone experimental rdfaparser	10 years ago
reger	11aa2edfe1	remove RDFa parser activation from frontend reason: experimental implementatin of RDFa parser not executed (limited to special urls) but may cause error on normal html parsing due to a inputstream.reset	10 years ago
reger	49b79987c9	remove obsolete searchfl work table was used to register urls with not complete words in snippet but is never accessed	10 years ago
Michael Peter Christen	d0aff91f23	fix for index import	10 years ago
Michael Peter Christen	34de1e8cbc	gzip compression will perform more efficient and with better compression level	10 years ago
Michael Peter Christen	98be59ce9c	full solr xml exports will now be automatically compressed during export. That makes it possible to export a solr xml dump even if disc space is low.	10 years ago
Michael Peter Christen	a1a8edfc0a	wrap HeaReader close() in a catch Throwable block to prevent that an excpetion during close blocks the whole shotdown process	10 years ago
Michael Peter Christen	b43811d38c	added surrogate import process for exported solr dumps. Just throw your solr dump file into DATA/SURROGATES/in/ and it will be imported!	10 years ago
Michael Peter Christen	b77537294d	prevent disc usage when showing tray animation	10 years ago
Michael Peter Christen	eec78e1b0c	added intensity option to graphics	10 years ago
Michael Peter Christen	a5007f345e	re-licensing some of my old visualization classes under LGPL 2.1	10 years ago
Michael Peter Christen	c99a665593	adding a 3-pixel font generator made some time ago..	10 years ago
Michael Peter Christen	c7576d6028	added a full solr export to the IndexControlURLs_p.html servlet. The export function is also now the default export option. The export file format for a full solr export is very similar to a solr search result xml, only the <lst name="responseHeader"> tag is missing. The exported xml has a special line termination feature: all documents will be exported into a single line without any CR in between. That means that every document is completely inside a single line. While this is not readable at all for humans, it is very useful for linux line processing scripts, like grep. Using grep it will be easy to select single documents which match for a given pattern. Such dumps shall be importable with the DATA/SURROGATE/in import function, but that import is not yet adopted to the new file format.	10 years ago

1 2 3 4 5 ...

7762 Commits (500cfa9457e15cb12bd8d4bac7c579883f876527)