yacy_search_server

Commit Graph

Author	SHA1	Message	Date
sgaebel	df9ea0a42a	removes some warnings: unused imports, params	4 years ago
sgaebel	9bc2297161	fixes deleting during recrawl	4 years ago
sgaebel	80785b785e	adds deleting during recrawl	4 years ago
Michael Peter Christen	e0ad8ca9da	replaced json library from JSON.org with libandroid-json-java This fixes https://github.com/yacy/yacy_search_server/issues/347	5 years ago
Michael Peter Christen	ea8df27e95	modified org.json.* library to fit into the YaCy environment as drop-in replacement. Also made some fixes and enhancements to the library.	5 years ago
Michael Peter Christen	60dc1241a3	added org.json.* library from https://android.googlesource.com/platform/libcore/+/refs/heads/master/json/src/main/java/org/json as a preparation step for https://github.com/yacy/yacy_search_server/issues/347	5 years ago
Michael Peter Christen	053e54a2c7	grand CORS for json files	5 years ago
Michael Christen	cfa27d2fd5	fixed links	5 years ago
Michael Christen	cb20aa7e54	removed donation message in search result column	5 years ago
Michael Christen	25227676ae	removed some warnings	5 years ago
luccioman	6b45cd5799	New optional crawl filter on the URL a doc must match to crawl its links For finer control over which parsed documents can trigger an addition of their links to the crawl stack, complementary to the existing crawl depth parameter.	6 years ago
luccioman	d16bc99835	Added "Show Metadata" links to the ViewFile.html links mode To conveniently follow parsed links in the file viewer	6 years ago
luccioman	a5771b1f14	Made SNI extension user configurable without the need for server restart TLS Server Name Indication (SNI) extension activation can now be configured with the new Settings_p.html?page=httpClient administration page. SNI extension is also now enabled by default, as in 2019 the unrecognized_name(112) alert is more properly handled by major web servers TLS implementations, following the RFC 6066 standard. Related YaCy issues : #153 #189 and #272 JDK 1.7 bug : https://bugs.java.com/bugdatabase/view_bug.do?bug_id=7127374 Apache httpd issue : https://bz.apache.org/bugzilla/show_bug.cgi?id=56241 RFC 6066 : https://tools.ietf.org/html/rfc6066#section-3	6 years ago
luccioman	e90405b6f0	Support parsing audio URLs without file extension Added also a Junit for the audio tag parser	6 years ago
luccioman	a8316c79da	Allow JS resorting of search results by unauthenticated users Acces rate limitations to this search mode by unauthenticated users are set low by default to prevent unwanted server overload but can be customized through the SearchAccessRate_p.html configuration page Fixes #291	6 years ago
luccioman	0ab2b49c31	Made /yacysearch access rate limitations user configurable With a new admin page at /SearchAccessRate_p.html in menu Network Access > Local Search > Access Rate Limitations	6 years ago
luccioman	5b7e41202a	Added Solr GSA writer support for responses from remote instances	6 years ago
luccioman	4d8a948455	Properly close PDF snapshots loaded with pdfbox library	6 years ago
luccioman	74e6d6e984	Added Solr GrepHTML writer support for responses from remote instances	6 years ago
luccioman	5e6501974d	Added Solr snapshots writer support for responses from remote instances	6 years ago
luccioman	384c37102c	Improve accuracy of total results count on latest pages in Stealth mode Previously, when mixing results from local RWI and local Solr (Stealth mode), total local Solr count could be ignored on last result pages, when the page offset was higher than local Solr count but lower than total RWI count.	6 years ago
luccioman	5e9a08355a	Improved logging for federated search - Do not use spaces in logger identifier name so the log level can be configured in yacy.logging - Hold the logger instance to avoid the logging system to look for it from its name at each appended log message	6 years ago
luccioman	9782a98a9c	Added the possibility to customize facets sort type and direction Previously search navigators/facets elements were sorted only by counts. Now from the ConfigSearchPage_p.html admin page, sort direction (ascending/descending) and type (on counts or labels) can be customized independently for each navigator.	6 years ago
sgaebel	c2398fd890	remove warnings: 'Statement unnecessarily nested within else clause'	6 years ago
sgaebel	811d40a6c4	taking care of closing inputstreams, HTTPClient	6 years ago
sgaebel	8d2e7262d9	Recrawl: - set the chunksize to 100 to meet the max of the embedded solr - re-enable sorting (the case where we switched it of should be away) - enable recrawling on remote-solr	6 years ago
sgaebel	8f58c1dcfa	extend the SolrServlet to be usable as remote solr (incl. update) this feature needs to be enabled by uncomment the url-pattern	6 years ago
luccioman	7223a2fdb1	Removed usage of now deprecated Jetty function	6 years ago
luccioman	440d9f2fa0	Exclude peers with empty or disabled RWI from remote RWI search	6 years ago
luccioman	08ea0b0397	Added a configurable timeout to wkhtmltopdf calls for pdf snapshots Necessary to prevent blocking the indexing workflow when some wkhtmltopdf renderings fail without terminating	6 years ago
luccioman	3fb449b3b6	Properly resolve relative URLs against document URL in html base tags Fixes issue #256	6 years ago
luccioman	73a6e45524	Extended detection of external tools used for Snapshots generation This enable detecting wkhtmltopdf and Imagemagick convert executables when they are at system Path in addition to common installation paths.	6 years ago
luccioman	7dc1f60619	Fixed detection of absolute data folder path on MS Windows	6 years ago
luccioman	595e144797	Trace a message on incomplete proper server finish when killing process	6 years ago
luccioman	9daeea823b	Fixed concurrency issue on cache used for circles rendering Without synchronization lock, concurrent rendering of images including circles could lead to glitches as reported in issue #248	6 years ago
Michael Peter Christen	c347e7d3f8	Merge branch 'master' of https://github.com/yacy/yacy_search_server.git	6 years ago
Michael Peter Christen	848e9304d9	evil bots may crawl harder	6 years ago
luccioman	a997133260	Fixed gzip decompression regression on index transfer APIs Processing of gzip encoded incoming requests (on /yacy/transferRWI.html and /yacy/transferURL.html) was no more working since upgrade to Jetty 9.4.12 (see commit `51f4be1`). To prevent any conflicting behavior with Jetty internals, use now the GzipHandler provided by Jetty to decompress incoming gzip encoded requests rather than the previously used custom GZIPRequestWrapper. Fixes issue #249	6 years ago
luccioman	e85f231bdf	Fixed termination of Host browser and link structure Solr query threads On some conditions (especially when reaching timeout), concurrent Solr query tasks used by the /HostBrowser.html and /api/linkstructure.json never terminated, thus leaking resources, as reported by @Vort in issue #246	6 years ago
luccioman	fcf6b16db4	Added new crawler attribute for finer control over Media Type detection New "Media Type detection" section in the advanced crawl start page allow to choose between : - not loading URLs with unknown or unsupported file extension without checking the actual Media Type (relying Content-Type header for now). This was the old default behavior, faster, but not really accurate. - always cross check URL file extension against the actual Media Type. This lets properly parse URLs ending with an apparently odd file extension, but which have actually a supported Media Type such as text/html. Sample URLs with misleading file extensions added as documentation in the crawl start page. fixes issue #244	6 years ago
luccioman	a83a56473e	Added suport for PDF snapshots generation when running on MS Windows	6 years ago
luccioman	8852c97cee	Added basic styling for cleaner rendering of missing image snapshots For the output of the Solr snapshots writer	6 years ago
luccioman	746e0e788d	Render a relevant HTTP status code on snapshot image rendering error Instead of a null response body which is not very helpful.	6 years ago
luccioman	50b6edfcf5	Updated Solr snapshots writer for a cleaner html head	6 years ago
luccioman	f366f43d6b	Made snapshots size customizable in Solr snapshots response writer	6 years ago
luccioman	7a62fc0e66	Fixed concurrency issue in custom classloader used for template classes As reported in issue #241, the problem is only critical (random but complete crash of the JVM) when upgrading to JDK11.	6 years ago
luccioman	61c337f29a	Decode blacklist entries for easier edition of non ascii chars Not using the JDK URLDecoder.decode() function, as it strips '+' characters when they occur after '?' (both characters having regular expression semantics when used in blacklist path patterns)	6 years ago
luccioman	ed93221fa1	Improved normalization of blacklist path patterns having non ascii chars Normalize blacklist path patterns using percent-encoding, at pattern edition in web interface and at loading from configuration files. Fixes issue #237	6 years ago
luccioman	2a73b63d9e	Use a constant default target file name for seed SCP upload method To make seed upload (in /Settings_p.html?page=seed page) with SCP easier when the user specify a remote target directory path. See report by @vikulin in issue #227	6 years ago
luccioman	b5eabb626f	Removed some dead code	6 years ago

1 2 3 4 5 ...

8733 Commits (df9ea0a42a6188246bc77ea98b49f2f6babf50c6)