yacy_search_server

Commit Graph

Author	SHA1	Message	Date
luccioman	4564541b3b	Fixed blacklist Regex containing '+' characters rendering. As reported on YaCy forum by shni (http://forum.yacy-websuche.de/viewtopic.php?f=5&t=5970) when a blacklist entry contained both '?' and '+' characters, the '+' chars were wrongly decoded and rendered as spaces.	8 years ago
luccioman	0612a8f4f2	Fixed the previously added link to scheduled dump operations.	8 years ago
luccioman	a87281b498	Added MediaWiki dump import scheduling feature. Checking the last modified date by default to prevent unnecessary long running operations.	8 years ago
luccioman	10c03c6c64	Improved MediaWiki dump import monitoring. When import thread is terminated : - now stop refreshing and stay on the monitoring page to give user a feedback after a long running import - added link to the next monitoring step : results from surrogates reader - added link to new import On the new import page, added a link on the eventual last import report.	8 years ago
luccioman	edd7ccac40	Added some JavaDoc	8 years ago
luccioman	79fdf14b0a	Fixed regression introduced by commit `9ad4d16` On MediaWiki dump imports, the SurrogateReader was trying to unread too many bytes, then failing with the following exception : "java.io.IOException: Push back buffer is full".	8 years ago
Michael Peter Christen	7678fd67e3	copied fix from yacy_grid_parser for wrong array type	8 years ago
Michael Peter Christen	200b100fb8	added patch to rewrite altered yacy grid schema into yacy schema This generates the stub and protocol parts of an url for inboundlinks, outboundlinks and images	8 years ago
reger	9ad4d16829	Add a responsHeader to the solr index export with a format identifier and export parameter (in accordance with response xml format) for easier format detection on import.	8 years ago
luccioman	9697209ef6	Fixed Index Export feature for compatibility with old indexed documents. This is a fix for mantis 682 (http://mantis.tokeek.de/view.php?id=682) and issue #116	8 years ago
luccioman	88c062639b	Added some JavaDoc	8 years ago
luccioman	8d288f5dba	Crawl results page : apply table lines number limit. Take into account the already existing default limit value (especially useful after a long crawl or surrogates import), or a custom one from parameter "count". Added a "Show all" link for convenience.	8 years ago
luccioman	31fff2c986	Extended WikiCode template inclusion syntax support. Wiki templates are not rendered but syntax support is improved, which greatly enhance snippets rendering on search results coming from a MediaWiki dump import. Tested on various dumps from Wikimedia at https://dumps.wikimedia.org/backup-index.html See also Wikipedia transclusion documentation at https://en.wikipedia.org/wiki/Wikipedia:Transclusion	8 years ago
Michael Peter Christen	973d74712f	added yacy grid flatjson surrogate parser	8 years ago
luccioman	b1da92648e	Fixed surrogates import monitoring page (/CrawlResults.html?process=7) This page was always empty, as described in mantis 740 (http://mantis.tokeek.de/view.php?id=740)	8 years ago
luccioman	527d494c1a	Fixed "Unchecked conversion" compilation warnings.	8 years ago
reger	2b03e40134	upd to jwat-1.0.5	8 years ago
reger	7a7da698d4	fix unit test MultiProtocolURL(file) assertion for Windows path with drive letter.	8 years ago
reger	c77e43a391	Take out mailto collect in internal parsed document As earlier plans to make use of mailto as separate webgraph entity didn't materialize (see http://forum.yacy-websuche.de/viewtopic.php?f=8&t=5726&p=32493&hilit=mailto#p32493) free the unused handling and resources.	8 years ago
Michael Peter Christen	335868edba	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	8 years ago
reger	bec34d3546	Add url input field as source for WarcImporter allowing to import warc from url without prior download.	8 years ago
reger	d3df8a46c4	fix unresolved_pattern on missing post parameter api/message.html	8 years ago
luccioman	f66438442e	Extended Mediawiki dump import to remote URLs. When using a public HTTP URL in /IndexImportMediawiki_p.html, the remote file now is directly streamed and processed, allowing import of several GB dumps even with a low memory remote peer, and without need to manually download the dump file first.	8 years ago
luccioman	e5c3b16748	Improved http client close time on stream processing errors.	8 years ago
luccioman	23775e76e2	Fixed endless loop case in wikicode processing. Detected when importing recent MediaWiki dumps containing some pages with script content in plain text format (see Scribunto extension https://www.mediawiki.org/wiki/Extension:Scribunto ). Further improvement : modify the MediawikiImporter to prevent processing revisions whose <model> is not wikitext.	8 years ago
luccioman	0bc868a819	Improved support for non ASCII chars in local file system URLs Creating a MultiProtocolURL instance from a File object and then retrieving a File with getFSFile() was inconsistent with file paths containing space or non ASCII chars.	8 years ago
luccioman	7edddd7b0d	Improved error reports on various wiki dump prerequisites failure cases. Also added some JavaDoc.	8 years ago
luccioman	dfe8d4139b	Used a text input for wiki dump import file selection. Using an HTML "file" input was confusing (as reported by promocore on YaCy forum : http://forum.yacy-websuche.de/viewtopic.php?f=5&t=5965) , and it only worked with MS IE/Edge on a local YaCy peer : - for security reasons some current major browsers such as Firefox or Chrome do not allow to send full file path information when using a file form input - the local file system selection popup doesn't make sense when you want to import a dump on a remote YaCy server	8 years ago
reger	3a71430030	Adjust ConfigSearchPage_p to activated hosts navigator as plugin	8 years ago
reger	7b80189bda	Activate hosts navigator plugin. This includes rwi results in the navigator count. This might be tangential related to http://mantis.tokeek.de/view.php?id=736 as the example includes a local index search, while rwi results are not counted.	8 years ago
reger	05a1b14b4a	add missing text from ConfigRobotsTxt_p to master.lng and link to Translation Editor to Translation News page.	8 years ago
reger	a39c00a93f	add servlet to list user in UserDB and made user editor available in separate servlet for a quick and easy overview of configured user and selection for edit.	8 years ago
reger	a4498e17c0	fix edit current user form to required post mehtod introduced with `cde237b687`	8 years ago
Michael Peter Christen	f5ad29edb1	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	8 years ago
Michael Peter Christen	76e9135526	added flatjson parser (stub, unfinished)	8 years ago
reger	46a4aaf09c	upd to Solr-5.5.4	8 years ago
reger	b7417ac329	Introduce a Keyword search navigator using the index field keywords. The keywords field string is split into words as navigator entries. A keyword navigator facet is essential for search appliance usage were documents and metadata use often specialized keyword vocabularies to filter search results. This navi can be used without custom index schema. As we don't have defined a search query command to filter "keywords" yet, the filtering is limited by adding the keyword to the search query.	8 years ago
reger	eddb7a9804	upd to pdfbox-2.0.5.jar and transient dependency xmpcore-5.1.3.jar required by metadata-extractor-2.10.1 (fix build.xml compiler warning)	8 years ago
reger	27884da1ff	add CookieTest_p.html text to master.lng	8 years ago
luccioman	665d087d76	Enforced access controls on a few more administration pages. - ensure use of HTTP POST method when performing server side effect operations - transaction token required to ensure the request has effectively been requested by user interaction	8 years ago
luccioman	0feded21dd	Escaped HTML eventually active content from recorded API call comments.	8 years ago
luccioman	09e72eb0a4	Set Config Portal as a private administration page. Consistently with its required action from submission credentials, and because external unauthenticated users do not need to access these settings.	8 years ago
reger	c19d60f06b	update master.lng with recent text changes to IndexExport_p.html, IndexImportWarc_p.html	8 years ago
reger	9339a6a4c5	use css error class for error msg in IndexImportOAIPMH_p.html, adjust to xhtml <p> usage rule	8 years ago
reger	777cb5b812	remove test case for Standard_MemoryControl which will always fail see https://github.com/yacy/yacy_search_server/pull/114	8 years ago
reger	ba339a2a45	Add servlet to import warc file from filesystem IndexImportWarc_p.html. Apply Importer interface to WarcImporter	8 years ago
Michael Peter Christen	1d81b8f102	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	8 years ago
Michael Peter Christen	69081bce00	added export to elasticsearch. The export dump can easily be imported to elasticsearch using the command curl -XPOST localhost:9200/collection1/yacy/_bulk --data-binary @yacy_dump_XXX.flatjson	8 years ago
reger	510f11d374	Implement surrogate import from Warc archives (as first option handle warc = Web ARChive File Format. Warc files with extension .warc or compressed warc.gz can be placed in the DATA/surrogate/in and contained responses are imported to the index. The used library is stream based so we can easily extend it later to use and load warc's from the net.	8 years ago
luccioman	5b5b9d5d96	URL Viewer : only display the link to metadata when metadata exists	8 years ago

... 7 8 9 10 11 ...

13569 Commits (bcbd0ae1a4ea4550089b82d7ae47b93a8524e787) All Branches Search

13569 Commits (bcbd0ae1a4ea4550089b82d7ae47b93a8524e787)

All Branches