yacy_search_server

Commit Graph

Author	SHA1	Message	Date
reger	10b0eb106f	fix link target on iframe list in CrawlProfileEditor	9 years ago
reger	5744342fec	handle image preview for url w empty file extension fix of commit `688f7b2a5c`	9 years ago
reger	43c27aa550	upd to solr/lucene 5.3.1	9 years ago
reger	688f7b2a5c	allow/display svg images in image results previews svg is not supported by awt but by most browser. Image content is delivered as received (without size adjustment)	9 years ago
Michael Peter Christen	225200194a	every time a crawl is started, the user expects a different search result behaviour. This requires that the search cache is flushed for each crawl start. TODO: this should also be done if a crawl is terminated.	9 years ago
reger	b92d81b073	remove double caching of inputstream in ViewImage	9 years ago
Michael Peter Christen	3c31bf845f	fix for latest merge	9 years ago
luc	5578886f6f	Merge branch 'master' of https://github.com/luccioman/yacy_search_server.git	9 years ago
reger	2951c9fc40	remove unused check for known fileextension in searchtrailer (check is done on add to filetype-nav)	9 years ago
reger	733d725dec	limit css scrolling to result/content window x from pull request #10	9 years ago
Burkhard	4c38083a11	Merge pull request #10 from Raegdan/raegdan-css-layout-fix Fixed CSS scrolling	9 years ago
luccioman	a7179138ce	Returned again to main repository location : does anyone want to consider mantis 597 ? (http://mantis.tokeek.de/view.php?id=597)	9 years ago
luccioman	199b2ce52d	Translator refactoring : to simplify locale files writing, process keys as simple string and no more as regular expressions. Updated all locale files to adapt to refectored Translator : removed useless escaped characters and did minor corrections. Performed minor syntax corrections on some html source files. Added an util to translate all html source files with all locales without launching full YaCy application. Corrected main arguments parsing on other translation utils.	9 years ago
luccioman	4dd9c0d5d9	Merge from main repository	9 years ago
Michael Peter Christen	0a37d8af89	in case that a site crawl is started for urls with file:// path, the host filter does not work because there is no host given in such urls. In that case, patch the filter to be a sub-path filter.	9 years ago
luccioman	9df249296a	Return to mai repository version	9 years ago
luccioman	c1d937a90c	Merge branch 'master' of ssh://git@github.com/yacy/yacy_search_server	9 years ago
reger	7c1da173e0	fix missing license in image search see http://mantis.tokeek.de/view.php?id=522	9 years ago
luccioman	918ef72bbe	Corrected br markup	9 years ago
luccioman	f88bb2277e	Corrected bookmark link title	9 years ago
luccioman	802ea66d19	Merge branch 'master' of ssh://git@github.com/yacy/yacy_search_server	9 years ago
reger	5297e80cda	fix missing onclick in ConfigPortal to enable checkbox	9 years ago
luccioman	70e483ecc6	Merge branch 'master' of ssh://git@github.com/yacy/yacy_search_server	9 years ago
sixcooler	87e4abe393	fight the fieldcache by usind DocValues: in Solr-5.x the fieldcache has moved and was not cleared anymore. This results in an huge fieldcache. (http://lucene.apache.org/#highlights-of-the-lucene-release-include https://issues.apache.org/jira/browse/LUCENE-5666) Here I try to use DovValues where it is possible. For this I used the Api-Scheme as new basis für the Solr-Schema. This needs at least a complete optimization of the Solr-Index to get a smaller FieldCache. Everything that is indexed with these setting will not use the Fieldcache at all.	9 years ago
luccioman	67799ce867	Updated translation of index.html, yacysearch.html and simpleheader.template, corrected some special characters not written as HTML entities.	9 years ago
Michael Peter Christen	df3314ac1a	added a new facet type based on a probabilistic classifier using bayesian filters. This can be used to classify documents during indexing-time using a pre-definied bayesian filter. New wordings: - a context is a class where different categories are possible. The context name is equal to a facet name. - a category is a facet type within a facet navigation. Each context must have several categories, at least one custom name (things you want to discover) and one with the exact name "negative". To use this, you must do: - for each context, you must create a directory within DATA/CLASSIFICATION with the name of the context (the facet name) - within each context directory, you must create text files with one document each per line for every categroy. One of these categories MUST have the name 'negative.txt'. Then, each new document is classified to match within one of the given categories for each context.	9 years ago
Michael Peter Christen	dbbad23e12	removed warnings	9 years ago
reger	9e4043731d	add missing ; in base.css	9 years ago
Michael Peter Christen	de8cfbe1d7	added export option to export the fulltext of the search index text only	9 years ago
Kirill Fomchenko	ab22a32c09	Fixed CSS scrolling When the sidebar on search page becomes scrollable, the scrollbar shrinks the sidebar and makes the search results weirdly scrollable on X axis by several pixels. Now the sidebar always have a scrollbar, and results are never X-scrollable.	9 years ago
Michael Peter Christen	785781253e	added jsonp to suggest servlet	9 years ago
reger	821262a179	add CommonPattern for multiple spaces to eliminate empty split words on following spaces	10 years ago
Michael Peter Christen	f901e7d3cf	fix for non-authorized view of IndexBrowser: show only the number of non-failure documents	10 years ago
Michael Peter Christen	3c4c69adea	fix for - bad regex computation for crawl start from file (limitation on domain did not work) - servlet error when starting crawl from a large list of urls	10 years ago
Michael Peter Christen	1fec7fb3c1	suppress access to solr when doing search suggestions in case that the index has more than two million documents. This protects the index from beeing flooded with search requests that cannot be resolved before the real search query has to be computet.	10 years ago
Michael Peter Christen	886fca2260	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	10 years ago
Michael Peter Christen	694b22f165	migration to Solr 5.2: huge benefits - this is a lot faster! This is a very complex migration: many classes had been renamed or removed, dependencies changed and the solr index type is now aligned to be a solr cloud repository. Together with the Solr 5.2 library update, one other dependent library had been updated as well: httpclient 4.4->4.4.1 Older indexes are migrated from 4_10 to 5_2. However, the new index structure is more efficient and we recommend to re-index everything. Please use the index export before you do the update to a large surrogate xml file. After the update, start with an empty index and then initialize this with your dump.	10 years ago
Michael Peter Christen	6c2e6f1f37	remove redundant code	10 years ago
Michael Peter Christen	9c12555be5	added link to Snapshots in search results if the snapshot exists and option is set in ConfigSearchPage_p (this is a stub: we also need a visualization of pdf files!)	10 years ago
reger	72f6a0b0b2	enhance recrawl job - allow to modify the query to select documents to process (after job has started) - allow to include failed urls (httpstatus <> 200)	10 years ago
Michael Peter Christen	e0a23c56c7	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	10 years ago
Michael Peter Christen	fb9e1dd3f5	servlet for latest commit	10 years ago
reger	7478338a40	remove augmented parsing activation from frontend experimental implementation not used and based on error prone experimental rdfaparser	10 years ago
reger	11aa2edfe1	remove RDFa parser activation from frontend reason: experimental implementatin of RDFa parser not executed (limited to special urls) but may cause error on normal html parsing due to a inputstream.reset	10 years ago
Michael Peter Christen	ff11ac89f7	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	10 years ago
Michael Peter Christen	5e2d23b7a0	removed the new index export method from the IndexControlURLs_p.html servlet and moved it to a new /IndexExport_p.html servlet. This servlet is now more prominent linked in the main menu under Production -> Index Export/Import	10 years ago
reger	49b79987c9	remove obsolete searchfl work table was used to register urls with not complete words in snippet but is never accessed	10 years ago
Michael Peter Christen	b43811d38c	added surrogate import process for exported solr dumps. Just throw your solr dump file into DATA/SURROGATES/in/ and it will be imported!	10 years ago
Michael Peter Christen	eec78e1b0c	added intensity option to graphics	10 years ago
Michael Peter Christen	c7576d6028	added a full solr export to the IndexControlURLs_p.html servlet. The export function is also now the default export option. The export file format for a full solr export is very similar to a solr search result xml, only the <lst name="responseHeader"> tag is missing. The exported xml has a special line termination feature: all documents will be exported into a single line without any CR in between. That means that every document is completely inside a single line. While this is not readable at all for humans, it is very useful for linux line processing scripts, like grep. Using grep it will be easy to select single documents which match for a given pattern. Such dumps shall be importable with the DATA/SURROGATE/in import function, but that import is not yet adopted to the new file format.	10 years ago
Michael Peter Christen	47682bf467	fix for unresolved pattern	10 years ago
Michael Peter Christen	197f7449e5	All entities of crawl profiles are now editable in the crawl profile editor.	10 years ago
reger	1d8e1e4bac	- Image search expand box, adjust javascript hs padtominsize parameter, to make sure expand box doesn't shrink on small images - asure ImageResult.imagetext has value for the link text (use filename if no alt text given)	10 years ago
reger	000dde9511	Eleminate duplication of values for search ResultEntry by instatiation from URIMetadataNode, by eleminating differentiation of ResultEntry/URIMetadataNode. - moved remaining ResultEntry functionallity to URIMetadataNode - for 1:1 functionallity added a function makeResultEntry() - removed ResultEntry - refactored related code Main difference is after makeResultEntry the text_t content is removed and alternative title/url strings for display are calculated. Main difference left is, that	10 years ago
reger	3d53da8236	refactor ResultEntry to be based on MetadataNode/SolrDocument to share/reuse common access routines	10 years ago
reger	17e820cfd7	use doctype() in ViewFile to choose display routines in preference of getfileExtension()	10 years ago
reger	aa83931765	Convert content charset for display via CacheResource_p Cached resource charset encoding might not fit to internal handling (using utf-8), convert resource to utf-8 see http://mantis.tokeek.de/view.php?id=576	10 years ago
reger	3e742d1e34	Init remote crawler on demand If remote crawl option is not activated, skip init of remoteCrawlJob to save the resources of queue and ideling thread. Deploy of the remoteCrawlJob deferred on activation of the option.	10 years ago
Michael Peter Christen	dbf9e3503d	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	10 years ago
Michael Peter Christen	8b1a30be50	removed a -UNRESOLVED_PATTERN-	10 years ago
Michael Peter Christen	9938c81378	fix for division by zero	10 years ago
reger	ace71a8877	Initial (experimental) implementation of index update/re-crawl job added to IndexReIndexMonitor_p.html Selects existing documents from index and feeds it to the crawler. currently only the field fresh_date_dt is used determine documents for recrawl (fresh_date_dt:[* TO NOW-1DAY] Documents are added in small chunks (200) to the crawler, only if no other crawl is running.	10 years ago
Michael Peter Christen	f810915717	added crawl start from a clone with very, very large url: they are now encoded as post submit form inside a javascript creation function.	10 years ago
reger	609c52e987	refactor getBookmark to consistenly check existance by != null (w/o throwing exception on not found)	10 years ago
reger	5f4d35437e	add bookmark.query to edit form	10 years ago
reger	89124335c4	update bookmark autosearch description - add german translation	10 years ago
Michael Peter Christen	213401a446	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	10 years ago
Michael Peter Christen	97930a6aad	added must-not-match filter to snapshot generation. also: fixed some bugs	10 years ago
reger	b47267b79c	precaution against NPE on createorgetBookmark on search result	10 years ago
Michael Peter Christen	75879e051b	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	10 years ago
reger	8a5b8f8789	on bookmaring of search result, remember orig. query in separate bookmark property (instead of using the description field) - adjust display and autosearch - don't overwrite existing bookmark but combine info	10 years ago
reger	cf1fc7f700	harmonize filesearch input box layout	10 years ago
Michael Peter Christen	e334a06370	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	10 years ago
reger	579303a04e	add additional links to crawl queue pages	10 years ago
Michael Peter Christen	99718dc09a	don't record dump generation calls since that - is not a change of the index - happens very often within self-backup strategies from the outside (i.e. cronjobs)	10 years ago
Michael Peter Christen	5b59477415	update to bootstrap.css 3.3.4	10 years ago
Michael Peter Christen	0d365e67a5	Merge pull request #2 from Scarfmonster/master English Synonyms and small fixes	10 years ago
Eugene Kuligin	8ae3229306	add vertical margin to the search cloud block	10 years ago
Eugene Kuligin	f9408dfa48	fix RSS icon displaying	10 years ago
reger	f7b0148f6a	fix NPE in Vocabulary_p servlet called w/o parameter	10 years ago
Ryszard Goń	ca1a70aec8	fix for Accept '?' URLs column in Crawl Profile List	10 years ago
Ryszard Goń	b0cd0212fd	SynonymLibrary status check fix for multiple files	10 years ago
Ryszard Goń	f3f1b2e899	added English synonyms	10 years ago
reger	296e97c78e	put https port in peers dna as we flag if a peer is accesible via https, we need to know the port if we want to use is (e.g. for interYaCy communication) start to provide / tansport the port by recording it in peers dna. - add https link on the Network.html lock symbol	10 years ago
Michael Peter Christen	088853c1e8	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	10 years ago
Michael Peter Christen	fed26f33a8	enhanced timezone managament for indexed data: to support the new time parser and search functions in YaCy a high precision detection of date and time on the day is necessary. That requires that the time zone of the document content and the time zone of the user, doing a search, is detected. The time zone of the search request is done automatically using the browsers time zone offset which is delivered to the search request automatically and invisible to the user. The time zone for the content of web pages cannot be detected automatically and must be an attribute of crawl starts. The advanced crawl start now provides an input field to set the time zone in minutes as an offset number. All parsers must get a time zone offset passed, so this required the change of the parser java api. A lot of other changes had been made which corrects the wrong handling of dates in YaCy which was to add a correction based on the time zone of the server. Now no correction is added and all dates in YaCy are UTC/GMT time zone, a normalized time zone for all peers.	10 years ago
reger	f6a55f9279	incoming connection count/text fix improvement on http://mantis.tokeek.de/view.php?id=570	10 years ago
Michael Peter Christen	3e338e0987	Merge pull request #1 from Scarfmonster/master Search navigation fix	10 years ago
reger	702c30e619	add info text icon next to Augmented Browsing check-box with hint to config page	10 years ago
reger	4c907bec89	show "Augmented Browsing" link in search result only if urlproxy allowed and option switched on in layout (AugmentedBrowsing_p.html, ConfigSearchPage_p.html) as user only gets a error page if the option is not enabled	10 years ago
Ryszard Goń	6d78a6d06e	Search navigation fix	10 years ago
Michael Peter Christen	a08a3c5f29	reverted json syntax for facet results to version from january	10 years ago
Michael Peter Christen	d8cc773d05	fix for not valid json in case that topics are switched off	10 years ago
Michael Peter Christen	1df6492019	enhanced suggestions	10 years ago
Michael Peter Christen	c7fdde3bd1	replaced "fork me" banner with github banner	10 years ago
Michael Peter Christen	876cdb083f	Merge branch 'master' of github.com:yacy/yacy_search_server	10 years ago
Michael Peter Christen	2e88028c1a	when selecting collections in navigation, do show the un-selected collections in search result. When selecting one of them in another search, switch off the previously selected collection. This actually turns the collection navigation modifier into a radio-button like behaviour	10 years ago
reger	2f592a8063	add SynonymLibrary status to DictionaryLoader_p servlet http://mantis.tokeek.de/view.php?id=564	10 years ago
reger	c59ebde083	show location nav as selectable nav in search page layout - switch automatically on upon load of geodata provider - but allow switch on also without geodata file (and display the location nav if search result has lat/lon location)	10 years ago
Michael Peter Christen	5bc1e5cfbf	use a cursor hand on facet headline to show that this is clickable	10 years ago
Michael Peter Christen	40389987ec	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	10 years ago
Michael Peter Christen	f9ba50379d	added an expansion option to search facets on result page: - if less or equal of 8 facet options are present, they are shown by default - if more facet options are present, they are hidden To view or hide all facets, just click on the facet header bar	10 years ago
reger	b1ec0644e5	fix NPE in location search on missing/empty PubDate in underlaying rss data	10 years ago
reger	2f84b04fa9	add err msg on failure during Load_rss	10 years ago
reger	96292cf3eb	shorten exception loggin on not available connection in Load_RSS_p servlet	10 years ago
reger	66d0b5046a	fix NPE on viewfile of url not in index	10 years ago
Michael Peter Christen	5789c96292	fix: banner did not show link and qph for portal mode	10 years ago
Michael Peter Christen	9bf0d7ecb9	added a new collection type 'dht' to all documents from the peer-to-peer interface to distinguish rich and poor document data. This also reverts some changes from commit `796770e070` because the firstSeen database is the wrong method to distinguish these types of data	10 years ago
reger	7fcf0d0b71	fix missing display of CrawlerMonitor -> robots.txt Monitor revert delete of file api/table_p.html see `3ffe19b85c` (still used in this menu)	10 years ago
Marc Nause	efadb710a4	Updated Git links from Gitorious to Github.	10 years ago
reger	65f8371163	fix link to DeReWo project page	10 years ago
reger	a5d19e2982	update configheuristics_p.html text to state current opensearch heuristic function	10 years ago
reger	74ed399180	remove unused statement	10 years ago
Michael Peter Christen	fd4e2c809a	Show dates in the content of a document in the search result: - if an eventDate is given in the search result, replace the document date with the event date and prefix it with the string "on ". - the document date is omitted if a date from the cent is shown Added also the date as fields in the json and rss result sets.	10 years ago
Michael Peter Christen	710a0efa1b	generalized time period computations	10 years ago
Michael Peter Christen	dcfc384eee	bugfix for fixed host/port	10 years ago
Michael Peter Christen	535f1ebe3b	added a new way of content browsing in search results: - date navigation The date is taken from the CONTENT of the documents / web pages, NOT from a date submitted in the context of metadata (i.e. http header or html head form). This makes it possible to search for documents in the future, i.e. when documents contain event descriptions for future events. The date is written to an index field which is now enabled by default. All documents are scanned for contained date mentions. To visualize the dates for a specific search results, a histogram showing the number of documents for each day is displayed. To render these histograms the morris.js library is used. Morris.js requires also raphael.js which is now also integrated in YaCy. The histogram is now also displayed in the index browser by default. To select a specific range from a search result, the following modifiers had been introduced: from:<date> to:<date> These modifiers can be used separately (i.e. only 'from' or only 'to') to describe an open interval or combined to have a closed interval. Both dates are inclusive. To select a specific single date only, use the 'to:' - modifier. The histogram shows blue and green lines; the green lines denot weekend days (saturday and sunday). Clicking on bars in the histogram has the following reaction: 1st click: add a from:<date> modifier for the date of the bar 2nd click: add a to:<date> modifier for the date of the bar 3rd click: remove from and date modifier and set a on:<date> for the bar When the on:<date> modifier is used, the histogram shows an unlimited time period. This makes it possible to click again (4th click) which is then interpreted as a 1st click again (sets a from modifier). The display feature is NOT switched on by default; to switch it on use the /ConfigSearchPage_p.html servlet.	10 years ago
reger	ba276d3e64	add description_txt to default query fields, Dublin Core Metadata field extracted by most parsers.	10 years ago
reger	ad1596f9ac	upd lucene api doc link	10 years ago
reger	1196ff01c8	revert: formatting fix eats also up highlighting need other solution for snippets with unwanted html code	10 years ago
reger	61f42a7928	fix formatting issue in search result display if description contains html code noticed e.g. for id=NmNdJ9uApLaQ http://hswong3i.net/blog/hswong3i/virtualmin-drupal-7-x-ubuntu-12-04-howto	10 years ago
Michael Peter Christen	6578ff3ddb	enhanced suggest function	10 years ago
reger	ab98f69592	fix: searchoption hint for heuristic	10 years ago
Michael Peter Christen	974d58b01f	IPv6 Fix for push interface	10 years ago
Michael Peter Christen	fe50e5aef6	fix for failed selection of terms in faceted search with vocabularies	10 years ago
Michael Peter Christen	1309619a71	remove remote indexing option in crawl start if not in p2p mode	10 years ago
Michael Peter Christen	6324db1213	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	10 years ago
reger	5cb05c3013	adjust table column width to not line wrap crawler traffic line	10 years ago
Michael Peter Christen	606d00c8f2	cloning a crawl now accepts the class name of vocabulary scapers	10 years ago
reger	11b21308c0	fix: malformed filename in image search fix for http://mantis.tokeek.de/view.php?id=533	10 years ago
reger	9e1ec5fec4	refactor: just some more useages of constant for term ":[* TO *]"	10 years ago
Michael Peter Christen	b5ac29c9a5	added a html field scraper which reads text from html entities of a given css class and extends a given vocabulary with a term consisting with the text content of the html class tag. Additionally, the term is included into the semantic facet of the document. This allows the creation of faceted search to documents without the pre-creation of vocabularies; instead, the vocabulary is created on-the-fly, possibly for use in other crawls. If any of the term scraping for a specific vocabulary is successful on a document, this vocabulary is excluded for auto-annotation on the page. To use this feature, do the following: - create a vocabulary on /Vocabulary_p.html (if not existent) - in /CrawlStartExpert.html you will now see the vocabularies as column in a table. The second column provides text fields where you can name the class of html entities where the literal of the corresponding vocabulary shall be scraped out - when doing a search, you will see the content of the scraped fields in a navigation facet for the given vocabulary	10 years ago
Michael Peter Christen	68c605d637	replace with CommonPattern.SPACE for split	10 years ago
Michael Peter Christen	1f5047b15f	using precompiled pattern CommonPattern.SEMICOLON for splits	10 years ago
Michael Peter Christen	a8a2b7a803	persistency for vocabulary facet switch	10 years ago
Michael Peter Christen	efbc9a3561	introducting a new getConfig method which parses comma-separated llists from setting fields; refactoring for all places where such lists are parsed	10 years ago
Michael Peter Christen	69eacdf4eb	applying precompiled CommonPattern.COMMA.split to all places where split(",") was used	10 years ago
Michael Peter Christen	5a060c9f26	refactoring of reindexSolr (just replaced constant string)	10 years ago
Michael Peter Christen	3d717b749a	fix for urlmaskfilter	10 years ago
Michael Peter Christen	2636582435	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	10 years ago
reger	0260d3d800	Allow to hide linkstructure graphic in crawl monitor using/setting the config param DECORATION_GRAFICS_LINKSTRUCTURE	10 years ago
Michael Peter Christen	bee5ee7cce	removed some warnings	10 years ago
Michael Peter Christen	6390454652	fix for vocabulary on/off setting	10 years ago
Michael Peter Christen	29f6e9db7a	write java version to status page	10 years ago
Michael Peter Christen	7db2888336	fixed font size and print page generation in pdf snapshots	10 years ago
reger	24f68a4eb7	refactor opensearch heuristic introduce FederateSearchManager handling search heuristic to external systems via specific FederateSearchConnectors, which provide the query() functionallity, the translation to YaCy schema .toYaCySchema() and the search() routine to deliver results to searchevents, which is generally implemented in Abstract connector. The manager enforces now a min 15s delay between calls to external systems. Besides the OpensearchConnector a SolrFederateSearchConnector is available. It uses a additional config file for fieldname translation. default heuristicopensearch.conf: - openbdb.com removed - seems not longer to deliver results - config via solrconnector to datacite.org added (large technical library archive)	10 years ago
Michael Peter Christen	3b51636ecb	fix for mediawiki import	10 years ago
Michael Peter Christen	8cafdb989a	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	10 years ago
reger	4214f250d0	Add option for extended search (Autosearch) to Bookmark.html asking all connected peers for the searchterm added as description to the bookmark created by the bookmark icon. Intended for searches/research projects with not sufficient results from local and DHT selected remote target peers. Function: the process checks newly created bookmarks for description starting with "query=..." and takes this to ask every peer for 20 search results and adds it to the local index in a background job. link to start/stop the process added to /Bookmarks.html	10 years ago
reger	bb37cb32e4	Add title import for bookmark icon if avail in index	10 years ago

1 2 3 4 5 ...

5369 Commits (bbcd9441bc9108802231bf452e547f96825b9253)