yacy_search_server

Commit Graph

Author	SHA1	Message	Date
Michael Peter Christen	0aa6fcf259	remove old vocabularies and synonyms before adding new	9 years ago
Michael Peter Christen	289018b559	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	9 years ago
Michael Peter Christen	7b412e8c07	added msg (text emails) format; should be handled by html parser.	9 years ago
reger	f91298d3b6	fix one implicit Integer/Long type conversion -> causes Java 1.8 compile error	9 years ago
reger	821262a179	add CommonPattern for multiple spaces to eliminate empty split words on following spaces	10 years ago
Michael Peter Christen	90f75c8c3d	added enrichment of synonyms and vocabularies for imported documents during surrogate reading: those attributes from the dump are removed during the import process and replaced by new detected attributes according to the setting of the YaCy peer. This may cause that all such attributes are removed if the importing peer has no synonyms and/or no vocabularies defined.	10 years ago
Michael Peter Christen	7829480b82	refactoring: separated condenser and tokenizer	10 years ago
Michael Peter Christen	593de05922	enhanced surrogate import process speed (dramatically!)	10 years ago
Michael Peter Christen	3c4c69adea	fix for - bad regex computation for crawl start from file (limitation on domain did not work) - servlet error when starting crawl from a large list of urls	10 years ago
Michael Peter Christen	1fec7fb3c1	suppress access to solr when doing search suggestions in case that the index has more than two million documents. This protects the index from beeing flooded with search requests that cannot be resolved before the real search query has to be computet.	10 years ago
Michael Peter Christen	694b22f165	migration to Solr 5.2: huge benefits - this is a lot faster! This is a very complex migration: many classes had been renamed or removed, dependencies changed and the solr index type is now aligned to be a solr cloud repository. Together with the Solr 5.2 library update, one other dependent library had been updated as well: httpclient 4.4->4.4.1 Older indexes are migrated from 4_10 to 5_2. However, the new index structure is more efficient and we recommend to re-index everything. Please use the index export before you do the update to a large surrogate xml file. After the update, start with an empty index and then initialize this with your dump.	10 years ago
sixcooler	e427efbe54	Next Try for a fix for upload-connection staying in blocked state. This was caused by reading via GZIP from close-wait connection an caused high cpu- and system-loads. Instat of implementing handling of the RedListener now I found a timelimeted 'get' "realy" solving this problem.	10 years ago
reger	0fab445b19	Resourceobserver log warning - deleting releases files - only on actual deletes instead of entering routine	10 years ago
sixcooler	ef6a64b2a4	Fix for upload-connection staying in blocked state. This was caused by reading via GZIP from close-wait connection an caused high cpu- and system-loads. Solved by implementing handling of the RedListener.	10 years ago
reger	c973f94936	add log entry on release file delete by ResourceObserver	10 years ago
reger	121972752c	implement deleteOldDownloads in RexourceObserver on low diskspace - direct assign sb.observer (skip redundant InitThread)	10 years ago
Michael Peter Christen	9c12555be5	added link to Snapshots in search results if the snapshot exists and option is set in ConfigSearchPage_p (this is a stub: we also need a visualization of pdf files!)	10 years ago
reger	72f6a0b0b2	enhance recrawl job - allow to modify the query to select documents to process (after job has started) - allow to include failed urls (httpstatus <> 200)	10 years ago
reger	7478338a40	remove augmented parsing activation from frontend experimental implementation not used and based on error prone experimental rdfaparser	10 years ago
reger	11aa2edfe1	remove RDFa parser activation from frontend reason: experimental implementatin of RDFa parser not executed (limited to special urls) but may cause error on normal html parsing due to a inputstream.reset	10 years ago
reger	49b79987c9	remove obsolete searchfl work table was used to register urls with not complete words in snippet but is never accessed	10 years ago
Michael Peter Christen	d0aff91f23	fix for index import	10 years ago
Michael Peter Christen	34de1e8cbc	gzip compression will perform more efficient and with better compression level	10 years ago
Michael Peter Christen	98be59ce9c	full solr xml exports will now be automatically compressed during export. That makes it possible to export a solr xml dump even if disc space is low.	10 years ago
Michael Peter Christen	a1a8edfc0a	wrap HeaReader close() in a catch Throwable block to prevent that an excpetion during close blocks the whole shotdown process	10 years ago
Michael Peter Christen	b43811d38c	added surrogate import process for exported solr dumps. Just throw your solr dump file into DATA/SURROGATES/in/ and it will be imported!	10 years ago
Michael Peter Christen	b77537294d	prevent disc usage when showing tray animation	10 years ago
Michael Peter Christen	eec78e1b0c	added intensity option to graphics	10 years ago
Michael Peter Christen	a5007f345e	re-licensing some of my old visualization classes under LGPL 2.1	10 years ago
Michael Peter Christen	c99a665593	adding a 3-pixel font generator made some time ago..	10 years ago
Michael Peter Christen	c7576d6028	added a full solr export to the IndexControlURLs_p.html servlet. The export function is also now the default export option. The export file format for a full solr export is very similar to a solr search result xml, only the <lst name="responseHeader"> tag is missing. The exported xml has a special line termination feature: all documents will be exported into a single line without any CR in between. That means that every document is completely inside a single line. While this is not readable at all for humans, it is very useful for linux line processing scripts, like grep. Using grep it will be easy to select single documents which match for a given pattern. Such dumps shall be importable with the DATA/SURROGATE/in import function, but that import is not yet adopted to the new file format.	10 years ago
Michael Peter Christen	197f7449e5	All entities of crawl profiles are now editable in the crawl profile editor.	10 years ago
reger	1d8e1e4bac	- Image search expand box, adjust javascript hs padtominsize parameter, to make sure expand box doesn't shrink on small images - asure ImageResult.imagetext has value for the link text (use filename if no alt text given)	10 years ago
reger	8b35656007	remove hard throw exception in makeResultEntry remove not used "share." peername.yacy url rewrite	10 years ago
reger	af57fbefad	use available mime (instead null) on imageresult from metadatanode	10 years ago
reger	dd7782bac0	revert deletion of BinSearch (accident)	10 years ago
reger	000dde9511	Eleminate duplication of values for search ResultEntry by instatiation from URIMetadataNode, by eleminating differentiation of ResultEntry/URIMetadataNode. - moved remaining ResultEntry functionallity to URIMetadataNode - for 1:1 functionallity added a function makeResultEntry() - removed ResultEntry - refactored related code Main difference is after makeResultEntry the text_t content is removed and alternative title/url strings for display are calculated. Main difference left is, that	10 years ago
reger	29c4aa3991	fix compiler notification of missing serialID from last commit	10 years ago
reger	3d53da8236	refactor ResultEntry to be based on MetadataNode/SolrDocument to share/reuse common access routines	10 years ago
reger	d882991bc5	Implement sharing of ioDispatcher for term & citation index as proposed in ioDispatcher description	10 years ago
reger	370ba9da71	On imageSearch prefere mime to sort out none-image documents Generalize the hack to prevent urls with just a img extension beeing returned improving http://mantis.tokeek.de/view.php?id=528	10 years ago
reger	cd31633369	improve MultiprotocolURL.getFileExtension() prevent string OOB while querypart contains a dot (return just "") see log snippet in http://mantis.tokeek.de/view.php?id=533	10 years ago
reger	c60ccdfbcf	Increase IODspatcher dumpQueue size to 2 to reduce risk of concurrent emergency dump, skip concurrent emergency merge dealing with/see http://mantis.tokeek.de/view.php?id=566	10 years ago
reger	8a9622c31c	fix string OoB on getImagelinks with long alttext in description calculation	10 years ago
reger	3e742d1e34	Init remote crawler on demand If remote crawl option is not activated, skip init of remoteCrawlJob to save the resources of queue and ideling thread. Deploy of the remoteCrawlJob deferred on activation of the option.	10 years ago
reger	13f013f64a	Limit extra sleep of BusyThread on LowMemCycle	10 years ago
reger	cd7c0e0aae	detail optimization of RecrawlThread	10 years ago
reger	ace71a8877	Initial (experimental) implementation of index update/re-crawl job added to IndexReIndexMonitor_p.html Selects existing documents from index and feeds it to the crawler. currently only the field fresh_date_dt is used determine documents for recrawl (fresh_date_dt:[* TO NOW-1DAY] Documents are added in small chunks (200) to the crawler, only if no other crawl is running.	10 years ago
reger	141cd80456	correct log msg text	10 years ago
reger	f3ce99bfb8	fix extract of inboundlinks_protocol_sxt url counter maybe > 999	10 years ago
reger	2bc9cb5828	fix early return in addToCrawler check / handle all supplied urls after error url	10 years ago
Michael Peter Christen	f5f88272e4	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	10 years ago
Michael Peter Christen	5c67c4d460	fix for latest commit, see `f810915717 (commitcomment-11145880)`	10 years ago
reger	c37dda8849	fix NPE on MultiProtocolURL on url with parameter value and '=' in getAttribute - added test case for it	10 years ago
Michael Peter Christen	f810915717	added crawl start from a clone with very, very large url: they are now encoded as post submit form inside a javascript creation function.	10 years ago
Michael Peter Christen	51de86c992	disabled debug thread dumps	10 years ago
Michael Peter Christen	d524a9d77c	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	10 years ago
Michael Peter Christen	0710648c31	enable api calls with very long urls	10 years ago
reger	31346e873b	upd library reference of missing jsch-0.1.21 in seeduploadscp.xml upd to jsch-0.1.52.jar	10 years ago
reger	609c52e987	refactor getBookmark to consistenly check existance by != null (w/o throwing exception on not found)	10 years ago
reger	1481a8ab56	add opensearch rss results to dht collection (due to text = snippet) which is used to differentiate meta from full data - make sure check for dht is not dependant on number of collection entries	10 years ago
reger	f134aa7f7f	persist bookmark timestamp on setTimeStamp()	10 years ago
reger	752eec6697	fix NPE in addToIndex when used outside searchEvent	10 years ago
Michael Peter Christen	fbf85a1561	added temporary debug output in http client	10 years ago
Michael Peter Christen	ff29b0e503	added option to re-index exported xml snapshot dumps to HTCACHE/snapshots by just placing them in the SURROGATES/in path	10 years ago
Michael Peter Christen	6f4fe4b175	revert of `8a7c68e4c7` keeping surrogates after processing is essential for some users. If the space they are taking is too high, please set up an automatic deletion process (like a cronjob).	10 years ago
Michael Peter Christen	97930a6aad	added must-not-match filter to snapshot generation. also: fixed some bugs	10 years ago
Michael Peter Christen	9d8f426890	adding a try-catch to link graph processing to prevent that a single malformed url interrupts the storage process	10 years ago
reger	8a5b8f8789	on bookmaring of search result, remember orig. query in separate bookmark property (instead of using the description field) - adjust display and autosearch - don't overwrite existing bookmark but combine info	10 years ago
reger	7224209486	break out of NormalizeDistributor loop on timeout	10 years ago
reger	47e61f8325	fix typo in image filter query (extra bracket)	10 years ago
reger	4b4ab6799f	fix String out of range in Collection Nav see http://mantis.tokeek.de/view.php?id=573	10 years ago
reger	572cfe8fd4	improve character encoding for urlproxy servlet for none utf-8 pages	10 years ago
reger	6bc8a9b11e	make Quality of Service Servlet available to prioritize requests from local host This assigns priorities to incoming requests. Higher priority numbers are served before lower. (disabled by default in defaults/web.xml, uncomment or copy entry to DATA/Settings/web.xml)	10 years ago
Ryszard Goń	ca1a70aec8	fix for Accept '?' URLs column in Crawl Profile List	10 years ago
reger	5408448a56	skip redundant add. of keywords to text search uses keywords as default search field	10 years ago
reger	296e97c78e	put https port in peers dna as we flag if a peer is accesible via https, we need to know the port if we want to use is (e.g. for interYaCy communication) start to provide / tansport the port by recording it in peers dna. - add https link on the Network.html lock symbol	10 years ago
Michael Peter Christen	fed26f33a8	enhanced timezone managament for indexed data: to support the new time parser and search functions in YaCy a high precision detection of date and time on the day is necessary. That requires that the time zone of the document content and the time zone of the user, doing a search, is detected. The time zone of the search request is done automatically using the browsers time zone offset which is delivered to the search request automatically and invisible to the user. The time zone for the content of web pages cannot be detected automatically and must be an attribute of crawl starts. The advanced crawl start now provides an input field to set the time zone in minutes as an offset number. All parsers must get a time zone offset passed, so this required the change of the parser java api. A lot of other changes had been made which corrects the wrong handling of dates in YaCy which was to add a correction based on the time zone of the server. Now no correction is added and all dates in YaCy are UTC/GMT time zone, a normalized time zone for all peers.	10 years ago
Michael Peter Christen	b060ba900d	added parsing of contentprop attribute in html tags for content='startDate' and content='endDate'. The value of these field is now written to new solr fields startDates_dts and endDates_dts.	10 years ago
Michael Peter Christen	4cb4f67f38	added parsing of dd, dt and article html fields. The parsed result is written to special solr fields which are deactivated by default.	10 years ago
reger	1395f10e95	fix typecast for css links	10 years ago
Michael Peter Christen	3288489fd2	more logging during start-up	10 years ago
Michael Peter Christen	abaaaef5f1	fix for filter queries	10 years ago
Michael Peter Christen	4d00175157	<experimental> added parsing of <article> html element. Whenever such an element occurs, the complete content of all article elements replaces the parsed <content> part of documents.	10 years ago
Michael Peter Christen	1df6492019	enhanced suggestions	10 years ago
Michael Peter Christen	ae02c92fd0	logging fix	10 years ago
Michael Peter Christen	5651713134	better debugging of fq	10 years ago
Michael Peter Christen	f5a032f293	split query into filter query and text query to get better ranking results and faster results	10 years ago
Michael Peter Christen	2e88028c1a	when selecting collections in navigation, do show the un-selected collections in search result. When selecting one of them in another search, switch off the previously selected collection. This actually turns the collection navigation modifier into a radio-button like behaviour	10 years ago
Michael Peter Christen	1de9b21c65	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	10 years ago
reger	5f4cd8d6f5	replace deprecated getIP with getIPs in AbstractRemoteHandler	10 years ago
Michael Peter Christen	fa7edc9f7a	refactoring of filter queries (several queries instead only one)	10 years ago
Michael Peter Christen	40389987ec	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	10 years ago
Michael Peter Christen	f9ba50379d	added an expansion option to search facets on result page: - if less or equal of 8 facet options are present, they are shown by default - if more facet options are present, they are hidden To view or hide all facets, just click on the facet header bar	10 years ago
reger	1f0f77bb77	make location facet return results for location nav facet of field coordinate_p does not return results, now using coordinate_p_0_coordinate as alternative to get facet counts. As the actual facet value is not used this should not harm any analysis (even if facet is a incomplete location). If facet value is used in future likely *_geohash field could be introduced (for facet and other ... as transport value)	10 years ago
reger	b1ec0644e5	fix NPE in location search on missing/empty PubDate in underlaying rss data	10 years ago
reger	c1dcc8c456	fix display and limit of max server connections after startup (on restart value returned to default=50) This has no effect on Jetty but the limit is still respected.	10 years ago
reger	839b962c20	correct percent encoding for '%' char	10 years ago
Michael Peter Christen	9bf0d7ecb9	added a new collection type 'dht' to all documents from the peer-to-peer interface to distinguish rich and poor document data. This also reverts some changes from commit `796770e070` because the firstSeen database is the wrong method to distinguish these types of data	10 years ago
reger	796770e070	prevent overwrite of crawled or received full documents by (newer) metadata To protect rich index data (full resource) from overwriting by metadata gathered during remote search, the newly introduced "firstSeen" index is used to differentiate between full-resource-doc and metadata, as a "firstSeen" entry is only added on store's of full-resource-docs (during crawl or remote search).	10 years ago
Michael Peter Christen	ee2490ab98	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	10 years ago
reger	431311df42	fix get fresh_date_dt to allow returned value to be date in future	10 years ago
otter	74c7e8b686	Fixes hanging FlushThread (see http://forum.yacy-websuche.de/viewtopic.php?f=5&t=5447) by replacing put() method by the more robust add() to add a merge job to the queue.	10 years ago
reger	f63fff9008	fix snippet containig number with comma as desmo point http://mantis.tokeek.de/view.php?id=344 to keep it as one word (by altering the split regex) - added sniipet test case with number - regex for word split to match multiple splitcars	10 years ago
reger	b241264632	fix error on *abc query input http://mantis.tokeek.de/view.php?id=486	10 years ago
reger	2ef8ffdb60	apply UTF-8 encoding copied from escape()	10 years ago
reger	7120ea42f1	fix for path with char code > 255 (causing index out of bound exception) + test cas for it	10 years ago
reger	1d81bd0687	fix url encoding for path see http://mantis.tokeek.de/view.php?id=559 So far we used same escape procedure for all parts of the url (which includes x-www-form-urlencoded for all url components) Added capability to use different encoding rules for the different url components (through specific bitset for each component). (this is inspired by org.apache.http.client and java.net.uri implementation). - Added test case for http://mantis.tokeek.de/view.php?id=559	10 years ago
reger	62087fb8b2	fix MultiProtocolURL mailto protocol detection	10 years ago
reger	2e8c24e02a	fix link to DeReWo download file	10 years ago
reger	706f75ddc2	try to fix hang on index blob merge on shutdown http://mantis.tokeek.de/view.php?id=505 It happens but not able to reproduce. This change makes sure terminate signal is catched at end of currently running merge jobs	10 years ago
reger	f94e34058c	fix url (path) %-decoding http://mantis.tokeek.de/view.php?id=519 - add test case for this	10 years ago
reger	7e09bff4a1	exclude default search fields from text copy to text_t for metadata index documents (reduce text redundance)	10 years ago
reger	86073a5ba3	For remote crawlReceipt add document abstract/description enhance the returned metadata returned to the originator by description_txt to improve fulltext search result hits.	10 years ago
reger	8af70950d9	harmonize snippet computation to considere description_txt always (solr hl & internal). For now just added desc to text list for computation, could be further equalized with hl computation.	10 years ago
Michael Peter Christen	fd4e2c809a	Show dates in the content of a document in the search result: - if an eventDate is given in the search result, replace the document date with the event date and prefix it with the string "on ". - the document date is omitted if a date from the cent is shown Added also the date as fields in the json and rss result sets.	10 years ago
Michael Peter Christen	893889bc7b	added special terms for on: - Date modifier: tomorrow, today; i.e.: search for: "Berlin on:tomorrow" to find events happening tomorrow in Berlin	10 years ago
Michael Peter Christen	710a0efa1b	generalized time period computations	10 years ago
Michael Peter Christen	d9d3111d10	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	10 years ago
Michael Peter Christen	535f1ebe3b	added a new way of content browsing in search results: - date navigation The date is taken from the CONTENT of the documents / web pages, NOT from a date submitted in the context of metadata (i.e. http header or html head form). This makes it possible to search for documents in the future, i.e. when documents contain event descriptions for future events. The date is written to an index field which is now enabled by default. All documents are scanned for contained date mentions. To visualize the dates for a specific search results, a histogram showing the number of documents for each day is displayed. To render these histograms the morris.js library is used. Morris.js requires also raphael.js which is now also integrated in YaCy. The histogram is now also displayed in the index browser by default. To select a specific range from a search result, the following modifiers had been introduced: from:<date> to:<date> These modifiers can be used separately (i.e. only 'from' or only 'to') to describe an open interval or combined to have a closed interval. Both dates are inclusive. To select a specific single date only, use the 'to:' - modifier. The histogram shows blue and green lines; the green lines denot weekend days (saturday and sunday). Clicking on bars in the histogram has the following reaction: 1st click: add a from:<date> modifier for the date of the bar 2nd click: add a to:<date> modifier for the date of the bar 3rd click: remove from and date modifier and set a on:<date> for the bar When the on:<date> modifier is used, the histogram shows an unlimited time period. This makes it possible to click again (4th click) which is then interpreted as a 1st click again (sets a from modifier). The display feature is NOT switched on by default; to switch it on use the /ConfigSearchPage_p.html servlet.	10 years ago
reger	d7259419f3	postpone raw snippet html encoding upon use instead of during init of snippet adressing http://mantis.tokeek.de/view.php?id=551	10 years ago
reger	de56d934b2	apply query parameter getQueryFields() to GSA servlet	10 years ago
reger	2d2299f484	fix mimetype of rss items in rss parser - remove self reference as anchor for items	10 years ago
Michael Peter Christen	b432049d59	enhanced date parsing time	10 years ago
reger	9b0de2de64	introduce getQueryFields to return default query fields (queryparamter QF) calculated from boostfields config, making sure title, description, keywords and content is always searched. - apply change to solrServlet makes sure every remote query uses at least all locally defined boost fields for search - apply to local solr search - simplify select query by using QF defaults	10 years ago
reger	a0f04db9ea	add extracted description/subject to pptParser	10 years ago
reger	8ec1db76ee	url unescape add check for inconsistent utf8 multibyte parsing If the url contains special chars (like umlaute äöü) it's interpreted as multybyte char and actually not converted at all (removed). Added a check if the multibyte convesion is not complete, just add the char as is. This fixes http://mantis.tokeek.de/view.php?id=200	10 years ago
reger	4b97ddb9ec	stop sending crawl receipts if receiver got offline	10 years ago
reger	7e35518787	add extracted description/subject to docParser	10 years ago
reger	f0a5188e11	replace depreciated HTTPClient setStaleConnectionCheckEnabled with setValidateAfterInactivity()	10 years ago
reger	7b569d2dbe	replace depriciated HTTPClient ALLOW_ALL_HOSTNAME_VERIFIER with NoopHostnameVerifier()	10 years ago
reger	fba34e12ef	fix formatting issue if snippet contains html code replacement for reverted commit `61f42a7928`	10 years ago
reger	e48720a58c	fix NPE in snippet computation	10 years ago
reger	eda0aeaf26	allow/recognize host in file: protocol crawl target This is useful in intranet indexing while crawling a intranet file server accessed via hostname while e.g. under Windows mapped to different drive letters on individual clients. Here you can crawl e.g. file://fileserver/documents having a valid uri in that intranet environment (while e.g. P:/documents might be client dependant).	10 years ago
reger	df83fcc4fc	disable optimistic GC assumption in StandardMemoryStrategy After several tests found that eom is not prevented. Major reason in testing was assumption future GC will free avg of last 5 GC. Disabeling this check improved eom exceptions. Added simplest testcase used for verification	10 years ago
Michael Peter Christen	8ff76f8682	the cleanup process experienced a 100% CPU load situation and the loop did not terminate: Occurrences: 100 at java.util.HashMap$KeyIterator.next(HashMap.java:956) at net.yacy.cora.protocol.ConnectionInfo.cleanup(ConnectionInfo.java:300) at net.yacy.cora.protocol.ConnectionInfo.cleanUp(ConnectionInfo.java:293) at net.yacy.search.Switchboard.cleanupJob(Switchboard.java:2212) at sun.reflect.GeneratedMethodAccessor12.invoke(Unknown Source) at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43) at java.lang.reflect.Method.invoke(Method.java:606) at net.yacy.kelondro.workflow.InstantBusyThread.job(InstantBusyThread.java:105) at net.yacy.kelondro.workflow.AbstractBusyThread.run(AbstractBusyThread.java:215) This tries to fix the problem; the problem should be monitored	10 years ago
Michael Peter Christen	1f5b5c0111	npe fix for latest scraper feature	10 years ago
Michael Peter Christen	ee97302a23	hack to make date detection faster (while it becomes a bit incomplete regarding language alternatives)	10 years ago
Michael Peter Christen	6578ff3ddb	enhanced suggest function	10 years ago
reger	fe6f5a395d	fix Umlaut handling in blekko heuristic search term http://mantis.tokeek.de/view.php?id=169 observation: blekko seams to block xxxbot agents (=0 results)	10 years ago
reger	23924348e2	url with semicolon or comma handling in proxy request apply patch supplied with bugreport http://mantis.tokeek.de/view.php?id=540	10 years ago
reger	9025fe3518	upd error message for proxy fix http://mantis.tokeek.de/view.php?id=539	10 years ago
Michael Peter Christen	97ba5ddbb7	configuration option for maxload limit for remote search	10 years ago
reger	c454ef69c6	add shortMemory check to heuristic search and skip operation on shortMemory (no request to remote openserch systems)	10 years ago
reger	9e1ec5fec4	refactor: just some more useages of constant for term ":[* TO *]"	10 years ago
reger	8c491f51a5	remove hardcoded initialization of language nav if not used	10 years ago
Michael Peter Christen	b5ac29c9a5	added a html field scraper which reads text from html entities of a given css class and extends a given vocabulary with a term consisting with the text content of the html class tag. Additionally, the term is included into the semantic facet of the document. This allows the creation of faceted search to documents without the pre-creation of vocabularies; instead, the vocabulary is created on-the-fly, possibly for use in other crawls. If any of the term scraping for a specific vocabulary is successful on a document, this vocabulary is excluded for auto-annotation on the page. To use this feature, do the following: - create a vocabulary on /Vocabulary_p.html (if not existent) - in /CrawlStartExpert.html you will now see the vocabularies as column in a table. The second column provides text fields where you can name the class of html entities where the literal of the corresponding vocabulary shall be scraped out - when doing a search, you will see the content of the scraped fields in a navigation facet for the given vocabulary	10 years ago
Michael Peter Christen	1cb290170e	refactoring of autotagging code (combined same code pieces)	10 years ago
Michael Peter Christen	c3b55455fc	enhanced initialization speed of vocabularies by using better normalization and by removal of unused data structures	10 years ago
Michael Peter Christen	68c605d637	replace with CommonPattern.SPACE for split	10 years ago

1 2 3 4 5 ...

7843 Commits (5d71fc70e3f21b46a87f4f3edd708196e11f8caf)