reger
3e742d1e34
Init remote crawler on demand
...
If remote crawl option is not activated, skip init of remoteCrawlJob to save the resources of queue and ideling thread.
Deploy of the remoteCrawlJob deferred on activation of the option.
10 years ago
Michael Peter Christen
dbf9e3503d
Merge branch 'master' of git@github.com:yacy/yacy_search_server.git
10 years ago
Michael Peter Christen
8b1a30be50
removed a -UNRESOLVED_PATTERN-
10 years ago
Michael Peter Christen
9938c81378
fix for division by zero
10 years ago
reger
13f013f64a
Limit extra sleep of BusyThread on LowMemCycle
10 years ago
reger
cd7c0e0aae
detail optimization of RecrawlThread
10 years ago
reger
ace71a8877
Initial (experimental) implementation of index update/re-crawl job
...
added to IndexReIndexMonitor_p.html
Selects existing documents from index and feeds it to the crawler.
currently only the field fresh_date_dt is used determine documents for recrawl (fresh_date_dt:[* TO NOW-1DAY]
Documents are added in small chunks (200) to the crawler, only if no other crawl is running.
10 years ago
reger
141cd80456
correct log msg text
10 years ago
reger
f3ce99bfb8
fix extract of inboundlinks_protocol_sxt
...
url counter maybe > 999
10 years ago
reger
2bc9cb5828
fix early return in addToCrawler
...
check / handle all supplied urls after error url
10 years ago
Michael Peter Christen
f5f88272e4
Merge branch 'master' of git@github.com:yacy/yacy_search_server.git
10 years ago
Michael Peter Christen
5c67c4d460
fix for latest commit, see
...
f810915717 (commitcomment-11145880)
10 years ago
reger
c37dda8849
fix NPE on MultiProtocolURL on url with parameter value and '='
...
in getAttribute
- added test case for it
10 years ago
Michael Peter Christen
f810915717
added crawl start from a clone with very, very large url: they are now
...
encoded as post submit form inside a javascript creation function.
10 years ago
Michael Peter Christen
51de86c992
disabled debug thread dumps
10 years ago
Michael Peter Christen
d524a9d77c
Merge branch 'master' of git@github.com:yacy/yacy_search_server.git
10 years ago
Michael Peter Christen
0710648c31
enable api calls with very long urls
10 years ago
reger
31346e873b
upd library reference of missing jsch-0.1.21 in seeduploadscp.xml
...
upd to jsch-0.1.52.jar
10 years ago
reger
609c52e987
refactor getBookmark
...
to consistenly check existance by != null (w/o throwing exception on not found)
10 years ago
reger
1481a8ab56
add opensearch rss results to dht collection (due to text = snippet)
...
which is used to differentiate meta from full data
- make sure check for dht is not dependant on number of collection entries
10 years ago
reger
5f4d35437e
add bookmark.query to edit form
10 years ago
reger
f134aa7f7f
persist bookmark timestamp
...
on setTimeStamp()
10 years ago
reger
752eec6697
fix NPE in addToIndex when used outside searchEvent
10 years ago
reger
a6daddbeaa
upd to commons-io-2.4.jar
10 years ago
reger
89124335c4
update bookmark autosearch description
...
- add german translation
10 years ago
Michael Peter Christen
fbf85a1561
added temporary debug output in http client
10 years ago
Michael Peter Christen
ff29b0e503
added option to re-index exported xml snapshot dumps to
...
HTCACHE/snapshots by just placing them in the SURROGATES/in path
10 years ago
Michael Peter Christen
6f4fe4b175
revert of 8a7c68e4c7
...
keeping surrogates after processing is essential for some users. If the
space they are taking is too high, please set up an automatic deletion
process (like a cronjob).
10 years ago
Michael Peter Christen
213401a446
Merge branch 'master' of git@github.com:yacy/yacy_search_server.git
10 years ago
Michael Peter Christen
97930a6aad
added must-not-match filter to snapshot generation.
...
also: fixed some bugs
10 years ago
Michael Peter Christen
9d8f426890
adding a try-catch to link graph processing to prevent that a single
...
malformed url interrupts the storage process
10 years ago
reger
b47267b79c
precaution against NPE on createorgetBookmark on search result
10 years ago
Michael Peter Christen
75879e051b
Merge branch 'master' of git@github.com:yacy/yacy_search_server.git
10 years ago
reger
8a5b8f8789
on bookmaring of search result, remember orig. query in separate bookmark property
...
(instead of using the description field)
- adjust display and autosearch
- don't overwrite existing bookmark but combine info
10 years ago
reger
7224209486
break out of NormalizeDistributor loop on timeout
10 years ago
reger
cf1fc7f700
harmonize filesearch input box layout
10 years ago
reger
4d73e9de06
upd to metadata-extractor-2.8.1
10 years ago
Michael Peter Christen
e334a06370
Merge branch 'master' of git@github.com:yacy/yacy_search_server.git
10 years ago
reger
0904a041a6
upd to poi-3.11.jar
10 years ago
reger
47e61f8325
fix typo in image filter query
...
(extra bracket)
10 years ago
reger
4b4ab6799f
fix String out of range in Collection Nav
...
see http://mantis.tokeek.de/view.php?id=573
10 years ago
reger
572cfe8fd4
improve character encoding for urlproxy servlet
...
for none utf-8 pages
10 years ago
reger
b161473cd0
upd to jsoup-1.8.2
10 years ago
reger
6bc8a9b11e
make Quality of Service Servlet available to prioritize requests from local host
...
This assigns priorities to incoming requests. Higher priority numbers are served before lower.
(disabled by default in defaults/web.xml,
uncomment or copy entry to DATA/Settings/web.xml)
10 years ago
reger
af2d66e3d8
correct typo in de.lng
10 years ago
reger
71bf95af8a
upd parser calls in test cases
10 years ago
reger
579303a04e
add additional links to crawl queue pages
10 years ago
Michael Peter Christen
99718dc09a
don't record dump generation calls since that
...
- is not a change of the index
- happens very often within self-backup strategies from the outside
(i.e. cronjobs)
10 years ago
Michael Peter Christen
5b59477415
update to bootstrap.css 3.3.4
10 years ago
Michael Peter Christen
016b4e58ac
Merge pull request #4 from dertuxmalwieder/master
...
Readme improvements
10 years ago