orbiter
5dfd6359cb
redesign of the QueryParams class: introduced QueryGoal which holds the
...
query string parser. This shall be used to create a proper full-string
matching which is handled then by QueryGoal.
13 years ago
Michael Peter Christen
5fd3b93661
added deletion of hosts during crawl start if deleteold option was given
13 years ago
Michael Peter Christen
d64445c3cb
because we have the inurl:<term> - searchmodifier, we don't actually
...
need regular expressions as search attributes. They had now been removed
from the advanced search page while they are still created internally.
The filter is then expressed against solr as regular expression filter
query. If the expression points out a selection of an specific protocol,
host or filetype this is then translated into a facetted query.
13 years ago
orbiter
b55ea2197f
- redesign of crawl start servlet
...
- for domain-limited crawls, the domain is deleted now by default before
the crawl is started
13 years ago
orbiter
1c66de4bd4
- removed scheduled crawling options in crawl start because it is
...
superfluous there; it can be changed in the scheduler servlet. It's also
confusing in the presence of the delete-option, which will be
implemented next.
- removed unused crawl start servlet
- some refactoring to make the time parser reusable
13 years ago
Michael Peter Christen
2e7219f9fd
removed hightlighting of search results within collections in GSA
...
interface
13 years ago
Michael Peter Christen
074dfd297b
added icons and a selection for hosts with urls pending for crawler or
...
with errors
13 years ago
Michael Peter Christen
4c4e0eece2
added new submenu 'Target Analysis' with three servlets which are useful
...
to analyse the target servers: robots.txt table, mass target analysis
and a regex tester
13 years ago
Michael Peter Christen
61995d508e
do the commit anyway before calling a search interface
13 years ago
Michael Peter Christen
86ec199126
using a better file name
13 years ago
Michael Peter Christen
5105256927
update to search result logging (this was a remaining issue from the
...
solr 4.0.0 migration)
13 years ago
Michael Peter Christen
570e42c4e3
fix for filetype naviagtor
13 years ago
Michael Peter Christen
71ed8e5e07
bugfixes for crawler
13 years ago
Michael Peter Christen
29fbbb49dc
better colors for host browser and corrected document count
13 years ago
Michael Peter Christen
6244b084cd
fixed wrong order of result count values
13 years ago
Michael Peter Christen
631b08e7e2
update to HostBrowser
13 years ago
Michael Peter Christen
51f420e4f5
removed location search because it is only working in special cases
13 years ago
Michael Peter Christen
15d1460b40
added information about the reason of pausing of crawls
13 years ago
Michael Peter Christen
2371ef031c
added solr faceted search support to YaCy search results
...
added solr highlighting / YaCy snippets to YaCy search results
- facets are now much more complete
- facets are computed and searched much faster
- snippet computation is done by solr if solr knows the snippet
13 years ago
Michael Peter Christen
d481abd087
added the visualization of error-urls to host browser
...
- only visible for admins
- a faceted search generates a huge list for all hosts in the host list
- the faceted search algorithms had to be modified for that
- within the browsing of the directory path, the error cause is written
to the url which is presented as error-url
- the errors are also accumulated for directory sums
13 years ago
Michael Peter Christen
a15819fbec
fix for some interface problems
13 years ago
Michael Peter Christen
791e1dcfdf
when a new crawl is started, delete all entries about error-urls for
...
crawl-start domains
13 years ago
Michael Peter Christen
c6a6f4c4e6
added a hack which makes the HostBrowser more performant when the given
...
host has a lot of urls. If the number of urls is > 1000, then the list
of documents is restricted to such which have no subpath, if the root
path is selected. However, this can cause a problem if no documents on
the root path exist but only on paths below that root path.
13 years ago
Michael Peter Christen
64ac2b7b7d
new submenu template
13 years ago
Michael Peter Christen
5e77801aac
update to web interface structure
13 years ago
Michael Peter Christen
8fb370d9f8
renovated the way how search results are count. should be correct now...
13 years ago
orbiter
354ef8000d
- added 'deleteold' option to crawler which causes that documents are
...
deleted which are selected by a crawl filter (host or subpath)
- site crawl used this option be default now
- made option to deleteDomain() concurrency
13 years ago
Michael Peter Christen
19d1f474ce
host browser now shows also number of pending files per subdirectory +
...
bugfixes
13 years ago
Michael Peter Christen
75dd706e1b
update to HostBrowser:
...
- time-out after 3 seconds to speed up display (may be incomplete)
- showing also all links from the balancer queue in the host list (after
the '/') and in the result browser view with tag 'loading'
13 years ago
Michael Peter Christen
e2c4c3c7d3
migration to solr 4.0.0
13 years ago
Michael Peter Christen
9330ad4838
- fixed the delete option in host browser
...
- added a delete method which can be used to delete a full subpath in
solr.
13 years ago
Michael Peter Christen
40df2fd193
added the host browser as link to search results. that means you can
...
select a browsing position after a search is done on the search results.
13 years ago
Michael Peter Christen
1168d09de8
more refactoring - integrated the code of SnippetProcess into
...
SearchEvent
13 years ago
Michael Peter Christen
6629e37685
tried to clean up the search process mess
13 years ago
Michael Peter Christen
c5f67a5d6d
fixed a problem with local search from solr results: now all results
...
from solr are shown (again)
13 years ago
Michael Peter Christen
f8f05ecba7
- added a delete button in host browser to delete a complete subpath
...
- removed storage of default collection name - default is now "user"
- made stacking of crawl start points concurrently
13 years ago
Michael Peter Christen
0716a24737
added more / all new crawl profile fields into crawl profile editor
13 years ago
Michael Peter Christen
4a14122ba7
in case that a crawl profile has a collection assigned, use the
...
collection to show a name in the web interface. This should prevent that
much too long names make the interface unusable.
13 years ago
Michael Peter Christen
0fe8be7981
enhaced data structures for balancer and latency computation which
...
should produce a bit better prognosis about forced waiting times.
13 years ago
Michael Peter Christen
ac9540dfb6
removed options for stopwords which are not used
13 years ago
Michael Peter Christen
ce3fed8882
added the Google Search Appliance (GSA) api interface to the main menu.
...
See:
https://developers.google.com/search-appliance/documentation/68/xml_reference#request_overview
13 years ago
Michael Peter Christen
0833937c1c
better balancing and duetime-cumputation also for no-delay intranet
...
hosts
13 years ago
Michael Peter Christen
c25d7bcb80
- added concurrency for robots.txt loading
...
- changed data model for domain counter
13 years ago
Michael Peter Christen
a87811bc38
more auto-commit calls when a search interface is opened, but not when a
...
search is done there to prevent blocking during search-time.
13 years ago
Michael Peter Christen
3d3d654e88
if a network configuration is choosed which does not allow DHT and no
...
P2P communication is in robinson mode) then some menu entries are
disabled which have no use in this mode.
13 years ago
Michael Peter Christen
2d9e577ad0
replaced the custom robots.txt loader by the standard http loader
13 years ago
Michael Peter Christen
799d71bc67
enhanced solr caching:
...
- increased cache size which is needed for longer solr commit time
- speed hacks on cache write code
13 years ago
orbiter
8952153ecf
update to Balancer algorithm:
...
- create a load list from the current list of known hosts
- do not create this list for each Balancer.pop access
- create the list from those hosts which have a zero-waiting time
- select 1/3 from that list which have the most urls waiting
- get hosts from the wainting list in random order
- fixes for some delta-time computations
- always load all urls from hosts which have never been loaded before
13 years ago
Michael Peter Christen
8e1248ffe3
force a commit in advance of a search for the administrator to get most
...
recent results even if commit time is high and an indexing is ongoing.
13 years ago
Michael Peter Christen
1baf498d59
- show more lines in online log
...
- reverse order is default now
13 years ago