Commit Graph

388 Commits (bdf30117c1b86d044f41d352e411a6d23eaf1c65)

Author SHA1 Message Date
theli bdf30117c1 *) Redesign of parser configuration
19 years ago
theli d4ac3e25b1 *) Bugfix for file system link bug during detection of invalid URLs
19 years ago
orbiter adf75bc9fa better logging for invalid file path detection
19 years ago
orbiter 40621a5663 anhancements in ranking preparation and fixed problem with parser/mime recognition
19 years ago
theli c650b112ea *) Bugfix for relative URL Bug in Crawler
19 years ago
theli 4e73035aef *) Bugfix for "too many open files" during index distribution
19 years ago
orbiter f57e2d67f5 shortened network overview (less columns fit easier on page)
19 years ago
orbiter 85282b1d98 enhanced YBR recognition and search result heuristics
19 years ago
orbiter b9cc9029e3 added ybr selection for remote search
19 years ago
orbiter 0e25020f51 added first generation and usage of YBR index-files. Enhanced overall ranking of search results.
19 years ago
theli 90d6c6223b *) Adding color codes to network graphic legend
19 years ago
orbiter bfe51c7228 added generation of domain-list
19 years ago
orbiter 0ec54d9c5f enhanced CR-file handling and added first RCI-evaluation tests
19 years ago
theli c2fe3a1670 *) Updating jMimeMagic Ruleset
19 years ago
orbiter 88e3234393 fine-tuning of rci-generation
19 years ago
orbiter a12759c1bf first try to implement a rci-computation from cr-files
19 years ago
orbiter 4a8e8f269e refactoring of cr-processing; new kelondro class to handle the attribute file format
19 years ago
orbiter 24dc0e0760 implemented cr-file processing and further transmission steps
19 years ago
orbiter 9d9a87f445 limited htcache storage length
19 years ago
theli d0dfccdb77 *) Making CrawlStacker pool configurable via GUI and config file
19 years ago
theli 3631cb1f6d *) deleting empty entities during index selection
19 years ago
theli ca26aab9b1 *) More debugging output for migrateWords
19 years ago
theli 9b35ae9027 *) Correcting wrong % values on IndexTransfer_p page
19 years ago
theli e6bf9d90a5 *) Fixing Problems with MalformedURLs during Word Selection
19 years ago
theli 86a9210264 *) indexing queue slots are now configurable via config file
19 years ago
theli 3c11d7b81c *) Bugfix for minimizeUrlDB
19 years ago
orbiter 9913049009 fixed outOfMemory bug caused by loops in kelondroTree during enumeration
19 years ago
theli bbb936b9ea *) Bugfix for not human readable content of PDFs while viewing the URL Content via GUI
19 years ago
theli 445e3a620f *) Avoid rejecting of html content by the crawler when the file extension is not set properly
19 years ago
theli 444a5a9368 *) Bugfix for Entries with null url in GlobalQueue
19 years ago
borg-0300 ebac51df52 restore defaultRemoteProfile
20 years ago
borg-0300 5778428455 move cutUrlText to nxTools,
20 years ago
borg-0300 9158845c3b bugfix for snippet text null bytes
20 years ago
orbiter f763923e0a added missing files for last commit
20 years ago
orbiter 79818a320f introduced citation-rank transmission protocol and activate transport for anonymisation
20 years ago
theli 7e0647f692 *) Bugfix for userDB usage during authentication
20 years ago
orbiter 02f8013013 auto-delete of corrupted word files during word-migration
20 years ago
orbiter d2731418bf added creation of global ranking files and changed url normal form usage
20 years ago
theli 6f9f8ed8f8 *) Automatic Reset of Stack Crawler DB on startup errors
20 years ago
theli fb766413d1 *) Changes on httpc dns caching
20 years ago
orbiter bc420c62f6 fixed htcache path generation (never change a running system)
20 years ago
theli dd24f0252f *) Searchword highlighting for info page
20 years ago
borg-0300 72cde1d894 getCachePath: no logging
20 years ago
borg-0300 1fbd72f9e0 rename "index.html" to "ndx"
20 years ago
borg-0300 cd1107d85e added support for URLs with '?&'
20 years ago
borg-0300 5fb2b017cb small change
20 years ago
borg-0300 544e4ea90e small change
20 years ago
borg-0300 00ab4d8723 cleaned, small change, Properties
20 years ago
theli b8ceb1ffde *) Adding better https support for crawler
20 years ago
borg-0300 e3179a6394 added getOwnSeedFile()
20 years ago