An open source corporate search engine
Arch is an open source extension of apache nutch (a popular, highly scalable general purpose search engine) for intranet search. Not happy with your corporate search engine? Not surprising, very few people are. To the best of our knowledge, there are no intranet engines that work as well as the google s global web search does.
There is a fundamental reason for this: the algorithms used by google on the global web (or similar) do not work nearly as well on intranets for the lack of statistical data. Arch (finally. ) solves this problem. It uses a novel method to deliver high precision search results that works great.
Don t believe it? Blind test evaluation tools are included. You can deploy arch and compare its performance to your current search engine and/or google (on the public part of your site) using a blind test methodology.
In addition to the excellent search quality, arch has many features critical for corporate environments:
Document level security
Users can find only documents that they are authorized to see
Inexpensive index updates
Arch is able to keep indexes up to date and avoid regular complete site recrawling
There is always a working index available, even if a crawl fails
Support for simultaneous indexing and search of multiple web sites, with ability to search and administer any site separately, if needed
Dynamic adding and removal of web sites is easy
An automatically generated site directory
Low cost support once deployed
Dual interface (php and java) for easy deployment and customization
Faceted search 'out of the box'
An extensive and extensible set of parsers for parsing a variety of file formats: html, php, pdf, ms office, open office, etc
A modular, plugin based architecture that can be easily customized and extended
The source code is included
High performance and scalability
Arch can run on computer clusters to index very large data sets