On Dec 16, 5:26pm, rishab@dxm.ernet.in wrote:
The alternative comes strangely enough from the US National Security Agency. It has always been suspected that the NSA searches through e-mail traffic for 'sensitive' material, which for the large volumes involved would require considerable help from computers.
I always imagined that the development of the algorithm itself predated email, and started back with cable and telex traffic.
Earlier this year, the agency began soliciting collaborations from business to develop commercial applications of their technique. It claimed to be able to quickly search through large quantities of text, in any language, for similarities to sample documents, and even automatically sort documents according to topics that it identifies. A similar though independently developed system is available from California-based Architext.
Has anyone got any further information about how this algorithm works? It sounds like Rishab has somewhat better info than was publicly available months ago when we last discussed this particular NSA "technology transfer". Ian.