Semantic Vectors: a Scalable Open Source Package and Online Technology Management Application
Dominic Widdows | Kathleen Ferraro
Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)

This paper describes the open source SemanticVectors package that efficiently creates semantic vectors for words and documents from a corpus of free text articles. We believe that this package can play an important role in furthering research in distributional semantics, and (perhaps more importantly) can help to significantly reduce the current gap that exists between good research results and valuable applications in production software. Two clear principles that have guided the creation of the package so far include ease-of-use and scalability. The basic package installs and runs easily on any Java-enabled platform, and depends only on Apache Lucene. Dimension reduction is performed using Random Projection, which enables the system to scale much more effectively than other algorithms used for the same purpose. This paper also describes a trial application in the Technology Management domain, which highlights some user-centred design challenges which we believe are also key to successful deployment of this technology.


Empirical Study of Predictive Powers of Simple Attachment Schemes for Post-modifier Prepositional Phrases
Greg Whittemore | Kathleen Ferrara | Hans Brunner
28th Annual Meeting of the Association for Computational Linguistics