This collection consists of ~20M web queries collected from ~650k users over three months.
The data is sorted by anonymous user ID and sequentially arranged.
Here at Google Research we have been using word n-gram models for a variety of R&D projects, such as statistical machine translation, speech recognition, spelling correction, entity detection, information extraction, and others. While such models have usu
X. Wang, Z. Wang, X. Han, W. Jiang, R. Han, Z. Liu, J. Li, P. Li, Y. Lin, und J. Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Seite 1652--1671. Online, Association for Computational Linguistics, (November 2020)
J. McAuley, C. Targett, Q. Shi, und A. Van Den Hengel. Proceedings of the 38th international ACM SIGIR conference on research and development in information retrieval, Seite 43--52. (2015)