Skip to content

Details

Talk 1: Large scale topic assignment on multiple social networks. Nemanja Spasojevic, Adithya Rao, Lithium

Abstract: Understanding data from social networks to infer user's topical interests and expertise is a challenging problem with applications in various data-powered products. In this talk, we present a full production system used at Lithium Technologies (Klout), which mines topical interests and expertise from multiple social networks and assigns over 10,000 topics to hundreds of millions of users on a daily basis.
The system generates a diverse set of features derived from signals such as user generated posts and profiles, user reactions such as comments and retweets, user attributions such as lists, tags and endorsements, as well as signals based on social graph connections. We show that using cross-network information with a diverse features for a user leads to a more complete and accurate understanding of the user's topics, as compared to using any single network or any single source.
Speaker Bios:

  1. Nemanja Spasojevic:Nemanja is the Director of Data Science at Lithium Technologies, and leads all the data science and engineering efforts around building scalable systems for topical text mining, user scoring and content relevance. Prior to his current role, he was previously at Google for 6 years, working on projects such as Google Books.

  2. Adithya Rao: Adithya is the Lead Research Engineer in the Data Science team at Lithium technologies. Some of the projects he has worked on includes the Klout score, topic extraction, user targeting and social content discovery. Before his current role, he graduated with a Masters degree from Stanford University, specializing in Machine Learning.

Talk 2: Topic Detection at Massive Scale, Stefan Gower, Topic Scout

Abstract: Topic Scout is radically different. Its web site focuses on the large number of topics it has already successfully built - up to 10,000 of them - and shows ample evidence that Topic Scout can indeed create large topic systems. The data is accessible on the web site, and the data speaks loudly.

To the best of my knowledge, this is the only web site in the world that does this. Every other web site I have ever seen relating to topics and classification just sells tools.
Topic Scout is not about selling tools. It is about providing topics, either pre-built or allowing customers to build their own, and do this very effectively and quickly. In Topic Scout, any person who can create a good Google query has the skill to define topics, and create topics in minutes, not days or weeks. Highly automated data mining does the rest.
How does this relate to this proposed talk? The proposed talk is about creating 100,000+ topics and classifying the web at web scale. The Topic Scout web site provides strong evidence of this capability. It shows very large numbers of highly relevant topic lexicons that have been created using Topic Scout's relevance engine. The topic lexicons are very impressive. Highly relevant items related to that topic. Take a good look at them, and see for yourself..

Speaker Bio: Graduate of UC Berkeley. Went to work at both small and large Bay Area companies including HP, Sybase, Oracle and GE. Was principal architect at Bridgestream, a security and data mining company bought by Oracle. Sole inventor on several patents related to transactions. Areas of expertise include data mining, distributed systems, expert system, internet of things and security.

Related topics

You may also like