Showing posts with label search engine. Show all posts
Showing posts with label search engine. Show all posts

Monday, August 06, 2007

Can we discover buzz patterns from Blogs?

The huge Consumer Generated Media ( CGM ) or User Generated Content ( UGC ) available in terms of blogs, social networks , public Wikis and other Internet based content stores have always inspired me to find out the answer of the following question

Can they be used to determine a trend or buzz for a specific business entity ( ex. Products like Apple iPhone or any TV soap like The Prison Break) ?

A typical example is the following blog post on a restaurant named “Mainland China”
http://bangalore.metblogs.com/archives/2006/12/dining_out_mainland_china.phtml
and comments from lots of blogosphere users on that post,
Now the question is whether this small fragment of user generated content can act as a piece of "gyaan" which can be effectively searched and therefore seamlessly discovered on the Web?

When I searched google for "Mainland China Bangalore" or "Dining Out: Mainland China" the above mentioned page came up as one of the top results. Now this is something that the internet users do regularly , though among all the internet users only a few come up with good search phrases that ensure contextual results. The popular search engines cannot take into account the meta context information which can possibly be best defined by the intent with what the author has written a specific blog post. Possibly the title of the post can be taken as one important context for any keywords that we index from a blog.

All those important information posted by the bloggers and other CGM providers will effectively be lost if we don't bundle them with a specific context ( As. Tagging ) or make them searchable
( Indexing ).

Some of us have already seen the Blog Buzz implementations , these are typically user driven classification and categorization of blogs and other UGC for better presentation patterns in search result. But definitely we can improvise context driven classification of contents and search upto a point where we can provide the users with structured information like www.wikipedia.com provides.

I looked into existing blog search engines like

1)Technorati
2)Google blog search
3)Blog pulse and
4) "Nielsen BuzzMetrics(www.nielsenbuzzmetrics.com)"

All these sites index blogs and provide search interfaces on them , when some of them has gone one step ahead in providing structured trend information out of the blog content. But looks like they have a long way to go.

Basically my idea is to come up with a very basic implementation that does the following

1)Crawl a set of blogs belonging to a very specific domain ( Ex. Restaurant or Movies )
2)Index them in the order of business entities they primarily talk about.
3)Present the information in a review oriented format on a brief , user friendly UI.

Saturday, March 10, 2007

Ypodia! – The yahoo search experience in the wikipedia way…...

Vipin asked me to define the problem …which was indeed an intimidating demand of his. Yes, how about user tuned web search or sharing the search query strings (on specific keyword) on the web, well as the web user count is increasing each and every day there is a increased need to mirror the searches that has happened on a specific keyword and sharing the web searches made by experienced web user to the newbie …isn’t that idea comprehensive enough. My idea is to have a wiki, where instead of static articles...It will contain links to different information sources grouped by different contexts and different search APIs like (Google, yahoo, technorati search and yahoo answers .Flicker photos etc.)

Introduction:

Web search has become the hottest application on the wire in recent years, but the online encyclopedia tools like wikipedia still prove their success over powerful web based search engines like Yahoo or Google when a user wants to find some specific information on the web with search keywords, like “second world war” or “Microsoft corporation”. Just to give an illustration, if you search Google or yahoo for “Microsoft corporation” will return you links from web like the following ( top 10 results )

1) www.microsoft.com

2) en.wikipedia.org/wiki/Microsoft

3) support.microsoft.com

4) home.microsoft.com

5) msdn2.microsoft.com

6) office.microsoft.com

7) office.microsoft.com/en-us/frontpage/default.aspx

8) www.research.microsoft.com

9) msdn.microsoft.com/xml

10)members.microsoft.com/careers/default.mspx

Now, lets see what happens when we go to wikipedia for the same

It takes us to the page of http://en.wikipedia.org/wiki/Microsoft_Corporation article, which comprises the following

* 1 History

o 1.1 1975–1985: The founding of Microsoft

o 1.2 1985–1991: The rise and fall of OS/2

o 1.3 1992–1995: Domination of the corporate market

o 1.4 1995–1999: Foray into the Web and other ventures

o 1.5 2000–2005: Legal issues, XP, and .NET

o 1.6 2005–2007: The road to Vista

* 2 Product divisions

o 2.1 Microsoft Platform Products and Services Divisions

o 2.2 Microsoft Business Division

o 2.3 Microsoft Entertainment and Devices Division

* 3 Business culture

* 4 User culture

* 5 Corporate affairs

o 5.1 Corporate structure

o 5.2 Stock

o 5.3 Diversity

o 5.4 Logos and slogans

* 6 Criticism

o 6.1 Corporate

o 6.2 Technical

* 7 Microsoft.com

* 8 See also

* 9 References and footnotes

* 10 External links

Isn’t that great, a single source for all your want to knows … you can know better about Microsoft corporation from this page than anything available on the web, as this is a very structured and globally contributed article which pretty much cover the entire Microsoft story, so now on you will always go to this page instead of doing popular search engine search.

But lets not stop here, lets take a third approach..What happens if I would have searched web using the globally shared contexts on “Microsoft Corporation”.

Like the following

search string
(contexts)

a)Microsoft corporation history:


##### results #### ##### source #####

en.wikipedia.org/wiki/Microsoft yahoo search
www.radessays.com/viewpaper/13230/Microsoft_History.html yahoo search
www.microsoft.com/billgates/bio.asp yahoo search
www.answers.com/topic/microsoft yahoo search
www.thocp.net/companies/microsoft/microsoft_company.htm yahoo search
www.csl.mtu.edu/winter98/cs320/micro/history yahoo search

The history of Microsoft Corporation? yahoo Answers search
The History and Development of Microsoft Corporation? yahoo Answers search


b) Microsoft products


www.microsoft.com/products yahoo search
www.microsoftproducts.net yahoo search
technet.microsoft.com yahoo search
office.microsoft.com yahoo search
msdn.microsoft.com/canada/academic/products yahoo search
members.microsoft.com/careers/careerpath/
marketing/product.mspx yahoo search
tech.msn.com/guides/msproducts/default.aspx yahoo search

Why do Microsoft products suck? yahoo Answer search
Why do some people hate Microsoft, but still use
Microsoft products? yahoo Answer search
Microsoft products? yahoo Answer search
IBM and Microsoft product compatibility? yahoo Answer search
How can I find who a Microsoft product is registered to? yahoo Answer search
Stupid Microsoft Product Activation Bullshit problem? yahoo Answer search

Now think of a wiki, which contains a log if all the searches made on specific keywords with different contexts. A place where you will find a collection of links grouped by contexts applied on the search keyword ( those linked are typically fetched by searching across the web using popular search apis like Yahoo , Google and Yahoo answer , flicker …basically a detailed snapshot of all the search query combinations available for the specific keyword ).

Gory Details:

How it works? Possibly that will an interesting question to answer at this point…

This is a new type of wiki driven by search , as said that basically creating a document here is something like this , you define a keyword on which you also create search contexts like if the keyword is a person , as Shahrukh Khan , the possible contexts will be “filmography” , “awards and nominations” , “scandals” . Then once you define the contexts, you will be given the options to choose the search channels , like Yahoo search , yahoo answers , de.icio.us tag search , technorati blog search and when you select the search channels and generate the results …you will be given control to choose the results. Now once you choose the results …it will wiki the whole thing, keyword as the document root …under that various contexts defined and under each context the links that you have chosen. Now when I say that the whole thing is going to be built on the wikipedia infrastructure …then its valid that anyone can add more contexts and refine and redefine any of the existing contexts.