Showing posts with label RS. Show all posts
Showing posts with label RS. Show all posts

Monday, February 20, 2012

Book out: Educational Recommender Systems and Technologies: Practices and Challenges


It's always nice to see a piece of work coming alive that one has been part of. In this case, I was involved in the reviewing process of the articles for the book called "Educational Recommender Systems and Technologies: Practices and Challenges", a topic that is very dear to me.

The editors of the book will collect reviews to share them with all, they will be compiled here.
Recommender systems have shown to be successful in many domains where information overload exists. This success has motivated research on how to deploy recommender systems in educational scenarios to facilitate access to a wide spectrum of information. Tackling open issues in their deployment is gaining importance as lifelong learning becomes a necessity of the current knowledge-based society. Although Educational Recommender Systems (ERS) share the same key objectives as recommenders for e-commerce applications, there are some particularities that should be considered before directly applying existing solutions from those applications.

Educational Recommender Systems and Technologies: Practices and Challenges aims to provide a comprehensive review of state-of-the-art practices for ERS, as well as the challenges to achieve their actual deployment. Discussing such topics as the state-of-the-art of ERS, methodologies to develop ERS, and architectures to support the recommendation process, this book covers researchers interested in recommendation strategies for educational scenarios and in evaluating the impact of recommendations in learning, as well as academics and practitioners in the area of technology enhanced learning.

Wednesday, July 15, 2009

Google parsing microformat, e.g. ratings

In May Google announced that they will start parsing microformats (on a small scale first), similar stuff came out from Yahoo! last year, but even on a smaller scale.

This is pretty huge for the end-user generated ratings! I must say that I did not see it coming in this way, which makes it even more exiting :)

..Google is releasing support for parsing and display of microformat data in their search results. .. anyone who marks their pages up with the appropriate microformat data will be able to make their information understandable by Google. This technology would allow you to explicitly search, for example, for only printers that had an average customer review of 3 stars or higher.

Holy smokes! This is cool, can't wait to see when it will first pop up in my search :)

So, since a long time it's been problematic to get enough ratings on items, this is a known problem especially in the field of Recommender systems. They talk about "sparse data". An example, you want to make a recommendation on music, but the item x has 3 ratings, item y 2 ratings, etc. This is way too little to be used to create recommendations using the algorithms that are out there. Take another example, a camera shop, they let users rate their cameras, but they get very little reviews from users.

Now, however, there are other camera shops who are struggling with the same problem. Essentially, they all are selling the same camera brands, and they all have only a few ratings on it, and at the end, non of them can do much fun with this small amount of rather anecdotal information.

There has been talk about a unique identifier set by industry, for example, so that all camera sellers could use them and thus aggregate all the reviews and ratings together. Yep, you guessed it, there's maybe that one shop down the blog who does not want to use it. I think a couple of years back Yahoo! came up with a very compelling paper reiterating the idea and trying to muster up enough consensus among industry and other players. Not much happened - and then, here is Google and microformats... beautiful :)

Why I'm interested in this is that with the idea of federating learning resource metadata across repositories, we face the same problem. As a result of sharing metadata, the same resource might end up used in many different repositories, where users might be allowed to rate them. But that metadata on ratings or evaluations is VERY seldom shipped back to the mother board.

The same with tags and bookmarking (other other tools that allow users to create collections or playlists). That could be valuable information for the repository who first federated the resource metadata out. By collecting back the varied annotations from different repositories, they could gain interesting information, and eventually overpass the sparse data problem. Moreover, they would gain data about what works and in which context, which makes me think of "travel well" resources.

In action

Here is an example of a search for Palm's new phone that I'm contemplating on. I search for reviews only and the result list shows the ratings, but I cannot yet make a query saying "palm pre" ratings grater than 3. Nice in any case.




I've had a few ideas on this with some colleagues and I really look forward to seeing what Google comes up with that. And how are they going to solve the issue of different rating scales used, and multi-attribute ratings.

Vuorikari, R., Manouselis, N., & Duval, E. (2007). Metadata for social recommendations: storing, sharing and reusing evaluations of learning resources. In D. H. Goh & S. Foo (Eds.), Social Information Retrieval Systems: Emerging Technologies and Applications for Searching the Web Effectively (pp. 87-107). Hershey, PA: Idea Group Inc. Retrieved from http://elgg.ou.nl/rvu/files/20/144/SIR_vuorikari_manouselis_duval_web.pdf.


Manouselis, N., & Vuorikari, R. (2009). What if annotations were reusable: a preliminary discussion. In M. Spaniol (Ed.), Advances in Web-Based Learning - ICWL 2009, Lecture Notes in Computer Science (Vol. 5686, pp. 255–264). Berlin Heidelberg: Springer-Verlag.

Monday, May 04, 2009

Link structure and anchor text

I read that Brin & Page (1998) paper again. A few guidelines to keep in mind:
..our notion of "relevant" to only include the very best documents since there may be tens of thousands of slightly relevant documents. This very high precision is important even at the expense of recall (the total number of relevant documents the system is able to return).


Two features to produce high quality precision:
  • Link structure is used to create objective measure of its citation importance that corresponds well with people’s subjective idea of importance. Well, it's that simple..

  • Anchor text:
    ..anchors often provide more accurate descriptions of web pages than the pages themselves. Second, anchors may exist for documents which cannot be indexed by a text-based search engine, such as images, programs,..
The point about the anchor text is so interesting, I wonder how well does it apply to tags? I bet really well..

I also found this interesting: "it has location information for all hits and so it makes extensive use of proximity in search"

Differences Between the Web and Well Controlled Collections
  • extreme variation internal to the documents: documents differ internally in their language (both human and programming), vocabulary (email addresses, links, zip codes, phone numbers, product numbers), type or format (text, HTML, PDF, images, sounds), and may even be machine generated (log files or out putfrom a database).
  • external meta information as information that can be inferred about a document, but is not contained within it. Examples of external meta information include things like reputation of the source, update frequency, quality, popularity or usage, and citations. Not only are the possible sources of external meta information varied, but the things that are being measuredvary many orders of magnitude as well.


http://www.scribd.com/doc/3208417/The-Anatomy-of-a-LargeScale-Hypertextual-Web-Search-Engine

Monday, March 23, 2009

Sneak preview: SIRTEL'09

Workshop on Social Information Retrieval for Technology-Enhanced Learning (SIRTEL'08) in the International Conference on Web-based Learning (ICWL) 2009 in Aachen, Germany, August 21, 2009
http://www.hkws.org/events/icwl2009/workshops.html


IMPORTANT DATES

Contribution Submission: June 14, 2009
Results Notification: July 13, 2009
Camera Ready Submission: July 31, 2009
Workshop date: August 21, 2009

CALL FOR WORKSHOP CONTRIBUTIONS

We are delighted to welcome exciting new contributions for the 3rd SIRTEL workshop
  • Research papers
  • System Demos
  • Hands-On proposals
  • Abstracts for "Pecha Kucha"
RATIONALE

Learning and teaching resource are available on the Web - both in terms of digital learning content and people resources (e.g. other learners, experts, tutors). They can be used to facilitate teaching and learning tasks. The remaining challenge is to develop, deploy and evaluate Social information retrieval (SIR) methods, techniques and systems that provide learners and teachers with guidance in potentially overwhelming variety of choices.

The aim of the SIRTEL’09 workshop is to look onward beyond recent achievements to discuss specific topics, emerging research issues, new trends and endeavors in SIR for TEL. The workshop will bring together researchers and practitioners to present, and more importantly, to discuss the current status of research in SIR and TEL and its implications for science and teaching.

The proceedings from the last years:
- SIRTEL'07 http://ceur-ws.org/Vol-307
- SIRTEL'08 http://ceur-ws.org/Vol-382


TOPICS OF INTEREST (but not limited to):

  • Recommender systems and collaborative filtering in educational settings
  • Defining the scope, purpose and objects of social information retrieval in TEL
  • Novel ways of generating input for recommenders (explicit and implicit methods)
  • Ranking of search results to support individualised learning needs
  • Integrating SIR services in existing educational platforms
  • Folksonomies, tagging and other collaboration-based information retrieval systems
  • Social navigation processes and metaphors for searching information related to teaching and learning
  • Social networks and interactions in learning communities to facilitate information sharing and retrieval
  • Approaches to TEL metadata reflecting social ties and collaborative experiences in the field of education
  • Pedagogic decisions, recommender systems and how to contextualise recommender system to support learning processes.
  • Interoperability of SIR systems for TEL
  • Visualisation techniques in learning and teaching
  • Semantic annotation and tagging for social information retrieval purposes
  • Evaluating the performance of SIR systems in educational applications
  • Measuring the effectiveness of SIR systems in supporting learning and teaching
  • Evaluation the user satisfaction with SIR systems in supporting learning and teaching

WORKSHOP SUBMISSIONS

The workshop invites several types of contributions which allow a wide level of participation:
  • Research papers (upto 8 pages)
  • System Demos (upto 2 pages)
  • Hands-On proposals (1-pager)
  • Abstract for Pecha Kucha (1-pager)

The workshop proceedings will be published as CEUR Workshop Proceedings online at http://ftp.informatik.rwth-aachen.de/Publications/CEUR-WS/. Copy rights will be reserved.

Please use the same template as the one for the main conference with details at http://www.hkws.org/events/icwl2009/submission.html. Workshop paper length is not limited.

All questions and submissions should be sent to: sirtelworkshop@gmail.com
The website with the call will be up shortly too.

PROGRAM COMMITTEE:
• Alexander Felfernig, University of Klagenfurt, Germany
• Brandon Muramatsu, MIT, USA
• Frans van Assche, European Schoolnet (EUN), Belgium
• John Dron, Athabasca University, Canada
• Lloyd Rutledge, Open University of the Netherlands, NL
• Markus Strohmaier, Technical University of Graz, Austria
• Markus Weimer, Technische Universität Darmstadt, Germany
• Martin Wolpers, Fraonhofer-Institut, Germany
• Miguel-Angel Sicilia, University of Alcala, Spain
• Olga Santos, UNED, Spain
• Rick D. Hangartner, MyStrands, USA
• Rosta Farzan, University of Pittsburgh, USA
• Wolfgang Reinhardt, Universität Paderborn, Germany
• Xavier Ochoa, Escuela Superior Politécnica del Litoral, Ecuador
• Yiwei Cao, RWTH Aachen University, Germany
• Zinayida Petrushyna, RWTH Aachen University, Germany

Friday, September 12, 2008

Cross-boundary ranking of learning resources

Based on the idea of Interest Indicators, like social bookmarks and ratings, I've looked at the data so that we can make the cross-boundary resources better available on the MELT portal.
The aim is that we can, based on previous users' behaviour :

a) make separate "travel well" lists of resources that have a potential to cross-borders better,

b) use this information to rank resources better in the normal search result list,

c) allow users search for resources that have a good "travel well" value (e.g. give me resources in math that can cross-borders)

This is the data that I'm using (table below) and this is how I've defined cross-boundary (e.g. cross-country and language) learning resources before. Using that definition I have manually verified the number of cross-country resources. In the dataset, about 82% of resources were cross-country.



Now, we have a problem, though. On our MELT portal we do not have information about the country where the resource originates from. Dah!

This is a big blunder (in my opinion) in our Application Profile, we have not defined the country where the resource originates. We do define the provider, and the country information could be inferred from the provider, but it does not always work.

For example, one of our providers has frequently metadata about resources that do not originate from the same country!

I've experimented with the data using the information that we have on the portal, which is LOM about the resource including the language of the resource. As we also know the mother tongue of the registered users, this gives us a kick.

In the table below we can see the coverage of cross-boundary actions that we can get on resources without using any manual labour or verification of the country or language. As a base-line, with manual verification I found that 82% of the actions concerned cross-border rating or bookmarking of a resource.



The first row represents the cross-language resources (i.e. user's mother tongue is different from the resource language). Just using this information, we get about 65% of resources right, as opposed using manual checking (82%). I think it's pretty good, I'd settle for that! (although I have to look what kind of material was left out!)

The two other comparisons in the table are based only using information about users' previous behaviour. These would be:
  • rating > 2
  • bookmark
Only using information regarding bookmarked and rated resources results in a lousy coverage of around 20%. The problem is that 25% of resources bookmarked or rated are on more than one resource, the data still is very sparse.

Anyway, I want to use that information to "cross-boundary rank" the resources. As we do not know the country where the resource comes from, my work-around is based on countries where these users come from.

Here is a visualisation about resources that have been bookmarked or rated by users (see also ManyEyes link below). We can see the orange node in the middle, a learning resource called "Five Days in New York..". We see 3 edges leading out to Finland, Belgium and Hungary. This means at least one user from each of these countries has bookmarked the resource!

So, even if we do not know the origin of the resource, we know that it has users from 3 different countries. I can infer that it is a cross-boundary resource.

As most likely one of these 3 users come from the same country than the resource comes from, I will minus one country out of the total of countries: (number of countries -1)

My cross-boundary rank will be the following:
  • Count the number of ratings grater than 2 and/or bookmarks for a resource (actions). Give each action one point
  • Count the number of these users and give each user one point
  • Count the number of user countries of origin. Give each country one point and then minus 1
  • Compare the mother tongue of each of these users to the language of resource. If they differ, give one point/mismatch.
Then, count the following:
number of users + number of actions + number of cross-language x (number of countries -1)
Let's take the above resource "Five Days in New York.." as an example
  • Count the number of ratings grater than 2 (3) and/or bookmarks for a resource (5). Give each action one point. (8)
  • Count the number of these users and give each user one point (5)
  • Count the number of user countries of origin (Hungary, Finland, Belgium). Give each country one point and then minus 1 ( 3-1=2)
  • Count the number of user mother tongue (hu, nl, fi). Compare the mother tongue of each of these users to the language of resource (en). If they differ, give one point/mismatch (3).

  • number of users (5) + number of actions (8) + number of cross-language (3) x (number of countries -1) (2) = 32 Travel well value
This way you can count a value of "travel well" for each resource that users have previously interacted with on the portal. The value will always be an integer, which is important from the technical implementation point of view (in Lucine index it apparently needs to be an integer).

The down side is that we'll have a huge cold-start problem. As I said, our data is very sparse. To seed the system, I actually still manually check the new resources that users have interacted with and make a fake bookmark on them so that it looks like it has at least two users from 2 different countries. This way the resource gets a "travel well" value counted and appears on the "travel well" list and is better ranked, etc.

Of course, at the end I will evaluate how this treatment affects on users, do we, for example, see a big amount of bookmarks on these resources that I have been able to count a travel well value?

You can see a visualisation here. This is based on on user's country of origin.

Friday, October 19, 2007

what every PhD should know: dinner discussion with a google guy

Just barely hanging out there. Today was lots of serious fun and intellectual challenges at the RecSys 2007 Doctoral Consortium . Interestingly, all the participants came from lots of different backgrounds from computer science, information retrieval to me from education. I think the diversity of backgrounds and focuses of studies represent the growth of the Field of Recommenders, it's not only about the best algorithm anymore, but a plethora of questions around.

Anyway, being somewhat a newbie here (yeah, I do not know all the people or study areas here, very eye opening!), it makes me think that all the PhD students should be exposed to the question " If you were to have dinner next to a main researcher in Google/Yahoo/or any other big name, what would you want to talk about?".

Well, as it happens to be, I never thought of that before. Neither was I prepped for that by my study programme. Nevertheless, I just spent my dinner next to Krishna Bharat, you know, the guy who greated Google News, nothing less, nothing more. Probably tomorrow I'll have like ten things I want to ask from him with no chance to get his attention anymore.

Bottom line: it is not only about the 1 minute elevator pitch, but about the life and such in general.

Sunday, September 16, 2007

SIRTEL'07: la raison d'etre

"We use people to find content. We use content to find people."*

On Sept 18 our SIRTEL workshop takes place. It's gonna be "Serious Fun"! Let me just outline why:

SIRTEL'07: Raison d'etre

Recommender systems, as well as social navigation, have been around since the popularisation of WWW, that's some 15-20 years now. The idea is to help people choose the right stuff from a potentially overwhelming set of choices. To facilitate that users could be helped with information from other users, the choices made before (by themselves or similar users), the ratings or reviews other people had done, etc. (Rescnik et al., 1997)

The field of learning technologies has seen recommenders of some sort being discussed and prototyped since the late nineteens. In the review of the field in Manouselis et al (2008) we identified about 10 recommenders, and even more conceptual papers of them, but very little has matierialised so far.

Since the last few years recommenders have made a second arrival into the discussion topics of technology, or network, enhanced learning. Undoubtedly, this has been influenced by the arrival "Web 2.0" with all its ideas:

- Collaborative tagging, for example, has changed lots of ideas of how metadata should be produced and how static a metadata record should be: it's not anymore one metadata record produced by a librarian, but lots of annotational and attentional metadata by lots of users.

- Other annotations by users that express their subjective judgements have seen a huge growth too, we don't only talk about ratings or reviews in their traditional sense, but also tumbs-up or down, giving pokes to people or objects, etc.

- Social bookmarking, which allows users to create easy references to their own collections of digital resources (photos, books, links, music,..), has given a new dimension to the concept of social navigations. The link between resource-user(-tag) allows users to navigate other people's collections and thus find novel resources. Also, the same resource-user-tag link gives researchers an itch to use this information to group similar users for recommendation purposes, as well as to study the emerging networks.

- Expressing social ties between people has also brought new possibilities along. We are not only seeing networks of friends, but there are new possibilities where people can express different networks, ones for professional use, others for personal, recreational, etc purposes. Also, portability of these networks has become an issue discussed for better designs (social-network-portability group, PeopleWeb ,..).

- Something else is also happening behind the scenes. Clicksteam and user behaviour on the Web is not anymore a property of the commercial portal on which users are, but users are starting to take seriously how their "attention" is being used, who owns it, etc. Attentional metadata is a huge source of information that educationalists are also starting to take more seriously and thinking how it could be used for better serving learners and teachers (Contextual Attention Matadata, Attention Profiling Mark-up Language, Attention Trust,..). Attentional metadata can also become crucial when it comes to better understanding the intentions of a user, why are they, for example, looking for some information and for what task at hand!

- Finally, content for educational use, or rather its production, is also seeing a change. Users generate more and more of the content on the Web in general, a trend which is also seen in the e-learning. Of course, traditionally teachers have always produced lots of their own material, but now its re-use also has been facilitated (e.g. repositories/referatories). Also, the collaboration aspect is facilitated by the Web, it has become easier for people to work together on things (e.g. wikis, collaborative platforms,..). Additionally, learners produce plenty of material which also should be seen and used as educational content.

To sum-up: two main topics evolve around social context and social content. Social context is how we express the who, where and with whom, and social content are the objects or digital artefacts that are in the center of the communication, exchange and networks.

All the above has hopefully also changed how we will see the future of social information retrieval for technology enhanced learning. This workshop will all be about that! Serious Fun!

-------

N. Manouselis, R. Vuorikari, F. Van Assche, “Collaborative Filtering of Learning Objects for Online Communities: An Experimental Investigation”, accepted for publication in Computers in Human Behavior, Special Issue on ‘Advances of Knowledge Management and Semantic Web for Social Networks’, 2008.

P.Morville, 2004

Resnick P. & Varian H.R., “Recommender Systems”, Communications of the ACM, 40(3),1997

Thursday, May 24, 2007

Massively multiplayer object sharing by R.Sinha

I've followed some stuff from Rashmi Sinha, and I think every once in a while she comes up with good ideas. Like I liked the stuff early on that she did on the recommenders and the focus on user-centric design. Sometimes I just don't like her stuff, it sounds very popularistic and her references are, well, not very academic. But then again, maybe she does not need to be either..

Anyway, this slideshow has cool ingredients. I like the idea of object/artefacts in the center of the social networks, that's why I'm a big fan of social bookmarking, for example. I really don't care that much about connecting to people that I don't know (mySpace) or even using LinkedIn (what's the point, you get a list of people, but no substance..), but when I can connect through items and tags to people's stuff that I find interesting, I find it useful.

In the slideshow Sinha talks about models of 2nd generation networks (the 1st g was only about people):
  • Model 1: Watercooler conversations
    (around objects e.g., Flickr, Yahoo answers)
  • Model 2: Viral sharing (passing on interesting stuff, e.g., YouTube videos)
  • Model 3: Tag-based social sharing (linked by concepts. e.g., del.icio.us)
  • Model 4: Social news creation (rating news stories, e.g., digg, Newsvine)
Then, further on, she talks about Cognitive Diversity, which I also find really important. It's related to the continuum of wisdom of crowds vs. stupidity of mobs. What I got out of the slide 29:
  • Good answers need many perspectives, thus many perspectives are needed otherwise groups become too homogenous, which might have its dangers also (stupidity of mobs, see Digg for that ;). If all the new members are too similar and like-minded, they don't bring anything new to the group (that's why we want serendipity from recommenders!). Diversity reduces groupthink (think of Digg again and how fast not favourable stuff gets buried), groupthink is bad and only way to fight that is diversity.
Moreover, she also talks about the importance of social influence condition and about Watt's study.

Lastly, some design principles:
  • Make system personally useful: For end-user system should have strong personal use; Self-expression (e.g., Newsvine);Social status: Digg
  • Don’t count on altruism: System should thrive on people’s selfishness


Friday, May 11, 2007

1st Workshop on Social Information Retrieval for Technology-Enhanced Learning

I hereby proudly present the call for the first ever workshop on Social Information Retrieval for Technology-Enhanced Learning!

The complete call can be found from here: http://ariadne.cs.kuleuven.be/sirtel

A few words on the raison d'être of this workshop, what are the drivers for it?

Everyone in the field of e-learning has their ears full of talks of Communities of Practice (CoP) and networks of users, but not very often do we see how they actually are leveraged in practical terms. This workshop focuses on one part of the process, namely on retrieval of useful resources, either learning resources or human resources, for that matter. The tag line could be as P.Morville said "We use people to find content. We use content to find people."

Take that a step further and think of using digital traces to find people, and also leaving digital traces so that you can be found by other people. In this workshop we are interested in both; social navigation systems and recommenders for retrieving resources to enhance learning and teaching.

Social information retrieval (SIR) refers to a family of techniques that assist users in obtaining information to meet their information needs by harnessing the knowledge or experience of other users. Examples of SIR techniques include sharing of queries, collaborative filtering, social network analysis, social navigation, social bookmarking and the use of subjective relevance judgements such as tags, annotations, ratings and evaluations.

SIR methods, techniques and systems open an interesting new approach to facilitate and support learning and teaching. There are plenty a resource available on the Web, both in terms of digital learning content and people resources (e.g. other learners, experts, tutors) that can be used to facilitate teaching and learning tasks. The remaining challenge is to develop, deploy and evaluate systems that provide learners and teachers with guidance to help identify suitable learning resources from a potentially overwhelming variety of choices.

Several questions are being researched around the application of SIR methods in Technology-Enhanced Learning (TEL) settings. The aim of the SIRTEL'07 Workshop is to bring together researchers and practitioners who are working on topics related to the application of SIR methods, techniques and systems in educational settings, as well as to present the current status of research in this area to interested researchers and practitioners. It aims to serve as a discussion forum where researchers will present the results of their work, and also establish liaisons between different groups that are exploring related subjects. In addition, it aims to outline the rich potential of emerging SIR methods, techniques and systems in order to better build TEL systems and services.

Feel free to involve yourself, submit a contribution, blog about this, social bookmark the call (tag sirtel07) and talk about this to your pals!

See you in Crete in Spetember!

Monday, May 07, 2007

Workshop on Social Information Retrieval in Technology-Enhanced Learning (SIRTEL07)

Good news! The workshop proposal for EC-TEL 07 was accepted, so I will be co-organising my first workshop on social information retrieval techniques in support of learning and teaching later this September.

The tag line will be "We use people to find content, we use content to find people" by Morville. On the other hand, maybe it should be "We use digital traces to find people, and we leave digital traces to be found"..

Two main focuses: Recommender systems and Social navigation

The list of topics will be LONG, but I put it in here as an appetiser:

  • Defining the scope, purpose and objects of social information retrieval in TEL
  • Recommender systems and collaborative filtering in educational settings
  • Novel ways of generating input information for recommenders in the area of learning and teaching
  • Ranking of search results to support individualised learning needs
  • Folksonomies, tagging and other collaboration-based information retrieval systems
  • Social navigation processes and metaphors for searching information related to teaching and learning
  • Analysing social interactions in learning communities and social networks on the Web to facilitate information sharing and retrieval
  • Approaches to TEL metadata that reflect social ties and collaborative experiences in the field of education
  • Interoperability of SIR systems for TEL
  • Integrating SIR services in existing learning management systems
  • Visualisation techniques to support social navigation in learning and teaching
  • Semantic annotation and tagging for social information retrieval purposes
  • Evaluating the performance of SIR systems in educational applications
  • Measuring the effectiveness of SIR systems in supporting learning and teaching
  • Evaluation the user satisfaction with SIR systems in supporting learning and teaching

The idea is that as this is the first European workshop on the topic, we will try to scout out who are there to work on this topic and set the ground for better future collaboration . Of course we wish to run the workshop again, not as a pre-workshop , but really as a part of the main show.

Voila, more info to come shortly and the website for the call!

Saturday, May 05, 2007

The LibraryThing Recommender

I knew that LibraryThing.com had plans to work on a recommender for books, and seems like its out now. It's called LibrarySuggester; you can type a name of any book that you own or have read and the systems spills out suggestions in different categories:
  • People with this book also have...(v 1)
  • Special sauce recommendations!
  • Books with similar tags
  • Books with similar library subjects and classifications..
  • Amazon recommendations
  • People with this book also have...(v2)
LibraryThing Suggester analyses the more than thirteen million books and sixteen million tags LibraryThing members have added, and comes back with reading suggestions. Amazon suggestions come from Amazon.com, not LibraryThing.

Crowdsourcing

13 million books and 16 million tags, holy cow! That's some serious amount of data that people have free-willingly entered into the system! Just imagine trying to do the same before the day when the Web was crowdsourced. It would have taken an enormous amount of man-hours to enter people's likes and dislikes in books into a recommeder system as input to compute a list of recommendations, let alone the ratings, evaluations and discussions people have added too.

This is exactly the same way we want to go down with learning resources; first create a tool for teachers to create their favourite collections of learning resources and then use those to better serve them in terms of recommendations.


















Transparency

When I look at the recommendations from LibrarySuggester, what I like is that they are clearly classified in different classes of recommendations and on what those are based on. It is nice, as a user, to get the reasoning behind, e.g. ah, I was recommended this book because other "people with this book also have.." or I know that it is based on similar tags, etc.

This kind of practice of being transparent about the recommendations has also been argued about in previous research in the field, and it seems to be something that people appreciate, as opposed to a "black box" recommendations where the user has no idea on what the recommendations are based upon (Swearingen, 2001; Rafaeli 2005).

List of recommendations

Also, what I like is that LibrarySuggester offers a list of recommendations, as opposed to one or a few to choose from. However, in my list there were 74 recommendations all together, which I find way too much!

There are also some really evident ones, like books from the same author, which is not really a salient recommendation. McNee, et al. (2006) talk about a "similarity hole" that item-item collaborative filtering algorithm can trap users into by only giving similar recommendations. They argue that the old-skool accuracy metrics should be taken with a caution, as they only are designed to judge the accuracy of individual items and not the list of items. Thus, "the recommendation list should be judged for its usefulness as a complete entity, not just as a collection of individual items."

Moreover, within the same framework, which is called Human-Recommender-Interaction, these folks talk about three aspects that should be improved in recommendations. They are similarity (discussed above), recommendation serendipity, and the importance of user needs and expectations in a recommender.

Serendipity

Take the list of "Special sauce recommendations" for Dune by F.Herbert. On the list of 20 books you can find on the top 2 of his other books, and 2 by B.Herbert, his son. This sounds rather dull and not really anything surprising, you could find that easily from a bookstore too. By serendipity, the authors mean how unexpected the recommendation is for the user and how novel it is. For me personally this is a very important factor and why I like the idea of recommenders, as opposed to just content-based retrieval of resources.

I won't discuss the importance of user needs and expectations in a recommender, as in this case it is pretty clear. In some other cases, though, like for learning resources, this comes very important, as teachers do have different tasks at hand when they are looking for learning resources. This is something I've blogged before about and will keep exploring in my context of research.



McNee, S.M. , Riedl, J. , and Konstan, J.A. (2006) "Being Accurate is Not Enough: How Accuracy Metrics have hurt Recommender Systems". In the Extended Abstracts of the 2006 ACM Conference on Human Factors in Computing Systems (CHI 2006) [to appear], Montreal, Canada, April 2006

Rafaeli S., Dan-Gur Y., Barak M. (2005), “Social Recommender Systems: Recommendations in
Support of E-Learning”, Journal of Distance Education Technologies, 3(2), 29-45,
April - June 2005.

Swearingen K., Sinha R. (2001). , “Beyond algorithms: An HCI perspective on recommender
systems”, ACM SIGIR 2001 Workshop on Recommender Systems, 2001.

Thursday, March 01, 2007

Connections between working items in EUN, KUL and COSL

Visiting The Center for Open Sustainable Learning (COSL) in Utah State University last week was a very rewarding experience. People there are working on many interesting things regarding the support for the use of open content in education. I had already checked their website (http://cosl.usu.edu/) before going there, but of course it's nothing compared to sitting down with people, talking about my projects and their projects, and, little by little, creating understanding on complimentary approaches and ideas to collaborate (hopefully).

To get the big picture on where I stand between all the work done in European Schoonet (EUN), KULeuven (KUL) and COSL, I created a complicated looking image that hopefully helps me to put pieces of the puzzle together.

The image here on the left represent a rather generic lifecycle of a learning resource (this is the connected part of the image) and the greenish bubbles around it are working/research items that people in EUN, KULeuven and COSL work on. With a quick look it seems that all the institutes work in the same field - providing access to learning resources - however, each has its own focus and flavour to give.

Who and what?

KUL is focused on searchability of resources across different repositories (SQI, OAI-PHM) and emphasises the generation (ALOCOM, AMG) and collection of metadata (CAM), both to describe the resource (LOM) and its context of use (Contextualized Attention Metadata).

COSL works on making the content accessible (EduCommons CMS), but is also focuses on tools to support the use of content (MOCSL). Moreover, the relations of content and content, content and users, and users and users seem to have become more important (Didly, Oz). They are also working on a recommender system.

EUN has interest in making the content available (federated search, Minor) and using the social aspects for its retrieval, the main methods being social bookmarking and user annotations (MELT).

Let's take it step by step following the scenario

1. Repositories to make content available

This image depicts a rather typical scenario of learning resources which end up in a repository; first the resource is "created", then"described" (i.e. metadata added) and "approved" (e.g. QA) to be part of a repository's collection where it is "published".

EduCommons, a plone-based CMS created by COSL folks has been developed for this reason.

eduCommons is an OpenCourseWare management system designed specifically to support OpenCourseWare projects like USU OCW. eduCommons will help you develop and manage an open access collection of course materials.
Whereas in EUN MINOR was developed. It's an open source learning resources repository that allows users to manage learning resources and their metadata. Minor comes with a built-in connector to the EUN's Federation of Internet Resources for Education, which greatly facilitates the work of connecting to other repositories and querying their resources.

2. Make the content available for search

Both EUN and KUL has been active in making the resources and/or metadata that reside in distributed collections and repositories accessible. Both have greatly contributed to Simple Query Interface (SQI). EUN and KUL, for example, demonstrated a real-time interoperability between Learning Object Repositories last summer.

Apart from SQI, folks in KUL are also looking into harvesting metadata, which seems like a suitable and complimentary solution for SQI.

3. Accessing the resources

Once the resources are made available in a way or another, there are the users who "discover" them using various methods. This is an area where quite a lot of work in done in all 3 places.

In EUN and KUL, within the MELT-project, the focus is on creating Social information retrieval (SIR) techniques for flexible access to large-scale collections of content. Here we are interested in social bookmarking and tagging, as well as on other annotations like comments, ratings, etc. This work can also lead to an implementation of a recommender system for learning resources within the federation of repositories.

COSL has also been busy thinking of how to use people, and relations that people and content have, to create serendipitous ways to discover resources. Ozmozr (Oz) is, or will be, an identity and content aggregator that leverages the power of social collaborative filtering through tagging. Currently it looks somewhat messy, but try to see for the grand idea that lies behind.

There is also Srumdidilyumptio.us, just didly among friends, that allows users to express relationships, let them between people and people, people and content or content and content, i.e. any two resources accessible by URLs. Didly might not look that cool, but one should think of it rather as an API to make use of that information in another context, I was told.

These two tools/concepts pretty well demonstrate the thinking that COSL folks take forward. One could say that they aggregate s!#t out of everything, to illustrate the Web 2.o thinking. Moreover, COSL will be working on topping up these two "relations aggregators" with a recommender system. The difference to the work in EUN is that they are planning to recommend anything that people have input in Oz, whereas EUN is only talking about recommending learning resources.

More in the same area, novel ways to access resources, one of the PhD students in KUL is working on creating a better ranking system for resources, namely the LearnRank. This will use the Contextualised Attention Metadata framwork (CAM). Another one is busy with information visualization techniques for the same purpose. We are also interested in experimenting on the combo of information visualization and social bookmarking in the future.

4. Support the decision making

Next step in discovering resources is to make the decision upon the piece of content. Usually users are exposed to a list of search results (sortable sometimes, ranked or not) or a list of recommendations. There are different ways the system can support users at this point.

EUN and COSL are both interested in user relevance judgments (AnnoRate; use of ratings and evaluations in EUN) and other annotations such as pedagogical comments, to help users to choose the content. EUN has also implemented personal collections of learning resources, so users could see in how many other collections the resource has been added in, to help them to judge its relevance.

After choosing the resource, users in EUN can add it to their favourites collection (bookmark) or choose to "obtain" the resource, its metadata or uri from the repository. Ozmozr, a COSL tool, allows users to bookmark items too, and to share them with other users. Additionally, at the same time, users can also save these bookmarks to their delicous account, for example.

It's good to make here a remark that CAM framework can be used to collect all kinds of user attention when interacting with a repository, for example, to save the search history, what items have been looked at or obtained, etc. This can be fed in again to help the discovery of resources, like is the case of LearnRank.

5. Re-using and re-purposing the content, content collaboration


Once the items have been obtained, there are the users, teachers or learners, who modify, "re-purpose", sequence or rip the resource to make it suitable for their needs.

In this sub-area ,some complimentary approaches have been taken too. Within an EUN-lead project a plone-based collaborative platform called LeMill has been developed by UIAH. LeMill is a web community for finding, authoring and sharing open and free learning resources.

In COSL some tools are under construction to support the use and re-use. Send2Wiki plug-in allows a user to send a webpage to a wiki and modify it there. It will automatically check on the CC licence and bring that info along. User can then modify the page in the wiki as they wish, for example to make a translation of it (hope to pilot this in EUN), or if it's from WikiPedia, modify the content to make it more suitable for K-12, for example (now some researchers in COSL say that Wikipedia is rather graduate level reading, thus the need for different versions).

Make a Path will help users in putting pieces of resource together for a coherent learning path. This tool is like creating your own music playlists. Once the users are using it (I hope EUN will pilot this, for example), some interesting patterns will most like start appearing. One possible way to take this forward to help sequence learning resources could be case based reasoning, like Claudio did for generating music play lists. Additionally, The INSTRUCTIONAL ARCHITECT allows teachers to find, use and share learning resources from the National Science Digital Library (NSDL) and the Web in order to create engaging and interactive educational web pages.

Moreover, there is research interest to better understand the collaboration aspects and communication patterns around content, how people work on it, share it, ask questions, help others, etc. In LeMill the collaboration aspect is important, one of the developers is currently conducting his PhD research (UIAH) on collaborative authoring of learning resources. Whereas in COSL, there is interest in looking into Yahoo! Answers, as well as there was some early work around supporting the users of MIT OpenCourseware.

I think KUL ALOCOM framework and AGM could be used here to make sure that a) the smallest possible pieces of content are saved in the repository, too, b) the correct metadata is sucked from the context. Also, CAM could be used to better track users interaction modes and patterns, etc.


6. Actual use of content and its evaluation


After ripping and mashing the content, it it usually "integrated" (e.g. using CMS, etc) to teaching and learning practices (pedagogical methods). "Use" of the learning resource takes place in its own context. Here, the CAM framework becomes useful to know more about how the resources have been used, in what context (e.g. entered by a university prof to a course, all this metadata can be sucked with the help of AGM and CAM).

Following the use, the users are given possibilities to evaluate or rate the content. The challenge here for a federation of repositories is that different repositories use different ways to rate and evaluate the content (different scale, different criteria (ped, technical, suitability in classroom,..), thus interoperability becomes cumbersome and there might be something lost in semantics (EUN). A rating service on a top of repositories like AnnoRate overcomes such hurdles.

To shortly summarise:


Everyone is interested in adding metadata, any kinds of it (LOM, attention, annotations, relations (bookmarks, sequence, ...), to the system












and using it for better retrieval of resources, and the new thing, to connect people.
















"We use people to find content. We use content to find people. Information seeking behavior and social network analysis go hand in hand." Peter Morville (2004)
The really interesting part of a folksonomy is not the content item being described, and not the tags that describe the item, but the person doing the tagging. by Greenchameleon

Wednesday, January 17, 2007

Notes on Combining Social- and Information-based Approaches for Personalised Recommendation on Sequencing Learning Activities

This paper, Combining Social- and Information-based Approaches for Personalised Recommendation on Sequencing Learning Activities, cames from OUNL and is related to a bigger schema of works that those guys are carrying out on competencies. Thus, the context is very related to lifelong learning, namely to higher ed and vocational training.

The aim of the paper is to describe a domain model for "way finding". By the term "way finding" is meant "selecting and sequencing learning activities". The raison d'etre is:
Learners' problems in way finding will decrease the efficiency of education provision (the ration of output to input) and increase the cost. The local context for this paper is Dutch Open University student who lacks adequate information on study possibilities at an early stage of study, and problems in getting a good overview of the number and best sequence to study modules.

The paper describes a personalised recommender system (PRS) model that combines social-based (i.e. completion data from other learners) and information-based (i.e. metadata from learner profiles and learning activities) data to recommend the best next learning activity. The system is currently under development, a limited implementation is running using learner profile metadata.

Note about learning activities; OUNL has been very active in developing IMS Learning Design. They (Tattersall et al. (in press)) have previously proposed IMS-LD as a candidate to model learning paths. Moreover, they argues that its selection and sequencing constructs appear suitable for learning activities (units-of-learning) as well as for higher levels of granularity (e.g. competence development programmes). Interesting. At one point of time one could look how IMS-LD information could be generated in attention metadata (CAM).

So, the idea of PRS approach is a hybrid recommender that uses
  • a) information from other learners and their completion of tasks (completion is understood like rating) in a collaborative filtering manner (in text called social-based approach), and
  • b) information from students profile and c) metadata about the resource in the spirit of a content-based system (in text called information based approach).

Authors also argue that it is not enough to find the most efficient learning paths (like the shortest route in GPS), but to explore which paths are most attractive or suitable (like routes suited for biking), thus personalisation needed (individualised needs, interest, preferences or circumstances).

Other key concepts are:
  • learner's start position in a given domain (prior learning history)
  • aimed competence profile for that domain
  • learning path towards that competence
To develop a PRS the authors identify the following pieces, that the paper defines:
  • uniform and meaningful description of formal and informal learning paths
  • learning activities that are addressable and meaningfully described
  • uniform learner profiles that define needs and preferences
  • uniform competence description that defines proficiency levels
  • a learning path processing engine
  • an engine recording completion of activities
  • information matching techniques to enable personalised recommendation

Related work

I will later post on my blog some excerpts from my own literature review in the field of learning to show other recommender ideas based on the same hybrid approach, as in this paper they mention that this approach has hardly been applied in learning. There was only a reference to Herlocker et al. (2004).

In the related work section it is mentioned that education field imposes some specific demands for recommender. Main differences sited between recommenders for books are the degree of voluntariness (learning is many times required to obtain some goals) and the possibility to establish an explicit completion (as most learning activities are to be assessed for successful completion). Hmm..I do agree with the statement, but had come up with different reasons myself. Goes to show, I guess, how the initial requirements for a recommender system differ from what I'm working on.

An interesting outcome is cited from Janssen et al. (in press): learners were offered a recommendation "most successful learner continued with Y after having completed X". I call this an "Amazon-like" recommendation (other people interested in this book also bought x, y, z) based on clustering behaviour. There were no personal characteristics taken into account in this study. They found out that this type of recommendation enhanced effectiveness in completion of the set of learning activities, but did not increase efficiency, the time it took to complete them.

Authors also acknowledge the problem of insufficient data that can be derived from existing log files, the same that my colleagues are working on with the view on capturing attention metadata (CAM).


Discussion

Authors state:
"From a self-organising point of view it would be ideal if way finding would emerge as a result of (in)direct interactions between members of the learning network, without being dependent of formalised descriptions in domain and user models."

I so agree: instead of investing time in describing all the information regarding the learner, his/hers existing and required competences; the resource; and the curriculum with goals and skills required, would be more interesting to tap onto existing knowledge from the masses and their previous experiences, the decisions they took to find the next suitable step, etc.

The authors also discuss the complimentary approach of controlled vocabularies or ontologies combined with annotations such as social tagging and rating, just in the same direction as we are doing in MELT (we talk about adding metadata a priori and a posterior) and what I'm interested in looking into.

Related to my work

The difference in what I'm looking into now and what this paper describes is, first of all, the context. I'm interested in a repository that is used mostly by K-12 teachers and learners (sometimes). The repository is not linked to formal learning requirements related to a curriculum, because it is used on the European level, where there are many curricula depending on a country or a local policy. However, each teacher who comes to that repository has his/hers own information seeking tasks, that I've talked about previously. Sometimes those tasks are related to covering a piece of a local curriculum, whereas some other times it is to find a piece of resource to support some generic learning goal, or find inspirational material, or something else.

Secondly, in my context of work sequencing learning resources is not the goal, rather just finding resources that fit to the search criteria and the task at hand. So I'm not so into this sequencing, however, I like the idea of playlists and using this type of expert knowledge of putting items together for learning purposes (the use of Case-based Reasoning like Claudio explore here).

Imagine if teachers could generate playlists of LOs as easily as I do playlists in iTunes (I'm NOT talking about automatically generating them, but hands-on deejaying). Then, those lists could be used as rules for generating new ones. In that case, we would not need LOM to know which item in a repository is described as "introductory item" or "motivational item" to start the lesson, or which one is good for "explaining a rule", but we could detect that information form playlists generated by teachers who knows through her domain knowledge that after this piece I put x,y and x. Let's see.

To check out from the paper:
- Koper (2005)
- Janssen et al. (in press) about the test
- Sicilia (2005) about ontologies to express competencies
- Van Setten, 2005 social-based approaches
- McCalla (2004), pragmatics-based paradigm of tagging learning activities with learner information

Tuesday, August 29, 2006

Notes and comments on: Accurate is not always good: How Accuracy Metrics have hurt Recommender systems

S.M. McNee, J. Riedl, and J.A. Konstan. "Being Accurate is Not Enough: How Accuracy Metrics have hurt Recommender Systems". In the Extended Abstracts of the 2006 ACM Conference on Human Factors in Computing Systems (CHI 2006) [to appear], Montreal, Canada, April 2006

The paper starts by informally arguing that "the recommender community should move beyond the conventional accuracy metrics and their associated experiment methodologies. We propose new user-centric directions for evaluating recommender systems".

The paper states that the current accuracy metrics, such as MAE (Herlocker 1999), measure recommender algorithm performance by comparing the algorithm's prediction against a user's rating of an item. They continue saying that this means, in essence, that a recommender that recommends places to a user where she has already visited would be rewarded rather than a recommendation on new places that might be of interest. Clearly, if that is the case, there is something rotten.

The paper proposes three aspects; similarity, recommendation serendipity and the importance of user needs and expectations in a recommender, and suggests how they could be improved.

A) Similarity

- the item-item collaborative filtering algorithm can trap users in a "similarity hole" only giving similar recommendations. This becomes more problematic when there is less data, for example, for a new user in a system.

The authors go on to discuss about the accuracy metrics that don't recognise this problem, because they are designed to judge the accuracy of individual items and not the list of items. However, the authors argue, "the recommendation list should be judged for its usefulness as a complete entity, not just as a collection of individual items." There was evidence in a user testing that the lists that had performed badly on conventional accuracy measures were the ones preferred by users. These lists had used the Intra-List Similarity Metrics and the process of Topic Diversification for recommendation lists (Ziegler 2005).

Authors go on saying that depending on the user's intentions, the makeup of items appearing on the list affected the user's satisfaction with the recommender. Here, in my opinion, it becomes important to remember the user intentions as provided by Swearingen & Sinha (2001)
  1. Reminder recommendations, mostly from within genre (“I was planning to read this anyway, it’s my typical kind of item”)
  2. More like this” recommendations, from within genre, similar to a particular item (“I am in the mood for a movie similar to GoodFellas”)
  3. New items, within a particular genre, just released, that they / their friends do not know about
  4. “Broaden my horizon” recommendations (might be from other genres)
Comment: So, taking all this into account, what can we think of the lists for learning resources?

B) Serendipity

This is how unexpected the recommendation is for the user and how novel it is. This is hard to measure. The authors approach the issue by its opposite: the ratability of received recommendations, and this, they say, is easy to measure by using the "leave-n-out" approach. However, the assumption that users are interested in the highest ratable items is not always true for recommenders. They give an example of recommending Beatle's White Album to users of a music recommender as a bad idea, as it almost adds no value.

The same example, I remember, was somewhere else on recommending to people buy bananas when they go shopping, however, apparently people almost always buy bananas anyway, thus no commercial value there..This could, though, have some value, when building people's trust on a recommender.

Moreover, the authors point out that different algorithms give different recommendations, and that people preferred one over another depending on their current task (think again about Swearingen/Sinha). To conclude on serendipity, the authors say that other metrics could be needed to judge a variety of algorithm aspects - no direction given on this one, though.

c) User experiences and expectations

New users have different needs from experienced users. Rashid (2001) has shown that the choice of algorithm for a new user greatly affects the experience (really?!) and also, apparently, the native language is greatly preferred (Torres 2004), wonder what kind of language groups were in question there..

- Moving forward

Authors don't suggest that the old-school metrics should be thrown away, but not only be used alone, we need to think of the users who want meaningful recommendations (!!).

Firstly, it is recommended that instead of looking at each item on the list of recommendations, one should pay more attention on the integrity of the list, using metrics like Intra-List Similarity metrics, and more of such kinds.

We should test more what kind of search algorithms users like and given them those.

Users have a purpose for expecting a recommendation, so we would need to know better what actually are the user needs when they come to see a recommendation (Zaslow 2002).

Well, well, if this is where we are at with recommender usability studies, it is not much. However, it is great that important people such as Grouplens researchers tell us this, so maybe it makes the general audience more susceptible for new things to come.

References:

Swearingen & Sinha (2001)

Torres, R., McNee, S.M., Abel, M., Konstan, J.A., and Riedl, J. Enhancing digital libraries with
TechLens+. In Proc. of ACM/IEEE JCDL 2004, ACM
Press (2004) 228-236.

Ziegler, C.N., McNee, S.M., Konstan, J.A., and Lausen, G., Improving Recommendation Lists through Topic Diversification. In Proc. of WWW 2005, ACM Press (2005), 22-32.


Zaslow, J. If TiVo Thinks You Are Gay, Here's How To Set It Straight --- Amazon.com Knows You, Too, Based on What You Buy; Why All the Cartoons? The Wall Street Journal, sect. A, p. 1, November 26, 2002.

Wednesday, August 23, 2006

Why Google is not a content-based recommender

Yesterday in the HMDB-bookclub, that we run in my unit, we read and discussed a paper on the recommender systems (Adomavicius & Tuzhilin, 2005, Toward the Next Generation of Recommender Systems). As this is my topic of research I was very eager to hear how my study-buddies perceived the issue and what did they have to say.

The discussion lingered into understanding the two main trends to produce recommendations: the content-based (CB) and collaborative recommendation. There were questions and attempts to answer them which left me unsatisfied after the session. Mainly, we left with the impression that Google, or any information retrieval system, would be, at the end of it, just a content-based recommender. I was somewhat troubled with this though and set my self for the quest to understand better what is there to discover.

Let’s go first by definition: Konstan et al. (2005) say:
Unlike ordinary keyword search systems, recommenders attempt to find items that match user's tastes and the user’s sense of quality, as well as syntactic matches on topic or keyword. For example, a music recommender will use an individual’s prior taste in music to identify additional songs or albums that may be of interest.

When Adomavicius et al (2005) talk about CB approach, they state that it has its roots in information retrieval and filtering research, but
the improvements over the traditional information retrieval approaches comes from the use of user profiles that contain information about user’s tastes, preferences, and needs. The profiling information can be elicited from users explicitly, e.g., though questionnaires, or implicitly- learned from their transactional behavior over time.

In the regular Google search there is no account, whereas to produce both content-based (CB) and collaborative filtering (CF) recommendations we need an account that we can assign to the user. An individual user profile is build based upon this.

In the CB recommendation a user is recommended items similar to the ones preferred in the past. This means that we need a search history, i.e. a user profile, where we can identify what the user has preferred in the past.

Thus, to generate a rather complete user profile that can find similarities between items (not people!) things like a history of viewed paged, bookmarked pages, the purchase history, “wish list”, and things like heurestic text analysis, etc. are important (implicit rating/input). Conventionally, especially with the first generation of recommenders the explicit ratings were the top notch:

...mid-1990s when researchers started focusing on recommendation problems that explicitly rely on the ratings structure. In its most common formulation, the recommendation problem is reduced to the problem of estimating ratings for the items that have not been seen by a user...Once we can estimate ratings for the yet unrated items, we can recommend to the user the items(s) with the highest estimated ratings(s) (Adomavicius, 2005)

Additionally, many times the CB systems would use additional information such as demographic, specific interests, location, etc that is part of the user’s self-manifested profile for the input. Maybe in the future this type of information could be extracted from some other sources, such as blog-postings, as were suggested during the session.

So, to get closer to the answer to the question, whether Google is just a content-based recommender, we can say that if used anonymously, it is not, although probably many of the techniques are the same. However, if we think of Google Personalized Search (beta) it for sure gets to be one.

The second somewhat baffling issues was the name of collaborative filtering, as it turns out, there is no collaboration between the users to produce any recommendations. In the CF recommendation the user is recommended items that people with similar tastes and preferences liked in the past. This means that we need a history for this person, too, in order to find out similarities within tastes and past experiences.

The strength of the CF approach at this stage is that even if you personally haven’t seen a link, product or what ever object we are talking about, or indicated the system what you liked about it, there most likely is someone in your nearest neighbourhood who has indicated that. Thus, in CF the values used to compute the recommendations are inferred based on similarities on the profiles, and you don’t need to have necessarily done it yourself. So, here lies the one main divider between CB and CF as for the input for the recommender: CB only uses YOUR history, whereas CF uses other users’ search history to better understand, or guesstimate, your history.

Well, this is stuff explained in short, more and better arguments are found in the papers and in my links at: http://www.furl.net/members/vuorikari/recommendation


Adomavicius & Tuzhilin, 2005, Toward the Next Generation of Recommender Systems

J.A. Konstan, N. Kapoor, S.M. McNee, and J.T. Butler. "TechLens: Exploring the Use of Recommenders to Support Users of Digital Libraries". A Project Briefing at the Coalition for Networked Information Fall 2005 Task Force Meeting, Phoenix, AZ, December 2005.
http://www.grouplens.org/papers/pdf/CNI-TechLens-Final.pdf