Monday, May 14, 2007

Notes on Everything is Miscellaneous interview

A pretty good interview on S.Weinberger's new book Everything is Miscellaneous with a Yahoo! guy Bradley Horowitz. It's actually way too long, he's a verbose guy, for sure, he's elevator pitch in the beginning would need an elevator ride in a skyscraper and back, and that would still not be enough! Some interesting bits to dive in or for fast-forwarding.
  • An interesting discourse over a unit of knowledge; he talks about group knowledge and how, for example, Wikipedia discussion pages or some mailing lists themselves are a unit of knowledge build through discussion, agreement and disagreement, and sometime arguments, too. Non of those individuals would not be able to construct that alone "the knowledge is not in anyone's head, it's literally in the discussion in the mailing list (28mins)
  • It was funny when they talked about Justin TV, how this guy carries a camera 24/7 and records his life. The interviewer comes up with something like what's the point, "you don't get a second life with which to review the first one (about at 31 min). This part is also related to gathering metadata about everything, in MIT they record individual's heartbeat throughout the day, so that you can later check when the heart rate was high, remember the moment and relive it! Dudes, get out!

  • Towards the end the whole thing gets more interesting. In about 45 min they talk about social filtering and Mr. Weinberger is very pro, he argues that it is hard for us to know what we are interested in. He goes "...the serendipity thing: we don't know what we are interested in. There are things that we can predict we are interested in, but largely not. The world is way more interesting than our interests, which is why social filtering is so important." I like that :)

  • Soon after they talk about the definition of discussion (48min) which makes me think of a lunch discussion with Teemu and his group in Medialab on how computer science has banalised terms like dialogue (a pop-up "yes" or "no"), interactive, etc..

  • At the very end (54min) they talk about folksonomies, and he says about how silly it would be to replace a taxonomy with a folksonomy. His argument was that "you don't want just one folksonomy", but many of them so that you can cluster and gather things so that it is relevant to what we care and interested in. That's good too!

Friday, May 11, 2007

1st Workshop on Social Information Retrieval for Technology-Enhanced Learning

I hereby proudly present the call for the first ever workshop on Social Information Retrieval for Technology-Enhanced Learning!

The complete call can be found from here: http://ariadne.cs.kuleuven.be/sirtel

A few words on the raison d'ĂȘtre of this workshop, what are the drivers for it?

Everyone in the field of e-learning has their ears full of talks of Communities of Practice (CoP) and networks of users, but not very often do we see how they actually are leveraged in practical terms. This workshop focuses on one part of the process, namely on retrieval of useful resources, either learning resources or human resources, for that matter. The tag line could be as P.Morville said "We use people to find content. We use content to find people."

Take that a step further and think of using digital traces to find people, and also leaving digital traces so that you can be found by other people. In this workshop we are interested in both; social navigation systems and recommenders for retrieving resources to enhance learning and teaching.

Social information retrieval (SIR) refers to a family of techniques that assist users in obtaining information to meet their information needs by harnessing the knowledge or experience of other users. Examples of SIR techniques include sharing of queries, collaborative filtering, social network analysis, social navigation, social bookmarking and the use of subjective relevance judgements such as tags, annotations, ratings and evaluations.

SIR methods, techniques and systems open an interesting new approach to facilitate and support learning and teaching. There are plenty a resource available on the Web, both in terms of digital learning content and people resources (e.g. other learners, experts, tutors) that can be used to facilitate teaching and learning tasks. The remaining challenge is to develop, deploy and evaluate systems that provide learners and teachers with guidance to help identify suitable learning resources from a potentially overwhelming variety of choices.

Several questions are being researched around the application of SIR methods in Technology-Enhanced Learning (TEL) settings. The aim of the SIRTEL'07 Workshop is to bring together researchers and practitioners who are working on topics related to the application of SIR methods, techniques and systems in educational settings, as well as to present the current status of research in this area to interested researchers and practitioners. It aims to serve as a discussion forum where researchers will present the results of their work, and also establish liaisons between different groups that are exploring related subjects. In addition, it aims to outline the rich potential of emerging SIR methods, techniques and systems in order to better build TEL systems and services.

Feel free to involve yourself, submit a contribution, blog about this, social bookmark the call (tag sirtel07) and talk about this to your pals!

See you in Crete in Spetember!

Monday, May 07, 2007

Workshop on Social Information Retrieval in Technology-Enhanced Learning (SIRTEL07)

Good news! The workshop proposal for EC-TEL 07 was accepted, so I will be co-organising my first workshop on social information retrieval techniques in support of learning and teaching later this September.

The tag line will be "We use people to find content, we use content to find people" by Morville. On the other hand, maybe it should be "We use digital traces to find people, and we leave digital traces to be found"..

Two main focuses: Recommender systems and Social navigation

The list of topics will be LONG, but I put it in here as an appetiser:

  • Defining the scope, purpose and objects of social information retrieval in TEL
  • Recommender systems and collaborative filtering in educational settings
  • Novel ways of generating input information for recommenders in the area of learning and teaching
  • Ranking of search results to support individualised learning needs
  • Folksonomies, tagging and other collaboration-based information retrieval systems
  • Social navigation processes and metaphors for searching information related to teaching and learning
  • Analysing social interactions in learning communities and social networks on the Web to facilitate information sharing and retrieval
  • Approaches to TEL metadata that reflect social ties and collaborative experiences in the field of education
  • Interoperability of SIR systems for TEL
  • Integrating SIR services in existing learning management systems
  • Visualisation techniques to support social navigation in learning and teaching
  • Semantic annotation and tagging for social information retrieval purposes
  • Evaluating the performance of SIR systems in educational applications
  • Measuring the effectiveness of SIR systems in supporting learning and teaching
  • Evaluation the user satisfaction with SIR systems in supporting learning and teaching

The idea is that as this is the first European workshop on the topic, we will try to scout out who are there to work on this topic and set the ground for better future collaboration . Of course we wish to run the workshop again, not as a pre-workshop , but really as a part of the main show.

Voila, more info to come shortly and the website for the call!

Saturday, May 05, 2007

The LibraryThing Recommender

I knew that LibraryThing.com had plans to work on a recommender for books, and seems like its out now. It's called LibrarySuggester; you can type a name of any book that you own or have read and the systems spills out suggestions in different categories:
  • People with this book also have...(v 1)
  • Special sauce recommendations!
  • Books with similar tags
  • Books with similar library subjects and classifications..
  • Amazon recommendations
  • People with this book also have...(v2)
LibraryThing Suggester analyses the more than thirteen million books and sixteen million tags LibraryThing members have added, and comes back with reading suggestions. Amazon suggestions come from Amazon.com, not LibraryThing.

Crowdsourcing

13 million books and 16 million tags, holy cow! That's some serious amount of data that people have free-willingly entered into the system! Just imagine trying to do the same before the day when the Web was crowdsourced. It would have taken an enormous amount of man-hours to enter people's likes and dislikes in books into a recommeder system as input to compute a list of recommendations, let alone the ratings, evaluations and discussions people have added too.

This is exactly the same way we want to go down with learning resources; first create a tool for teachers to create their favourite collections of learning resources and then use those to better serve them in terms of recommendations.


















Transparency

When I look at the recommendations from LibrarySuggester, what I like is that they are clearly classified in different classes of recommendations and on what those are based on. It is nice, as a user, to get the reasoning behind, e.g. ah, I was recommended this book because other "people with this book also have.." or I know that it is based on similar tags, etc.

This kind of practice of being transparent about the recommendations has also been argued about in previous research in the field, and it seems to be something that people appreciate, as opposed to a "black box" recommendations where the user has no idea on what the recommendations are based upon (Swearingen, 2001; Rafaeli 2005).

List of recommendations

Also, what I like is that LibrarySuggester offers a list of recommendations, as opposed to one or a few to choose from. However, in my list there were 74 recommendations all together, which I find way too much!

There are also some really evident ones, like books from the same author, which is not really a salient recommendation. McNee, et al. (2006) talk about a "similarity hole" that item-item collaborative filtering algorithm can trap users into by only giving similar recommendations. They argue that the old-skool accuracy metrics should be taken with a caution, as they only are designed to judge the accuracy of individual items and not the list of items. Thus, "the recommendation list should be judged for its usefulness as a complete entity, not just as a collection of individual items."

Moreover, within the same framework, which is called Human-Recommender-Interaction, these folks talk about three aspects that should be improved in recommendations. They are similarity (discussed above), recommendation serendipity, and the importance of user needs and expectations in a recommender.

Serendipity

Take the list of "Special sauce recommendations" for Dune by F.Herbert. On the list of 20 books you can find on the top 2 of his other books, and 2 by B.Herbert, his son. This sounds rather dull and not really anything surprising, you could find that easily from a bookstore too. By serendipity, the authors mean how unexpected the recommendation is for the user and how novel it is. For me personally this is a very important factor and why I like the idea of recommenders, as opposed to just content-based retrieval of resources.

I won't discuss the importance of user needs and expectations in a recommender, as in this case it is pretty clear. In some other cases, though, like for learning resources, this comes very important, as teachers do have different tasks at hand when they are looking for learning resources. This is something I've blogged before about and will keep exploring in my context of research.



McNee, S.M. , Riedl, J. , and Konstan, J.A. (2006) "Being Accurate is Not Enough: How Accuracy Metrics have hurt Recommender Systems". In the Extended Abstracts of the 2006 ACM Conference on Human Factors in Computing Systems (CHI 2006) [to appear], Montreal, Canada, April 2006

Rafaeli S., Dan-Gur Y., Barak M. (2005), “Social Recommender Systems: Recommendations in
Support of E-Learning”, Journal of Distance Education Technologies, 3(2), 29-45,
April - June 2005.

Swearingen K., Sinha R. (2001). , “Beyond algorithms: An HCI perspective on recommender
systems”, ACM SIGIR 2001 Workshop on Recommender Systems, 2001.

Thursday, April 19, 2007

D.Watts on Social influence and popular songs

This study is pretty interesting: there were 14,000 participants who were asked to listen and rate songs by bands they had never heard of. The point was to study the social influence, i.e. how seeing cues from other people, like Top10 downloads, no of downloads, etc. would influence on people's choice.

The set-up of this study is pretty neat, the participants were sliced into eight parallel “worlds” so that participants could see the prior downloads of people only in their own world. Everyone started from the same line, zero downloads — but because the "worlds" were kept separate, they subsequently evolved independently of one another.
What we found....In all the social-influence worlds, the most popular songs were much more popular (and the least popular songs were less popular) than in the independent condition. At the same time, however, the particular songs that became hits were different in different worlds, just as cumulative-advantage theory would predict. Introducing social influence into human decision making, in other words, didn’t just make the hits bigger; it also made them more unpredictable.

Our experimental design has three advantages over both theoretical models and observational studies. (i) The popularity of a song in the independent condition (measured by market share or market rank) provides a natural measure of the song's quality, capturing both its innate characteristics and the existing preferences of the participant population. (ii) By comparing outcomes in the independent and social influence conditions, we can directly observe the effects of social influence both at the individual and collective level. (iii) We can explicitly create multiple, parallel histories, each of which can evolve independently. By studying a range of possible outcomes rather than just one, we can measure inherent unpredictability: the extent to which two worlds with identical songs, identical initial conditions, and indistinguishable populations generate different outcomes. In the presence of inherent unpredictability, no measure of quality can precisely predict success in any particular realization of the process.

This makes me want to test and set up experiments, too. In the project that I'm part of, and where I will get my data, we are planning some experiences on the input part of the tags to see how social influence in terms of seeing other users' tags when inserting own ones, will effect on the nature of tags, their number, their convergence, etc.

But, on the retrieval side of things this would be very interesting too! To have two different interfaces to see the search result list, where on the one there would be all the social cues for social influence (no of downloads, no of bookmarks, others' tags), and on the other one there would be nothing. The experiment would test whether the users, in this case teachers, would be viewing the metadata of similar resources and what would they actually download, bookmark and rate, if they did any.

Well, actually the latter is the situation as it is now. So maybe I can just compare the data from this year and the year after, when we actually start implementing the social navigation part.


Link:
In NYTimes

Science 10 February 2006:
Vol. 311. no. 5762, pp. 854 - 856
DOI: 10.1126/science.1121066
http://www.sciencemag.org/cgi/content/full/311/5762/854

Supporting material:
http://www.sciencemag.org/cgi/content/full/311/5762/854/DC1

Thursday, April 12, 2007

What tasks teachers have on a learning resources portal?

The attitude of "if we build it, they will come" has resulted in national and regional learning resources portals where the offer, no matter how many learning objects or assests, does not necessary match the need of teachers. Why is that? What is it that teachers look for?

Curriculum coverage

I first started by looking at 29 different learning resources portals that national and regional educational authorities offer for K-12 teachers in Europe. My task was to find out how many of them offer curriculum related material, i.e. so that a teacher, knowing that tomorrow he has to teach an area that covers a certain goals of the national curriculum, can just go to the portal and pick a resource that actually goes through this particular area.

This type of standards-based curriculum material seem to have been on the offer in the US for some time. Two examples could be the DLESE http://www.dlese.org, focusing on Earth Science and IDEAS, http://ideas.wisconsin.edu a repository held by Wisconsin educators. In both teachers can search for a curriculum coverage.

I found that in about 10 out of 29 examples in Europe teachers can explicitly look for curriculum coverage on learning resources, however, it was not always very clearly indicated which goals or skills a resource aims to attain.

On rest of the portals teachers were able to find resources that were categorised by the school level and subject, thus, by using teachers internal knowledge of curriculum, they would, with little poking around, find the matching material. But the problem with this case is that there is nothing that makes the teachers' tacit knowledge externalised for other users, no one else can take advantage of the fact that this teacher knows how well this piece of learning material covers certain areas and goals of the local curriculum.

In MELT, we are trying to get to the core of this problem by encouraging teachers to add tags to resources that they find in the repository. We'll be eager to see whether tags can externalise any of that tacit knowledge that teachers have, and that could help other teachers in finding and using the material.

In Calibrate, another project that I have a minor involvement, we have an opposite approach to the problem. A few topic areas have been selected where 3 EU-countries share a similar curriculum. We try to map the different curriculum through its goals and skills, and match that with the material available. A complex task!

Tasks at hand

Then, I started thinking whether learning material that clearly covers the standard-curriculum is what teachers want? One could think that a busy teacher would really appreciate it, however, there might be other reasons why do they come to a learning portal. So I made a little poll with some tasks that I thought teachers could have, and asked some teachers to vote. I only got 36 votes so far (you can still vote here), but it seems that, at least for these teachers, the curriculum coverage was not what they preferred!









The majority of the respondent teachers (I don't know who they are) seem to be leisurely window-shopping while at the learning resources portal (I browse around to see what is available) or looking for material that could support the lesson that they are planning. Only one teacher is looking for learning resources with an exact curriculum coverage!

This rises two questions in my mind: either teachers have given up to look for curriculum covered material, because a) it never was on the offer, b) it was never successfully on the offer or c) they have better sources for that, like school books, etc.

Secondly, maybe teachers don't really know what to expect from a learning resources portal or a repository. Maybe it was never clear for them what was the intended goal of a learning resources portal and they just come there to see what's up, what's in there and maybe they return if something good is found.

Seriously thinking, do we even know for what tasks and goals the resources repositories are build? If we think of libraries, we know they have a clear goal, or a school book, yes, it has a clear goal. But a LOR (for learning object repository), isn't it something in-between, kind of pretends to be one or the other, without being either of them.

I just looked at the sneak-preview for Yahoo!'s new service for teachers. Pretty neat. They clearly aim it to be to create learning resources, and re-use the ones from other teachers. They also offer neat peer-networking possibilities. It's about using and re-using material that is in the center, not searching for it! A very different focus.

So, I think what should be build in on learning portals is a better support for the tasks that teachers and learners have at hand. So, when they come to a portal,
  • if they clearly are just window-shopping, let's provide them with first-grade Champs-Elysees shop-windows. Make nice pre-views available of resources that are there with added value lesson plans and case studies how others have used them in their lessons. Allow browsing other users' collections of learning resources with annotations and comments. These other users can be from any part of the continent, as the goal is to inspire and show how things can be done. If we know anything about the user, let's try to match them with like-minded peers!

  • if they look for material with curriculum coverage, let's lead them to an area where they can either search for material with curriculum coverage, or browse bookmarks and tagged resources for cues from other teachers on attaining certain skills and goals. For the latter, it could be useful to first show "traces", e.g. bookmarks, tags, pedaggical annotations, from teachers who come from the same area, like Yahoo! peer-network allows getting close to teachers from the same area (=same curriculum).

  • etc.
The point that I'm making is to be clear about the task, and the information seeking patterns in general that teachers have, and then match it with the best way providing search, social navigation, recommendations, shop windows, etc., but always thinking what would yield the benefits for the user and task at hand.

That's something to study deeper, and don't worry, I'm on it ;) We don't know yet if the best benefits can stem from using underlying social networks (location) or more implicit ones (profile, tags, similar bookmarks), or from using a search or what?

Wednesday, April 11, 2007

Social Navigation as seen a decade ago

I'm always fascinated when I find "old" papers that still resonate today. Well, I'm not talking about the manuscripts from the Library of Alexandria, but I just came across a paper on "Design Principles for Social Navigation Tools", written in 1998. The principals still seem rather relevant!

That makes me think; what was social navigation about a decade ago and how differently we perceive it today, after all, there was no social bookmarking nor tags out there, like we know them today.

Defining different flavours of social navigation

Social navigation can happen in many different forms... One may distinguish between direct and indirect social navigation [Dieberger, 1998, Svensson, 1998]. In direct social navigation, we talk directly to other users. In indirect social navigation, we can see the traces of where people have gone through the space, as for example in the Footprints system [Wexelblat and Maes, 1998].

In my context of use, direct social navigation could be seen to happen through networks of friends and colleagues, that the users of a social tagging system have established. Or, it can also be sort of "ask the expert", or ask your colleague type of thing. Direct social navigation could most likely take place in sharing pedagogical practices, for example.

Indirect social navigation, on the other hand, would be like following other users collections of bookmarks, browsing them through " xx other user have this in their collection" or through common tags, for example in the personal or common tag cloud.

Furthermore, social navigation may be intended or unintended by the advice-giver.. distinction can be made for when the advice-giver is one particular person, known to us, or when it is just a group of anonymous users that have happened to navigate through the same space as us. In-between these two extremes, we may have groups of users that are similar to the navigator in terms of interests, profession, knowledge or task. The advice-giver may also be an agent [Foner, 1993]1.

The idea of having intended or unintended advice-giver is intriguing, many times in collaborative environments for learning, we are very occupied in setting up advisors whose intend it to give advices. But in the real world, people are somewhat shy asking a specific "advisor" for hints or help, peers seem to work better, or many people look for "traces" for that purpose. So, it seem to me that it is very important to design a lot of unintended advice-givers to help social navigation in learning and teaching contexts, they can be anonymous crowds, agents, traces, annotations, hints of task orientation, shared interests, or what ever we can think of.

I wonder if this table makes sense like it is now: (I'm not talking about blanc spaces...)






Unintended advice-giverIntended advice-giver
Indirect social navigationTag cloud guiding the navigationA recommendation (person, agent,recommender,..)
Direct social navigationBookmarks from network of friends to navigateA friend, expert, peer gives advice or recommendation


The 6 design principles

The principles according to Forsberg et al. (1998) are Integration, Trust, Presence, Privacy, Appropriateness and Personalisation. The examples are pretty hilarious, really like from 10 years ago, but nevertheless, the principles stand like they should!

Integration:
It is emphasised that social navigation should be part of everyday tools to make best use of it. This is what we see nowadays a lot, more and more things are integrated in our workflow, for example the browser with all the add-ons and blug-ins has become a central tool.

To make the best use of social bookmarking, it is of utmost importance to make the process easy and part of what one was doing right at that moment. The delicious-bookmarklet added in the browser toolbar is a good example, it is so easy to use that you hardly even have to stop what you are doing at that moment to bookmark. Thus, you leave more traces for others, besides arranging your own information space.

The idea of Attention Metadata is another one, if you use Slogger to generate Attention Metadata on what all you do on your web-browser, you don't even have to think of doing it.

Presence:
This refers to how do you make the presence of others shown to users, how do they know that someone has been here before. Annotations (tags, comments, evaluations, opinions, ratings...) are a perfect way to do that (Kilroy was here!), you know that someone passed through that space.

Or showing how many other users are online at the moment, think of how much fun is it to log into Skype at 1am and see that despite the quietness of your work room, some other people are out there still slaving away.

Trust:
In order to take a note of someone else's doings, it helps if you know whether you can trust them, are they a reliable source, do they like the same things as you do, etc. When reading film critics, one quickly picks up the critic who is like-minded, and disregards the other one who always seem to have too much of a mainstream thinking. Similarly, to trust the source for online social navigation can become crucial, if there are many ways to go forward.

Appropriateness:
In some situations one design choice is more appropriate than the other one, for example indirect social navigation can be suitable for finding inspiring information, whereas you might rather turn to someone to talk to when you have a specific question at hand. This is in my opinion related to finding a means that fits the purpose of the information seeking task at hand, something that I've been mulling around a lot and is related to the Human-Computer Interaction framework.

Privacy:
This concerns the issues of making users aware of the traces that they leave, data that is logged and used for personalising their searches, etc. This is also related to some codes-of-conducts that some systems have. Very important for things like Attention Metadata, Google personalized, etc.

Personalising Navigation:
"Social navigation provides excellent opportunities for tailoring navigational advice to individual user's task, knowledge or abilities". How I see this, it is related again to the information seeking tasks, interests, etc.



Forsberg, Mattias and Höök, Kristina and Svensson, Martin (1998) Design Principals of Social Navigation. In: 4th ERCIM Workshop on User Interfaces for All, Stockholm, Sweden.

@InProceedings{sicsprint95,
author = {Mattias Forsberg and Kristina Höök and Martin Svensson},
year = {1998},
title = {Design Principals of Social Navigation},
address = {Stockholm, Sweden},
url = {http://eprints.sics.se/95/01/designprincip.pdf},
booktitle = {Proceedings of 4th ERCIM Workshop on User Interfaces for All}

Thursday, March 29, 2007

the Anti-Social Software Suite..you see me, not..

If you've been wondering why I haven't been seen online or blogging lately, here it comes:
I've been there all the time, but just wearing my new tech-ware, the Anti-Social Software Suite!

Friday, March 02, 2007

Notes on " When tags work and when they don't: Amazon and LibraryThing"

There is an excellent posting at the LibraryThing's blog on a study on tagging, it compares tags in LibraryThing and in Amazon.com. What is also a great reading are the comments, there are many of them which are very deep and reflective, too.

Lots of issues in this posting touch upon what we are currently thinking in MELT, where we are implementing a social tagging tool for teachers to tag digital learning material. A few points that I will elaborate upon later:
  • Tag categories (factual, subjective, personal); how there are so many opinion tags in Amazon compared with LT.
  • The idea of separate "buckets" vs. "works" (LT)
  • How the tagging interface works and supports the workflow (to show tags from others or not? how does this effect on convergence of tagging?)
  • Tags are not reviews or ratings! they serve different purpose.
  • Reasons and drivers for tagging
  • Number of tags to become important
  • Why implement tags in the first place, what is there for Amazon? they already have all sorts of ways to recommend content, do they still need tags? it's just another feature on top of all other whistles, it's not part of the design for the site to work.

Thursday, March 01, 2007

OLPC social features

For quite some time I've been reading about the One Laptop Per Child initiative and the new GUI that they've developed. It's been hard to get my head around how it actually would work based on screen shots.

Today I came across some YouTube clips on OLPC. Most of them were awful! Get this, they show the Linux booting up in verbose mode - as if anyone's interested!! Rest of the stuff is showing some features in too small font without commentary.

Luckily there was also this one, it pretty nicely presents the social features as part of design.



I would really love to be a fly on the wall when these laptops are introduced to kids!

I wonder what kind of sense they make for users who most likely have not seen conventional computers, nor used the Web.

A funny anecdote: I remember my fist time playing with Mac in summer of '84. I kinda liked the graphical interface (until then I had turned my head away in much of dislike from my brothers Commodor), but I was really quite baffled with icons like the trash can for delete. In Finland, the trash cans didn't look like that at all, so it was rather hard for me to make the connection. Of course one learns pretty quickly...

Connections between working items in EUN, KUL and COSL

Visiting The Center for Open Sustainable Learning (COSL) in Utah State University last week was a very rewarding experience. People there are working on many interesting things regarding the support for the use of open content in education. I had already checked their website (http://cosl.usu.edu/) before going there, but of course it's nothing compared to sitting down with people, talking about my projects and their projects, and, little by little, creating understanding on complimentary approaches and ideas to collaborate (hopefully).

To get the big picture on where I stand between all the work done in European Schoonet (EUN), KULeuven (KUL) and COSL, I created a complicated looking image that hopefully helps me to put pieces of the puzzle together.

The image here on the left represent a rather generic lifecycle of a learning resource (this is the connected part of the image) and the greenish bubbles around it are working/research items that people in EUN, KULeuven and COSL work on. With a quick look it seems that all the institutes work in the same field - providing access to learning resources - however, each has its own focus and flavour to give.

Who and what?

KUL is focused on searchability of resources across different repositories (SQI, OAI-PHM) and emphasises the generation (ALOCOM, AMG) and collection of metadata (CAM), both to describe the resource (LOM) and its context of use (Contextualized Attention Metadata).

COSL works on making the content accessible (EduCommons CMS), but is also focuses on tools to support the use of content (MOCSL). Moreover, the relations of content and content, content and users, and users and users seem to have become more important (Didly, Oz). They are also working on a recommender system.

EUN has interest in making the content available (federated search, Minor) and using the social aspects for its retrieval, the main methods being social bookmarking and user annotations (MELT).

Let's take it step by step following the scenario

1. Repositories to make content available

This image depicts a rather typical scenario of learning resources which end up in a repository; first the resource is "created", then"described" (i.e. metadata added) and "approved" (e.g. QA) to be part of a repository's collection where it is "published".

EduCommons, a plone-based CMS created by COSL folks has been developed for this reason.

eduCommons is an OpenCourseWare management system designed specifically to support OpenCourseWare projects like USU OCW. eduCommons will help you develop and manage an open access collection of course materials.
Whereas in EUN MINOR was developed. It's an open source learning resources repository that allows users to manage learning resources and their metadata. Minor comes with a built-in connector to the EUN's Federation of Internet Resources for Education, which greatly facilitates the work of connecting to other repositories and querying their resources.

2. Make the content available for search

Both EUN and KUL has been active in making the resources and/or metadata that reside in distributed collections and repositories accessible. Both have greatly contributed to Simple Query Interface (SQI). EUN and KUL, for example, demonstrated a real-time interoperability between Learning Object Repositories last summer.

Apart from SQI, folks in KUL are also looking into harvesting metadata, which seems like a suitable and complimentary solution for SQI.

3. Accessing the resources

Once the resources are made available in a way or another, there are the users who "discover" them using various methods. This is an area where quite a lot of work in done in all 3 places.

In EUN and KUL, within the MELT-project, the focus is on creating Social information retrieval (SIR) techniques for flexible access to large-scale collections of content. Here we are interested in social bookmarking and tagging, as well as on other annotations like comments, ratings, etc. This work can also lead to an implementation of a recommender system for learning resources within the federation of repositories.

COSL has also been busy thinking of how to use people, and relations that people and content have, to create serendipitous ways to discover resources. Ozmozr (Oz) is, or will be, an identity and content aggregator that leverages the power of social collaborative filtering through tagging. Currently it looks somewhat messy, but try to see for the grand idea that lies behind.

There is also Srumdidilyumptio.us, just didly among friends, that allows users to express relationships, let them between people and people, people and content or content and content, i.e. any two resources accessible by URLs. Didly might not look that cool, but one should think of it rather as an API to make use of that information in another context, I was told.

These two tools/concepts pretty well demonstrate the thinking that COSL folks take forward. One could say that they aggregate s!#t out of everything, to illustrate the Web 2.o thinking. Moreover, COSL will be working on topping up these two "relations aggregators" with a recommender system. The difference to the work in EUN is that they are planning to recommend anything that people have input in Oz, whereas EUN is only talking about recommending learning resources.

More in the same area, novel ways to access resources, one of the PhD students in KUL is working on creating a better ranking system for resources, namely the LearnRank. This will use the Contextualised Attention Metadata framwork (CAM). Another one is busy with information visualization techniques for the same purpose. We are also interested in experimenting on the combo of information visualization and social bookmarking in the future.

4. Support the decision making

Next step in discovering resources is to make the decision upon the piece of content. Usually users are exposed to a list of search results (sortable sometimes, ranked or not) or a list of recommendations. There are different ways the system can support users at this point.

EUN and COSL are both interested in user relevance judgments (AnnoRate; use of ratings and evaluations in EUN) and other annotations such as pedagogical comments, to help users to choose the content. EUN has also implemented personal collections of learning resources, so users could see in how many other collections the resource has been added in, to help them to judge its relevance.

After choosing the resource, users in EUN can add it to their favourites collection (bookmark) or choose to "obtain" the resource, its metadata or uri from the repository. Ozmozr, a COSL tool, allows users to bookmark items too, and to share them with other users. Additionally, at the same time, users can also save these bookmarks to their delicous account, for example.

It's good to make here a remark that CAM framework can be used to collect all kinds of user attention when interacting with a repository, for example, to save the search history, what items have been looked at or obtained, etc. This can be fed in again to help the discovery of resources, like is the case of LearnRank.

5. Re-using and re-purposing the content, content collaboration


Once the items have been obtained, there are the users, teachers or learners, who modify, "re-purpose", sequence or rip the resource to make it suitable for their needs.

In this sub-area ,some complimentary approaches have been taken too. Within an EUN-lead project a plone-based collaborative platform called LeMill has been developed by UIAH. LeMill is a web community for finding, authoring and sharing open and free learning resources.

In COSL some tools are under construction to support the use and re-use. Send2Wiki plug-in allows a user to send a webpage to a wiki and modify it there. It will automatically check on the CC licence and bring that info along. User can then modify the page in the wiki as they wish, for example to make a translation of it (hope to pilot this in EUN), or if it's from WikiPedia, modify the content to make it more suitable for K-12, for example (now some researchers in COSL say that Wikipedia is rather graduate level reading, thus the need for different versions).

Make a Path will help users in putting pieces of resource together for a coherent learning path. This tool is like creating your own music playlists. Once the users are using it (I hope EUN will pilot this, for example), some interesting patterns will most like start appearing. One possible way to take this forward to help sequence learning resources could be case based reasoning, like Claudio did for generating music play lists. Additionally, The INSTRUCTIONAL ARCHITECT allows teachers to find, use and share learning resources from the National Science Digital Library (NSDL) and the Web in order to create engaging and interactive educational web pages.

Moreover, there is research interest to better understand the collaboration aspects and communication patterns around content, how people work on it, share it, ask questions, help others, etc. In LeMill the collaboration aspect is important, one of the developers is currently conducting his PhD research (UIAH) on collaborative authoring of learning resources. Whereas in COSL, there is interest in looking into Yahoo! Answers, as well as there was some early work around supporting the users of MIT OpenCourseware.

I think KUL ALOCOM framework and AGM could be used here to make sure that a) the smallest possible pieces of content are saved in the repository, too, b) the correct metadata is sucked from the context. Also, CAM could be used to better track users interaction modes and patterns, etc.


6. Actual use of content and its evaluation


After ripping and mashing the content, it it usually "integrated" (e.g. using CMS, etc) to teaching and learning practices (pedagogical methods). "Use" of the learning resource takes place in its own context. Here, the CAM framework becomes useful to know more about how the resources have been used, in what context (e.g. entered by a university prof to a course, all this metadata can be sucked with the help of AGM and CAM).

Following the use, the users are given possibilities to evaluate or rate the content. The challenge here for a federation of repositories is that different repositories use different ways to rate and evaluate the content (different scale, different criteria (ped, technical, suitability in classroom,..), thus interoperability becomes cumbersome and there might be something lost in semantics (EUN). A rating service on a top of repositories like AnnoRate overcomes such hurdles.

To shortly summarise:


Everyone is interested in adding metadata, any kinds of it (LOM, attention, annotations, relations (bookmarks, sequence, ...), to the system












and using it for better retrieval of resources, and the new thing, to connect people.
















"We use people to find content. We use content to find people. Information seeking behavior and social network analysis go hand in hand." Peter Morville (2004)
The really interesting part of a folksonomy is not the content item being described, and not the tags that describe the item, but the person doing the tagging. by Greenchameleon

Wednesday, January 17, 2007

A really bad idea of crowd-sourcing the Web

12,000 Estimated number of illegal immigrants apprehended at the U.S.-Mexico border in November

10 Number of illegals apprehended in November thanks to a $200,000 experimental website allowing anyone to monitor the border and alert authorities. Viewers sent 14,800 e-mails through the site.


Oh my god, what's gonna be next?? It's not enough that we can practice neighbour-watching in our own neighbourhood, but now, thanks to the Web, we are able to big brother our neighbouring countries, practice border patrolling, and renounce illegal immigrants. A little respé, svp!!

http://www.time.com/time/magazine/article/0,9171,1576859,00.html

Notes on Combining Social- and Information-based Approaches for Personalised Recommendation on Sequencing Learning Activities

This paper, Combining Social- and Information-based Approaches for Personalised Recommendation on Sequencing Learning Activities, cames from OUNL and is related to a bigger schema of works that those guys are carrying out on competencies. Thus, the context is very related to lifelong learning, namely to higher ed and vocational training.

The aim of the paper is to describe a domain model for "way finding". By the term "way finding" is meant "selecting and sequencing learning activities". The raison d'etre is:
Learners' problems in way finding will decrease the efficiency of education provision (the ration of output to input) and increase the cost. The local context for this paper is Dutch Open University student who lacks adequate information on study possibilities at an early stage of study, and problems in getting a good overview of the number and best sequence to study modules.

The paper describes a personalised recommender system (PRS) model that combines social-based (i.e. completion data from other learners) and information-based (i.e. metadata from learner profiles and learning activities) data to recommend the best next learning activity. The system is currently under development, a limited implementation is running using learner profile metadata.

Note about learning activities; OUNL has been very active in developing IMS Learning Design. They (Tattersall et al. (in press)) have previously proposed IMS-LD as a candidate to model learning paths. Moreover, they argues that its selection and sequencing constructs appear suitable for learning activities (units-of-learning) as well as for higher levels of granularity (e.g. competence development programmes). Interesting. At one point of time one could look how IMS-LD information could be generated in attention metadata (CAM).

So, the idea of PRS approach is a hybrid recommender that uses
  • a) information from other learners and their completion of tasks (completion is understood like rating) in a collaborative filtering manner (in text called social-based approach), and
  • b) information from students profile and c) metadata about the resource in the spirit of a content-based system (in text called information based approach).

Authors also argue that it is not enough to find the most efficient learning paths (like the shortest route in GPS), but to explore which paths are most attractive or suitable (like routes suited for biking), thus personalisation needed (individualised needs, interest, preferences or circumstances).

Other key concepts are:
  • learner's start position in a given domain (prior learning history)
  • aimed competence profile for that domain
  • learning path towards that competence
To develop a PRS the authors identify the following pieces, that the paper defines:
  • uniform and meaningful description of formal and informal learning paths
  • learning activities that are addressable and meaningfully described
  • uniform learner profiles that define needs and preferences
  • uniform competence description that defines proficiency levels
  • a learning path processing engine
  • an engine recording completion of activities
  • information matching techniques to enable personalised recommendation

Related work

I will later post on my blog some excerpts from my own literature review in the field of learning to show other recommender ideas based on the same hybrid approach, as in this paper they mention that this approach has hardly been applied in learning. There was only a reference to Herlocker et al. (2004).

In the related work section it is mentioned that education field imposes some specific demands for recommender. Main differences sited between recommenders for books are the degree of voluntariness (learning is many times required to obtain some goals) and the possibility to establish an explicit completion (as most learning activities are to be assessed for successful completion). Hmm..I do agree with the statement, but had come up with different reasons myself. Goes to show, I guess, how the initial requirements for a recommender system differ from what I'm working on.

An interesting outcome is cited from Janssen et al. (in press): learners were offered a recommendation "most successful learner continued with Y after having completed X". I call this an "Amazon-like" recommendation (other people interested in this book also bought x, y, z) based on clustering behaviour. There were no personal characteristics taken into account in this study. They found out that this type of recommendation enhanced effectiveness in completion of the set of learning activities, but did not increase efficiency, the time it took to complete them.

Authors also acknowledge the problem of insufficient data that can be derived from existing log files, the same that my colleagues are working on with the view on capturing attention metadata (CAM).


Discussion

Authors state:
"From a self-organising point of view it would be ideal if way finding would emerge as a result of (in)direct interactions between members of the learning network, without being dependent of formalised descriptions in domain and user models."

I so agree: instead of investing time in describing all the information regarding the learner, his/hers existing and required competences; the resource; and the curriculum with goals and skills required, would be more interesting to tap onto existing knowledge from the masses and their previous experiences, the decisions they took to find the next suitable step, etc.

The authors also discuss the complimentary approach of controlled vocabularies or ontologies combined with annotations such as social tagging and rating, just in the same direction as we are doing in MELT (we talk about adding metadata a priori and a posterior) and what I'm interested in looking into.

Related to my work

The difference in what I'm looking into now and what this paper describes is, first of all, the context. I'm interested in a repository that is used mostly by K-12 teachers and learners (sometimes). The repository is not linked to formal learning requirements related to a curriculum, because it is used on the European level, where there are many curricula depending on a country or a local policy. However, each teacher who comes to that repository has his/hers own information seeking tasks, that I've talked about previously. Sometimes those tasks are related to covering a piece of a local curriculum, whereas some other times it is to find a piece of resource to support some generic learning goal, or find inspirational material, or something else.

Secondly, in my context of work sequencing learning resources is not the goal, rather just finding resources that fit to the search criteria and the task at hand. So I'm not so into this sequencing, however, I like the idea of playlists and using this type of expert knowledge of putting items together for learning purposes (the use of Case-based Reasoning like Claudio explore here).

Imagine if teachers could generate playlists of LOs as easily as I do playlists in iTunes (I'm NOT talking about automatically generating them, but hands-on deejaying). Then, those lists could be used as rules for generating new ones. In that case, we would not need LOM to know which item in a repository is described as "introductory item" or "motivational item" to start the lesson, or which one is good for "explaining a rule", but we could detect that information form playlists generated by teachers who knows through her domain knowledge that after this piece I put x,y and x. Let's see.

To check out from the paper:
- Koper (2005)
- Janssen et al. (in press) about the test
- Sicilia (2005) about ontologies to express competencies
- Van Setten, 2005 social-based approaches
- McCalla (2004), pragmatics-based paradigm of tagging learning activities with learner information

Wednesday, January 03, 2007

Fighting the read/write web-fatigue with interoperability

My problem with social software sites has long been my short attention span. I love to log-in - but not fully create my full profile as it is soo timeconsuming - I play around for some time to test and understand how some of the features work, and then I forget about it. Some random emails from even more random people wanting to make me their friend sometime remind me of the service. However, it's hard to go back as I can't even remember the password or which email I used to sign up.

This post, ..(cuz losing passwords is common amongst teens), really made me laugh about how teens use the Web. According to that teens would forget the password to enter to the social network service, and without any hesitation, they start a new profile, and email for that reason. I wonder if this is more the nature of teens than a new trend emerging among young users of Web?

Maybe teens just don't think the whole thing (i.e. social network sites) is that meaningful or worth saving. Or better, maybe they just really don't think about building a consistent profile of themselves, yet. Hell no, I would hate if all the stuff that I once did on the Web would be still available and indexed in Google! Everyone needs to start once in a while from scratch, cuz old habits stick (and stink)! But, maybe once those kids think that it's meaningful enough for them, and they will start remembering the passwords.

Or better, when they are old enough that they really care and want to keep all the digital pieces together, hopefully there are better ways to keep track of "yourself" than separate services with no portability of content that only rely on stupid passwords. There really should be better ways...

Anyway, the two points that I have seeing since I've opened my computer after vacations (yeah, I know how to log off) is Web-fatigue and portability of content, contacts and profile information in social networking sites, but I would really want to bring it to the whole field of social software.

Could 2007 be the year of social network fatigue? by ZDNet's Steve O'Hear --Another driving force for social networks in 07, will be the increasing
number of niche networks which are highly targeted to particular
interest groups or social activities. The question that still remains
however, is how many social networks any one user is likely to join and
remain active in? This is where Read/WriteWeb's prediction of fatigue
has more weight. Unless the time required to sign-in, post to, and
maintain profiles across each network is reduced, it will be impossible
for most users to participate in multiple sites for very long.
Therefore I think it will be essential for social networks to open up, through embracing open standards which allow for greater interoperability between networks.

Hell yeah, that would be lovely! That's what I've been wanting for some time now, well, ever since I started using social bookmarking services. I really like Furl, I think it's far nicer service to use than delicious, which I only use occasionally, because of peer-pressure, everyone else is there - well, to leverage on the crowds. But I just don't like it, for reasons that I won't go in this post.

Nevertheless, what I would like, is that having the profile and bookmarks that I've accumulated in Fulr could be taken advantage of also in delicious. There should be some way of updating my profile at the same time in both services, or that the services would have some way to make a personalised federated search across the services - sort of meta search across social bookmarking sites that would

a) allow me to browse other similar people's profiles,

b) would make active matching of my bookmarks to others in each service and recommend me links. Moreover, there should be some

c) possibility also to tap on my networks and contacts throughout the services, like that my delicious network would be notified of my new bookmarks in Furl.

Of course, now there are ways to do all that, subscribe manually users from one place to another, etc, but it would take ages to do it, and I would never be up-to-date in any of the places, I reckon.

So, this is to say that it would be really important to work on different types of interoperability between social software services. There's been posts regarding portability of contacts information, but also just plain user profile stuff (why can't I still even upload my vCard??), portability of social bookmarks (I once checked how about importing my personal links from delicious and furl to a repository of learning resources and found out that not even the RDF or XML was standard), etc.

The current way of wanting to lock-in people to a social software/service that they've started investing in (I'm thinking of investing time, knowing/inviting people and friends, creating and enriching profile, etc) is ridiculous. Only us, say, 30-years and +, are silly enough to stick around in places where we've started building up our personal portfolio of digital artefacts. We escape behind excuses like "I'm too busy to start a new blog and transfer my blogroll", or "I've lost my password to my domain name server to change it to a new one".. We should start seriously asking the providers for these services, to give us better portability of our data, do more open standards based communication layers to enable federated searches across services, etc.

Voila! My wish for 2007 is better interoperability for social software! Or otherwise, I'll just start behaving like those teens ;)





Tuesday, November 21, 2006

A Case study on 5S

5S, a fundamental abstraction of digital libraries, stands for Streams, Structures, Spaces, Scenarios and Societies. 5S proposes a formal language and model to describe digital libraries.

The advantage of using a formal language and model, according to its creators (Goncalves, et al. ), is that they are precise and unambiguous when defining semantics of specific abstractions of a knowledge field. Thus, 5S could be used as an instrument for
1) building and interpretation of a DL taxonomy,
2) informal and formal analysis of case studies of digital libraries and utilisation as a formal bases for a DL description language.
Moreover, 5S could be used for requirements analysis in Digital Library development.

An example of a case study to describe a digital library using the main elements of 5S is presented below, taking Learning Resources Exchange as an example.

5S, a fundamental abstraction of digital libraries, stands for Streams, Structures, Spaces, Scenarios and Societies. Streams are sequences of arbitrary items used to describe both static and dynamic content. Structures impose organisation. Spaces are stets with operations on those sets that obey certain constraints. Scenarios consist of sequences of events that modify states of a computation in order too accomplish a functional requirement. Societies are sets of entities and activities and relationships among them. The below case study illustrates of what each “S” is comprised of.


Societies
  • primary community: European teachers and learners
  • repository providers: public authorities who make the decision about joining
  • repository maintenance and running, editorial control
  • some pedagogical support, etc.
  • commercial providers

Scenarios (services, different corresponding scenarios)
  • trainings to help authorities to hook up to the federation
  • training to help teachers to use the system and submit material to it and creating the metadata
  • scenarios to maintain the service and content?
  • Scenarios on how to access the site, through browsing and searching, in one or federation (SQI)

Spaces
  • physical location of members (a metric space)
  • vocabularies and metadata used in different services (conceptual space). In addition to this, also manual, semi-automatic, and automatic indexing and classification methods to relate repositories to the conceptual space of ELR.
  • User interfaces (APIs?) to relate various software routines (like LMS etc)

Streams
  • simplest level they are streams of characters for text, and streams for pixels for images; audio, digital files. Challenges for quality of system if in real-tine or storage problems at the local level if downloaded and locally played
  • Network protocols, transmissions of serialised streams over the network such as federated search, harvesting, hybrid services using protocols like Dienst, Z39.50, OAI-PMH,...

Structures
  • database management system at the heart of the software for submission and workflow management.
  • Xml to store and exchange the resources' metadata
  • EUN Application profile
  • structures in the form of semantic networks, any?

S5 as a description language.

The paper, using S5 Descriptive Language, defines a digital library as follows:

A digital library is a 4-tuple (R, DM, Serv, Soc), where
  • R is a repository;
  • DM= {DM c1, DM c2, ...DM ck} is a set of metadata catalogs for all collections {C1, C2,...Ck}
  • Serv is a set of services containing at least services for indexing, searching and browsing;
  • Soc is a society.

“ We should stress that the above definition only captures the syntax of a digital library, that is, what a digital library is. Many semantic constraints and consistency rules regarding the relationships among the DL components (e.g. How the scenarios in Serv should be built from R and DM and from the relationships among communities inside the society Soc, or what the consistency rules are moment digital objects in collections of R and metadata records in DM) are not specified here. Those will be a subject of future research. “

Tuple: in a database, an ordered set of data constituting a record; a data structure consisting of comma-separated values passed to a program or operating system.

GONCALVES, M., A., FOX E., A,, WATSON, L., T. and KIPP, N.,A. (200 ).
Streams, Structures, Spaces, Scenarios, Societies (5S): A Formal Model for Digital Libraries
Virginia Polytechnic Institute and State University
http://www.dlib.vt.edu/projects/5S-Model/

Monday, November 20, 2006

Workshop: Impact of Social Software on Society: PROWalk Event

Last Friday I participated in a BlogWalk event in Bonn. Crowing intensively weary of conferences where participants sit down on their asses making the best use of wireless Internet-connection to catch up with all the mails they missed because of the travel, I took a personal engagement to actively participate in this one.

Already for some time I've been aware of the idea of BlogWalk, but only now had a chance to have my first experience. I must say that I missed the "walking" part, we did not leave the conference venue to explore other areas. Nevertheless, mentally we did and it was interesting 3h of interactions, introductions, supporting questions, personal experiences and accounts, monologues, and post-its clued on the window (i.e. window-wiki, see the pic below).


The theme was "Impact of Social Software on Society". The discussion started from personal accounts and feelings; how information acquisition has changed in last few to 10 years. Many of us had similar experiences of following tens to hundreds of blogs through Web-feed reader, instead of reading a few authoritative newspapers and catching a number of news channels to get a different point of view on affairs.

Currently, instead of getting a "pre-filtered" view, one aggregates and filters (through friends) news, opinions, digital prints and artefacts by choosing whose blogs, what news and whose pictures/videos to aggregate and actually read. There are many different context one works in during the day (different way of working, that knowledge worker-stuff).

Overdose of all this was also discussed, how time consuming it is to filter all this information, read and mull it over, etc. Check out "continuous partial attention", but it's not only about email flows or instant messaging, it's increasingly about Web-feed loads, too! From my own point of view, just to balance out all the rants and not-so-well-formed opinions on the blogosphere, there is nothing better than reading The Economist, one of the best source of journalism. I sometime think of it as an ultimate opposite to blogs: impersonal (you have no idea of who has authored the article), although views on economics, and the world (in that order) are very pre-set, they know how to be self-critical, and so on. All this coined with a previous discussion with a pal about how he, previously an avid blog reader/writer is gonna only start reading peer-reviewed articles. Of course it was just a half jokingly said, but there is a little truth there too. We are probably soon gonna see some blog-fatigue in a way or another.

While talking about all this, we posted notes on the window with issues that we thought relevant for the BlogWalk. The following categories emerged:

1) Usage (personal level)
2) Effects and Fall-out (group level)
3) Contexts of usage (organisational/societal level)
4) Convergence (deep changes behind it all)
5) On-line interaction and/vs face to face
6) Design and evaluation of tools

"These clusters could be combined into one general narrative, starting from the personal level, moving upwards towards the group/organisational and societal level. This leads to a view on the deeper changes behind it all. At the same time from the personal level, moving downwards, we can cover social affordances and technical affordances."




One gap that was identified is how little research is actually done on social practices on how the tools and systems are used. We were discussing about many different ways of how and for what reasons people aggregate Web-feeds, for example. A cognitive point of view on this is missing, which, in a way, hinders us to see and understand the phenomena that social software can be part of in our society. Some empirical descriptions are missing, or like it was put on the wiki:

To understand the social implications we need more storytelling/ethnographic/anthropologic research that starts with the individual, in stead of just looking at the 'big numbers' and large patterns. (More stuff like by Efimova and Ben Lassoued, and my blog articles on my personal information strategy: 1, 2 and 3)


Nowadays many skeptics laugh about blogs that they have only one reader, the author, attempting to put them down. Hec, blogs are great for self-reflection, too. How needs an audience for that? If we had better understanding of a diverse ways people use these tools, we could maybe make them have a better support for all the diver ways. Which made me think of many of the read/write web tools and systems being on their perpetual beta, and whether that is a good or a bad thing. It could be good, in case the developers actually followed how people use their tools and developed them accordingly. However, it can also mean that they have no interest in supporting any new features and functions, and they just leave it hanging. Thus, better exploratory field studies and different contexts are needed to be studied, which would lead to more rigorous field and lab experiments to understand this all better.

For once again, attention metadata was discussed, on the one hand from a privacy point of view, and on the other hand, as a way to better understand how users use the wide palette of tools and systems.

Wednesday, November 15, 2006

Where tags and structured vocabularies co-exist, case for tagging learning resources

NOTE: this is a draft and lot of thoughts put together

Context

In the Learning Resources Exchange-portal end-users, i.e. K-12 teachers from all over Europe can access digital learning resources that are made available by European Schoolnet or by its partners from different learning repositories from a number of European country. All of the learning resources have associated metadata to them, all of the partners use Learning Object Metadata, and many of them an application profile specifically geared to European K-12 education. The indexing keywords come from a multilingual Thesaurus that currently exists in 14 languages. Additionally, teachers can add their own keywords to learning resources when saving them to their personal collections, that function like bookmarks. Also, in their personal collections teachers can evaluate their learning resources, rate them and add comments. This paper focuses on vocabularies, both unstructured and structured ones, and proposes a complimentary approaches to their use to enhance learning resources retrieval. We think that tagging has the potential to enhance the retrieval, both through the complementary features that it can offer for controlled vocabularies, as well as through its underlying social structures.


Structured and Unstructured vocabularies to access resources


The table "axes of organizational vocabularies" depicts the dimensions of structured and unstructured vocabularies. Learning resource repositories, for example, often use an authoritative vocabulary to index the material that is based on their national or regional educational needs such as requirements in curriculum or educational levels. Some of these vocabularies are structured domain taxonomies with hierarchies (LRE thes, EUN, Italy; Mobus thesaurus, France;,..) whereas others might be unstructured glossary type lists of terms (eg....). These vocabularies are used in educational repositories and portal most often to provide access to resources by classification or allowing hierarchical browsing, search functions are equally applied, however, teachers seem to prefer browsing features (based on focus groups and some Calibrate logs).



















StructuredUnstructured
AuthoritativeDomain TaxonomyGlossary
Personal/OpportunisticHierarchical FilesystemFolksonomy

Axes of organizational vocabularies, Iverson (2006), CC some rights reserved.


As opposed to authoritative vocabularies, personal tagging of resources provides an unstructured vocabulary for resources. These end-user generated keywords serve, in the first place, personal knowledge management needs. For example by bookmarking resources one creates a link to the resources to access them later, and the use of personal keywords allows naming things in a meaningful way, grouping them, etc. Depending on the user-base and behaviour, tagging can start providing a community based vocabulary for the given community of users. This is commonly known as folksonomy. It is important to note that not all the actions of tagging and personal vocabularies allow a community folksonomy to emerge, thus the division to broad and narrow folksonomies (Vander Wal, 2005).

Classified categories or domain taxonomies offer a way to access resources, either through keyword based retrieval or browsing through categories. A solid indexing strategy based on controlled vocabularies guarantees that the searcher does not need to face all the uncertainties of natural language (synonymy, polysemy, homonymy), combined with all the uncertainties of a full text search (no relevance control on the retrieved occurrences) (Trigari, 2002). Both searching and browsing, based on metadata and controlled, structured taxonomies, are methods most used to access resources on educational portals.

Browsing and searching serve two fundamentally different information seeking tasks. It is similar to the difference between exploring a problem space to formulate questions, as opposed to actually looking for answers to specifically formulated questions (Mathes, 2005). One could say that searching happens at the stage where user already have a clearly defined task, a well formulated question, and an idea how and by what information sources to fulfill it (reference), whereas browsing many times occurs when the user has not yet a clear question or task at hand, rather just a hunch of what to look for.

Second important advantage of folksonomies is that they can serve the community by promoting a serendipitous access to resources through browsing interlinked related tag sets, also known as "tagclouds". A tagcloud can be a personal one comprised only of the personal keywords, it can be a local one to display the most common tags by a given community of uses (e.g. based on domain interest, language, etc.) or even a global one allowing an access to all resources that are available through the service.

Complicity of relationships between thesauri and folksonomies

An interesting hybrid of vocabularies will emerge within the ELR context, where users add personal keywords to learning resources that are already indexed by an authoritative, structured domain taxonomy such as the LRE Thesaurus. As opposed to most Web-based applications where no top-down vocabulary exist and where users tag content with their personal keywords (del.icio.us, Flickr,..), the situation is different in ELR. Here, a learning resource is indexed using a 1 to 3 thesaurus terms that further allow search and browsing in a multilingual context, added also with multilingual tags by end-users. This will open a number of interesting research possibilities in the twilight zone of structured and unstructured vocabularies.

In the LRE Thesaurus the structure of the descriptors follows the classical semantic relationships such as the following: intra-language equivalence (USE-UF), inter-language equivalence, hierarchical BT/NT relationship and the associative relationship (RT). The research attention will be directed to study the relations between terms in all the areas. For the revision purposes of the thesauri, we will be interested in looking whether folksonomies could provide any new descriptors for the LRE Thesaurus.

Maybe the most potential area of complicity between folksonomies and thesaurus will be in the area of inter-language equivalence (USE-UF) between the thesaurus terms and folksonomies, which could be seen as a good source for non-descriptors. Non-descriptors are part of thesaurus, they provide the intra-language equivalence that facilitates the access to resources that are indexed by using the descriptors (i.e. thesaurus term) that do not translate well to the language that the end-user uses. For example, a user could search for antiquities, but the preferred term in the thesaurus, the one used for indexing, is "ancient history". The use of antiquities as a non-descriptor thus facilitates the access to the resource.

Example of inter-language equivalence (USE-UF)

Antiquities
USE ancient history
ancient history UF Antiquities

Moreover, an equally interesting area will be associative relationships (RT) and folksonomies. The associative or related term relationship is expressed in thesauri between terms that are mentally associated to such an extent that it's useful to make the link between them explicit, however, they are not members of equivalence set, neither are subordinated or superordinated to another (Trigari, 2001). We think that folksonomies could help to define and re-define a number of associative relationships in thesauri.

Example of associative (RT) relationship
NT1 old age
. . . . RT elderly person

As the ELR Thesaurus is also multilingual, we will be interested in investigating what kind of inter-language equivalence folksonomies could offer between languages. As for the area of hierarchical broad terms and narrow terms (BT/NT relationship) between the thesauri terms, we assume that folksonomies have very little to offer.

All in all, additionally to the above, for the revision purpose of the thesaurus, we think that folksonomies can be a great source to verify whether the scope of the current descriptors actually serves the need of its users for indexing and retrieval purposes.


To better understand the nature of folksonomies and tagging

A number of interesting research papers have emerged in the area of tags and folksonomies. Mathes (2005) described the phenomena, whereas the term and its main characteristics such as broad and narrow folksonomies were described by Vander Wal (2005). Broad folksonomies can be defined as social tagging systems where many people tag a few items (e.g. del.icio.us), whereas narrow folksonomies are the ones where content provider, usually the owner of the content, gives tags to items (e.g. flickr). Mainly, folksonomies emerge from broad ones, as there are many users and also niche user groups, who give the item the name that they know it for. Through the recent research we also know more about the evolution of tagging and how users provide them (Sen et al., 2006). Marlow et al., (2006) propose a model and taxonomy of different tagging systems. Moreover, Bogers et al. (2006), Lin et al. (2006) and Tonkin (2006) have provided interesting insights into the issues between classification and tagging.

By studying tagging and folksonomies more closely in this specific hybrid setting of learning resource,s that are indexed by experts and tagged by end-users, might yield some more interesting insights into its nature, both from qualitative and quantitative point of view. More so, we are interested in finding best ways to have social tagging systems to co-exist with more top-down indexing with structured vocabularies, and find out ways that can leverage such a system to better serve its users' needs. Also, research in this type of setting might allow us to better understand the many downfalls of folksonomies, such as ambiguity of tags, their synonym (different word, same meaning) and homonym (same word, different meaning) control (Guy, 2006). Moreover, we are interested in using the folksonomies to tap into the social structures and networks to, first of all, study and analyse the charasteristics and the main behavior of our user-base, and secondly, to use the social networks to enhance the information retrieval and serendipitous access to resources.


References


Bogers, T., Thoonen, W., & van den Bosch, A. (2006). Expertise classification: Collaborative classification vs. automatic extraction. Social Classification: Panacea or Pandora? 17th Annual SIG/CR Classification Research Workshop. Retrieved from http://www.slais.ubc.ca/users/sigcr/sigcr-06bogers.pdf.

Guy, M., & Tonkin, E. (2006). Folksonomies: Tidying up Tags? D-Lib Magazine, 12(1). Retrieved November 10, 2006, from http://www.dlib.org/dlib/january06/guy/01guy.html.

Iverson, L. (2006). Thomas Vander Wal on Folksonomy. Blog posting. Retrieved from http://www.ece.ubc.ca/~leei/weblog/2006/03/thomas_vander_wal_on_folksonom.html.

Lin, X., Joan, B., Yen, B., et al. (2006). Exploring characteristics of social classification. Social Classification: Panacea or Pandora? 17th Annual SIG/CR Classification Research Workshop. Retrieved from http://www.slais.ubc.ca/users/sigcr/sigcr-06lin.pdf.

Long, K. (2005). Leveraging folksonomy - flickr clusters at ExperienceCurve. Blog posting. Retrieved November 10, 2006, from http://blog.experiencecurve.com/archives/leveraging-folksonomy-flickr-clusters.

Marlow, C., Naaman, M., boyd, D., et al. (2006). HT06, Tagging Paper, Taxonomy, Flickr, Academic Article, . Hypertext 06. Retrieved from http://www.danah.org/papers/Hypertext2006.pdf.

Mathes, A. (2004). Folksonomies - Cooperative Classification and Communication Through Shared Metadata. Retrieved November 13, 2006, from http://www.adammathes.com/academic/computer-mediated-communication/folksonomies.html.

Sen, S., Shyong K., L., Cosley, D., et al. (2006). tagging, community, vocabulary, evolution. Proceedings of CSCW 2006. Retrieved from http://www.grouplens.org/papers/pdf/sen-cscw2006.pdf.

Tonkin, E. (2006). Searching the long tail: Hidden structure in social tagging . Social Classification: Panacea or Pandora? 17th Annual SIG/CR Classification Research Workshop. Retrieved from http://www.slais.ubc.ca/users/sigcr/sigcr-06tonkin.pdf.

Trigari, M. (2001). Multilingual thesaurus, why? . European Schoolnet, ETB project.. Retrieved November 13, 2006, from http://etb.eun.org/eun.org2/eun/en/etb/content.cfm?lang=en&ov=3813.

Vander Wal, T. (2005). Explaining and Showing Broad and Narrow Folksonomies :: Personal InfoCloud. Blog posting. Retrieved November 13, 2006, from http://www.personalinfocloud.com/2005/02/explaining_and_.html.