Wednesday, November 14, 2007

Tagging in different e-learning environments

In the last days we've had a few discussions about tagging in e-learning environments. My environment, where the tagging takes place, is a portal for learning resources.

Today I came across this nice graph that displays "power law of participation". Now, haven't looked at the scientific background of it yet, so no comments on that. Anyway, it kinda rang the bell with what I'm doing when looking into levels of user engagement on the portal.

According to this graph, adding things to favourites (e.g. bookmarking) and tagging them represents a pretty low threshold to participate in the activities of that given community.




















I'm also looking at leMill environment, which on the other hand, demands a pretty hight level of user engagement, as it is about collaborative authoring of digital learning resources. About a year down with users, there is somewhat little collaborative authoring that actually takes place, Hans told me yesterday.

Maybe tagging in some way could help the participant to take the first steps? well, they can already tag and favourite things in LeMill, so maybe the issue is rather to see if similar levels of engagement appear in that community.

So, along with that, I am interested in looking at the tags in LeMill from the same point of view that I'm doing for tags in our learning resources portal. The difference is that our case is clearly what is called broad folksonomies, whereas leMill should be a rather classical narrow folksonomy. Or is it? Maybe once we start looking at those tags as a triple {user, resource, (tags)} with a timestamp on them, it appears that participants first start by bookmarking and tagging resources from other users, before the user takes a step to create her own resources and finally collaboratively work on other's resources.

Update:

So, some data to back-up was found:





















Social Technographics®
Mapping Participation In Activities Forms The Foundation Of A Social Strategy
by Charlene Li
http://www.forrester.com/Research/Document/Excerpt/0,7211,42057,00.html
with Josh Bernoff, Remy Fiorentino, Sarah Glass

This is a document excerpt EXECUTIVE SUMMARY
Many companies approach Social Computing as a list of technologies to be deployed as needed — a blog here, a podcast there — to achieve a marketing goal. But a more coherent approach is to start with your target audience and determine what kind of relationship you want to build with them, based on what they are ready for. Forrester categorizes Social Computing behaviors into a ladder with six levels of participation; we use the term Social Technographics® to describe a population according to its participation in these levels. Brands, Web sites, and any other companies pursuing social technologies should analyze their customers' Social Technographics first and then create a social strategy based on this profile.

Wednesday, November 07, 2007

notes on "Context, (e)Learning, and Knowledge Discovery for Web User Modeling: Common Research Themes and Challenges"

"Context, (e)Learning, and Knowledge Discovery for Web User Modeling: Common Research Themes and Challenges" by B.Berendt

This paper is about context and how to define it or how it is defined differently. The following is related to "Context in Web usage mining and eLearning"

2.1 Context as data and as metadata

"In order to evaluate whether intended and actual usage coincide or not, and in order to obtain a more fine-grained picture of actual usage, it is of course interesting to measure aspects of actual usage. "

- This is also one thing that we are interested to find out in MELT, and partly also in my PhD. As we have very little access to "actual use" we try to infer this type of information from usage logs. E.g. We have a teacher who has said in his profile that he teachers students from 12 to 13 year olds. If he bookmarks LOs that have intended audience of 14-18, we can maybe infer that this LO can also be used for younger students. Especially, if we start seeing this taking place a lot, we might want to update the LOM on intended audience: instead of 14-18 we could say 12-18.

- My interest is also to see if tags can give us any hints of this.


2.2 Context and model parts

"context representations can form and/or enrich (a) user models, (b) material/environment models, or (c) interaction models."

- EUN uses a) in one search to rank resources, but we are still only implementing it and we don't know how users react to it. That is related to my own PhD, as are how different search methods are used. In general, we do way too little with user modeling (I guess bigger issues are still more imminent)


2.3 Context: parameters of the (inter)action

- For my PhD I'm looking into user logs to create "levels of user interaction", e.g. what does it mean if a user views a page vs. makes a bookmark on it. We want to use this as an input for a recommendation system, for example.

- I'm also interested in the type of search that the user has chosen and its relation to the search task that the user has at hand.

- Tags were mentioned in this context, that is also a huge part of what I am looking. There are different questions around them, one most interesting related to search is how they can be used for discovering resources.

Need to look into these papers:

- B. Berendt, G. Stumme, and A. Hotho. Usage mining for and on the semantic web. In H. Kargupta, A. Joshi, K. Sivakumar, and Y. Yesha, editors, Data Mining: Next Generation Challenges and Future Directions, pages 461–480. AAAI/MIT Press, 2004.

- Claus-Peter Klas, Hanne Albrechtsen, Norbert Fuhr, Preben Hansen, Sarantos Kapidakis, L aszl o Kov acs, Sascha Kriewel, Andr as Micsik, Christos Papatheodorou, Giannis Tsakonas, and Elin Jacob. A logging scheme for comparative digital library evaluation. In Julio Gonzalo, Costantino Thanos, M. Felisa Verdejo, and Rafael C. Carrasco, editors, ECDL, volume 4172 of Lecture Notes in Computer Science, pages 267–278. Springer, 2006.

- Totally agree with his observation, not the method: Tanimoto [53] emphasizes that may be difficult to conclude, from a mere clicking event, that there was indeed attention paid to (specific) content of the requested page.


2.4 Context: background knowledge

Tags, tags, tags. multiple views.


2.5 Context: Activity structure
" This metadatum can provide important information about a visitor’s intention or expectation (e.g., whether they followed a prescribed link from a course page, or whether they found a material by actively searching with a very detailed search phrase)."

For me this is important, I guess using terms from this paper, I'm interested in user's intentions and expectations and finding out the ways the users choose to access or discover resources in our portal. I'm also interested in seeing whether one method is more useful to a given task, e.g. if people like browsing to find inspirational material and some other method (social information retrieval vs. information retrieval) for another task. If we know what kind of method is useful for a given task, I think we can help our users a lot.

2.6 example

An example is given using the three aspects of context; activity structure, parameters of the (inter)action and background knowledge. The type of analysis allows answering questions like: which search options are popular and are there differences between users? Which content areas were frequented, and how did people navigate between then; did they go back to the search options, or did they use the inter-content links? Did certain content areas become hugs for navigation and thus served to organise the domain and the presentation of the domain? On the other hand, questions like; were there differences between users with high verbal and users with high visuo-spatial competencies; did certain textual or pictoral material become hub?

These are also questions that I am looking at in lre portal and am getting a good idea of them. However, I have not been able to link them with the task at hand yet, which is something that I'm interested in.

Schooling for Tomorrow scenarios and Science-Fiction

Last autumn I heard a few e-learning keynotes with a heavy science-fiction emphasis. That was wild, I totally loved it. I can not agree more on the idea that working with future scenarios, or science-fiction for that matter, is actually helping the future to come along. It is about helping to shape the future after having peeked in to the future with a positive or a negative outlook. And, I'd like to say that looking or creating scenarios in many perspectives is useful too, as it helps you to see whether this is where you want to end-up or not.

There's been some work going on since the beginning of the millennium regarding scenarios for the future of schooling, we also in EUN worked on that. The OECD report Schooling for Tomorrow, is out too. It's way less exiting than talking about 2048 in a science-fiction scenario, I'm afraid, but still aiming at the same goal - see how technology and network enhanced learning could be used in the days to come.

Of course I picked upon the scenario called "Learning in Networks replacing schools"













This scenario imagines the disappearance of schools per se, replaced by learning networks operating within a highly developed “network society”.

Networks based on diverse cultural, religious and community interests lead to a multitude of diverse formal, non-formal and informal learning settings, with intensive use of ICTs.

How about that for science-fiction?

Well, if it is up to me, I would like to see learning networks in schools even if schools per se are not facing the extinction. You know, before we need to go to a dinosaur museum to see a replica of a teacher.

I don't mind the idea of "networked society", but somehow, when wearing those gray classes, I'm thinking of the efficiency of terrorist cells workings, and how religious and local interest groups could manipulate their own learning interests on me, while at the same time monitoring with whom do I want to learn on the international scale. Yak!

The Schooling for Tomorrow scenarios by OECD are a real tool set for policy-makers, and why not others, to work on. I like practical things like these are. Lately, I've been toying with the idea of making people, who work with education, technology and networks, to write science-fiction short stories of learning in the future. I also think that this would be a very helpful exercise for other PhD students to open up their thinking and not be stuck with what we got now.

How to get started with your own science-fiction story on Schooling for Tomorrow

It's important to understand that it's the journey that is important, not the destination. To say it in other words, I think the whole thinking process to come up with a plot for your story is what counts! Don't worry about picking up the right publisher now ;)

This is what I found out about writing science-fiction, how to get started
  • idea: the premise, the basic thought around which the story will turn
  • a setting: worl for your story to take place in; it can be familiar, or wholly new
  • Characters: two or three, at least, to people your story
  • Aliens: (optional) strange and mysterious beings for your characters to encounter
  • Problem: something your characters want, need, must escape from, etc.
Hmm, I really like to think of Aliens and education...

Next, when you have those sorted out, think about the five dimensional framework from Schooling for Tomorrow. What are:
  • “attitudes, expectations, and political support”,
  • “goals and functions” of education systems
  • “organisations and structures”,
  • the “geo-political aspects”
  • “the teaching force”.
Then, just let it flow!

My own attempt on this is in a wiki. I've already had quite a few engaging and hilarious dinner discussions with my friends about what should the plot be. We never got very far, but it's always been a lot of fun! I only wish we could somehow get the discussion transcripts on the wiki....

Tuesday, November 06, 2007

Notes on "Collaborative tagging and Semiotic Dynamics"

By Gattuto, C., L.Vittorio and L.Pietronero (2006).

Firstly, I must say that I was glad to read this paper. Lately, I've been seeing many papers talking about the properties of folksonomies, like co-occurrence, etc., which have intrigued me quite a lot. This paper explains the process pretty well and underlines an important point - they factor out the users and only deal with streams of tagging events and their statistical properties!

I must admit that this makes the whole area of Semiotic Dynamics less attractive to me. I think it is important to study tags and their properties, but not in isolation from the user. I see (barely) the point to explain tagging activity and the growth of tags in separation from the users. But fair enough.

Problem statement: Uncovering the mechanisms governing the emergence of shared categorisatioins or vocabularies in absence of global coordination is a key problem with significant scientific and technological potential. Collaborative tagging provides a precious opportunity to both analyze the emergence of shared conventions and inspire the design of large agent systems.


Semiotic Dynamics study how populations of humans or agents can establish and share semiotic systems, typically driven by their use in communication. The author argue that the emergence of a folksonomy exhibits dynamical aspects also observed in human languages, such as the crystallisation of naming conventions, competition between terms, takeovers by neologisms, and more.

  • Users interact with a collaborative tagging system by using tags or adding new resources to system
  • Basic unit of information in collaborative tagging systems is a (user, resources, {tags}) triple, which they refer as post in this paper. Tagging event is a tri-partite graph (with partitions corresponding to users, resources and tags, respectively) and can be used as a navigation aid in browsing tagged information
    • Comment: I like the tri-partite graph as navigation aid, yes!, but as the authors mention just above, they don't think of other users and those networks as navigational aid. In contrary, they omit the users just to study the properties, which strikes bizzarre to me.
The authors cite the "rich get richer" model (Yule-Simon's stochastic model) and propose to enhance it with a "fat-tailed memory kernel". This original model is related to the construction of text from scratch:
At each discrete time step one word is appended to the text: with probability p the appended work is a new workd, never occurred before, while with probability 1-p one work is copied from the existing text, choosing it with a proability proportional to its current frequency of occurrence. This simple process ields frequency-rank distribution that display a power öaw tail with exponent alpha = 1-p, lower than the exponent we observe in actual data. This happends because the Yule-Simon process has no notion of "aging", i.e., all positions within the text are regarded as identical ..
This all leads to a model of users' behaviour: the process by which users of a collaborative tagging system associate tags to resources can be regarded as the construction of a "text", build one step at a time by adding "words" (tags) to a text initially comprised of n 0 words. There is also that same Yule-Simon model with long-term memory (about inventing new tags or using existing ones), but recent tags are used more often than old ones.

Also, "in our model,.., the average user is exposed to a few roughly equivalent top-ranked tags and is translated to mathematically into a low -rank cutoff of the power law, i..e., the observed low-rank flattening".

Conclusion: It seems that users of collaborative tagging system share universal behaviour which, despite the intricacies of personal categorisation, tagging procedures and user interactions, appear to follow simple activity pattern.

There is also something about the co-occurrence between high-rank and low-rank tags: it says: "This suggest that high-frequency tags partition - or "categorize" - the resources marked by tags of lower frequency. "
Comment: This all sounds interesting and important, but will need to look into that later.


Monday, November 05, 2007

My PhD research

Notes on "Aspects on Broad Folksonomies"

Aspects on Broad Folksonomies by M.Lux and M. Granizer (2007)

This paper continues the trend in studying and analysing the underlying statistical properties of broad folksonomies that aims to identify laws and characteristics which allow inferring those properties. A few notes on what I found interesting related to the emerging notion of quality of tags, something that I've also spared a few thoughts on.

First, though, on some other issues. The paper talks about the emergence of power law distribution in folksonomies. They describe which approach they took to fit the sample to a power law, which was something that I've sometimes contemplated on the how-part of things. The paper aims at analysing whether one can find similar term distribution in folksonomies as in classical term retrieval (e.g. Zipf. note: Zipf's law with an exponent between 1 and 2). The dataset is that of delicious (uh, with about 800 000 bookmarks and about 27 000 users- I got a way to go with my MELT bookmarks).

Tag co-occurrence
They are able to show that "for around 80% of the tags of a folksonomy the co-occurring tags follow a power law distribution, which approves Cattuto's assumption. We found that for about 90% of the estimated power law exponent B xxx [-1.5, -0.5], which shows that for most tags co-occurrence follows a model with similar parameters. "
Resource and user based tagging characteristics
Secondly, they looked into frequently used tags (more than 30 users).
  • For resources statistics they (frequency of users tagging the resource with a tag) found that around 18,4% of resources followed a power law distribution.
    • assigned by lot of users to few resources (head) and to a lot of different resources by a few users (tail)
  • For user statistics (frequency of resources tagged with a tag), around 13% are following a power law.
    • few users tag a lot, whereas lot of users tag a few
  • i.e. the characteristics of the user statics are similar to the characteristics of the resource statics.
  • They argue that those tags, which follow a power law w.r.t users and resources are high quality tags (i.e. tags describing resources with high accuracy [no misspellings and meaningful tags] ) for most of the users involved in the investigated social bookmarking system.
  • A small fraction of tags have overlapping user groups, which points towards sub communities (user groups sharing the same link selection and tagging behavoiur) in the tail of the power law distribution.
    • this was found through splitting resources in 3 (high, mid and low rank resources)
They also looked at the big chunk of tags that were not following the power law.
  • Unique assignments. More than half (57%) of less frequently tags are used only once. They think that they can be seen as "shortcuts" for a user to a resource or a misspellings. They argue that these tags are useless from retrieval point of view (hmm..).
  • Personal vocabulary. especially in less frequently used tags (19%) of tags were only used by one user but assigned to many resources. They are useful for personal retrieval but useless for the rest of the community.
  • Unpopular vocabularies. between 1/5 and 2/5 of tags are assigned to different resources by different users only once. Unpopular vocs used by a small fraction of users.
  • they conclude that from retrieval point of view (e.g. inverted indices, TF*IDF) a large fraction of tags are good for single or sub-communities, and only the power law distributed tags are good for that.
    • They don't say anything about how to include the large fraction of tag not distributed by power law into IR methods.
Retrieval Aspects
Q: Do tags add information to further to description and title for retrieval purposes? This is a lot along the lines that I am also interested in, although I will look more into the networks of users. They say that for retrieval tags can be seen as an additional resource. Moreover, about 50% of available description contain information similar to the information described by tags, whereas the remaining 50% can be seen as orthogonal information.

Comment. This all is treating tags only as additional keywords that can be useful for conventional retrieval purposes. I think the connection tag-resource-user is more interesting. Just the fact that even if the tag is misspelled or hooks to a small user community is less important to me, because I know that the fact that this resource was tagged shows that the user has an interest to this resources, thus it is a vote. This aspect has an immense potential for retrieval (recommender point of view), but is seldom regarded in papers with very conventional retrieval approach.

Open social and education

I wonder who is going to come up with the first OpenSocial app or widget for educational use? We certainly are talking about it, for example for our eTwinning platform. It could be cool to be able to use information about teachers collaborative networks to allow, say, better retrieval of learning resources relevant for the project, purpose or task that teachers are undertaking; link with some other sources that teachers are working on through cool widgets, etc.

I never thought that Facebook, which has lately become really popular among my friends (not early adapters), would be the seul app that would "take it all". I was glad to read this:
"The market has already decided that there's going to be a long tail of social networks, and that people are going to belong to more than one. As soon as you belong to more than one, this kind of interoperability is critical," Dash says. "Open standards win every time." wired

Hurray for open standards!

Tuesday, October 30, 2007

Multilingual tags and the language of LO

I've looked into tagging in different languages before. An interesting thing came out of our little pilot: teachers, non of whom mother tongue was English, still had about 20-30% of tags in English. We had two different thoughts on this,
  • either tag is in English because teacher wanted to share these tags with other teachers, or
  • tag was in English because it is related to the language of the learning resources that was bookmarked
I was now interested in the second possibility, and took a look at a sample of 136 bookmarks with tags in multiple languages related to them.
  • The LOs were in English, Hungarian, Polish and Estonian.
  • The users (43) were Hungarian, Polish, Estonian and Lithuanian
Fair enough, all the English tags were related to the English resources! In close to 30% bookmarks (39 out of 136) this was the case (which also means that 30% of LOs were in English).

Moreover, it seems that for about 1/3 of the times the language of the LO was the same as that of the tag, whereas 2/3 of the cases it varies according to the language of the user. In about 3% of bookmarks one was able to observe multi-lingual tags.

Thursday, October 25, 2007

Radioheads making €, good for them

Cutting the record lable out of the equation seems to be a good deal. I'm glad to see the bold decision to go directly to fans has not only shown a great example, but also proved profitable! I've always liked the idea of Magnatune.com, although never bought anything..
En trois jours, Radiohead avait vendu 1,3 millions d’albums. Le prix moyen aurait été de 6 euros - un chiffre qui semble être tombé à 4 euros après que les premiers fans aient passé commande. Avec l’élargissement de l’audience a un plus grand public, la moyenne du prix d’achat s’est tassé : on estime entre 1/4 et 1/3 le nombre d’internautes qui auraient choisis de ne rien débourser. Wired estime néanmoins que le groupe aurait déjà pu récolter entre 4 et 8 millions de d’euros.

Même avec une moyenne basse de près de 3 euros par album vendu souligne Guillaume Champeau sur Ratiatum, c’est près de 4 millions d’euros que le groupe aurait gagné en quelques jours. A une dizaine de pourcent de rémunération par album, “dans les circuits classiques, Radiohead aurait du vendre 2,5 millions d’albums pour gagner l’équivalent”. Internet actu

Friday, October 19, 2007

what every PhD should know: dinner discussion with a google guy

Just barely hanging out there. Today was lots of serious fun and intellectual challenges at the RecSys 2007 Doctoral Consortium . Interestingly, all the participants came from lots of different backgrounds from computer science, information retrieval to me from education. I think the diversity of backgrounds and focuses of studies represent the growth of the Field of Recommenders, it's not only about the best algorithm anymore, but a plethora of questions around.

Anyway, being somewhat a newbie here (yeah, I do not know all the people or study areas here, very eye opening!), it makes me think that all the PhD students should be exposed to the question " If you were to have dinner next to a main researcher in Google/Yahoo/or any other big name, what would you want to talk about?".

Well, as it happens to be, I never thought of that before. Neither was I prepped for that by my study programme. Nevertheless, I just spent my dinner next to Krishna Bharat, you know, the guy who greated Google News, nothing less, nothing more. Probably tomorrow I'll have like ten things I want to ask from him with no chance to get his attention anymore.

Bottom line: it is not only about the 1 minute elevator pitch, but about the life and such in general.

Tuesday, October 16, 2007

New acquitance: Semiotic Dynamics

Pretty exiting, I came across this new area of Semiotic Dynamics, which is described as "a new field that studies how semiotic relations can originate, spread, and evolve over time in populations, by combining recent advances in linguistics and cognitive science with methodological and theoretical tools from complex systems and computer science." One topic of this study field is folksonomies, which draw my attention. The stuff can look like this.

Everyone nowadays repeat the same mantra of web 2.0, but somehow this project managed to say things sets it apart:

..users are no longer limited to consuming or creating online content, they also provide the semantic scaffolding holding together such content, thus taking on an active role in shaping the architecture of online information. The collaborative character underlying many Web 2.0 applications puts them in the spotlight of complex systems science,..

"Semantic scaffolding holding together .. content", that's a pretty awesome way to put it!

The paper "Vocabulary growth in collaborative tagging systems" investigates the temporal evolution of a tagging vocabulary size (of delicious) both on a
  • global level (the number of distinct tags in the entire system) and
  • local level (the growth of the number of distinct tags used in the context of a given resource or user).
It asks questions like how does the number of tags grow?; what is the rate of invention of new tags? is the asymptotic number of tags finite (uugh, a nice way to say it)? etc...

The paper finds out that the growth behaviours are remarkably regular throughout the entire history of the system with power-law behaviours with exponent smaller than one (non of that "fat head and long thing tail"!) and across very different resources being bookmarked.

Moreover, they find that there are some intrinsic characteristics of the system which do not depend strongly on the size of the dataset, like that the average number of tags is about 3.4 (local level). If I get it all right, they conclude on this that on the local scale (resource or user) "all curves tend to lie along a "universal" growth curve with an exponent close to 2/3".

The authors of this paper also highlight that the tools and concepts from complex system science may prove valuable for understanding the structure and dynamics of folksonomies.

Some interesting papers towards this direction: http://www.furl.net/members/vuorikari/semiotic_dynamics

Wednesday, October 10, 2007

Google goes micro-blogging

All the roads lead to ...Google. I guess we could re-phrase the old saying. Just received a notification from Jaiku that they are joining Google. In the other words, Google acquired them. Good for those guys, I hope. I wonder how many Finnish SMEs have become part of Google in the past?

I kinda enjoy Jaiku even if I don't micro-blog from my phone. I like it as an aggregator of feeds and to check what my pals are doing. Unfortunately the Facebook app. does not work that well, but hey, maybe Google will fix this one?

Saturday, September 29, 2007

Notes on Smart Indicators on Learning Interactions

Smart Indicators on Learning Interactions by Clahn et al. (2007) discusses how indicators can be used to help learners, or groups of learners, to organise, orientate and navigate through learning environments by providing contextual information that is relevant for performing learning tasks. Indicators are part of the interaction between a learner and a system (social or technical).

Indicator system is defined as a system that informs a user on a status, on past activities or on events that have occurred in a context; and helps the user to orientate, orgaise or navigate in that context without recommending specific actions.

So, it is not:
  • a feedback system (analyse user interactions to inform learners on thier performance on a task and to guide the learners though it) or
  • a recommender system (analyses interactions in order to recommend suitable follow-up activities),
  • instead it provides information about past actions or the current state of the learning process.
  • Moreover, smart indicator systems adapt their approach of information aggregation and indication according to a learner's situation and context.

The paper draws heavily on the notion of social navigation, interaction history and footprints, and offers a good review of this literature (ToRead).

The paper offers an architecture of smart indicators, where different layers are defined to support user modeling (first two) and helping the system to adapt to better decision making process (last two). Four layers:
  • sensor layer
  • semantic layer
  • control layer, where a strategy defines the conditions according to learner's context
  • indicator layer, presents aggregated information to the learner.

This approach of smart indicators adapts the strategies on the control layer (as opposed to semantic layer) to meet the changing needs of a learner.

SENSOR AND SEMANTIC LAYER

The paper further presents the information aggregates of sensor and semantic layers. The idea is to classify and organise the user's engagement (interaction foot prints) with the system, e.g. contributions, tagging activities. In the sensor layer, there is a division between "learner interaction" and "contextual sensors", e.g. location tracker, tagging activities (in my case this is considered direct) and contributions of peer-learners.

I am doing the same with my research data, and I call it the "user engagement" following the Yahoo!'s idea on STAR-metadata (kind of attentional and explicit metadata about users actions).

I tried to apply the classes of Chlan's prototype to my research data (learning repositories) that I collect using our CAM framework. Our focus being somewhat different, it did not really work out that well. The attempt below, though:

Direct: accessing resources through browsing, tag cloud, search result list, other user's favourites (implicit interest)
  • user views metadata
  • user views tags
  • user views resource ("entry selection sensor")
  • (timestamp on everything)

Direct: higher level interaction with a resource (explicit interest)
  • user adds a resource to favourites and tags it ("entry contribution sensor", "tag selection sensor", "tagging sensor" or"tag tracing sensor", hard to say in my case)
  • user rates the resource ("entry contribution sensor")
  • user comments on the resource ("entry contribution sensor")
  • shares resource with network ("entry contribution sensor")
  • (timestamp on everything)

Contextual sensors could be (here I'm blending them with user information):
  • context of a project within which the user access resources
  • the information about the country and school from where the user is from

SEMANTIC LAYER

The semantic layer users the information from Sensor layer and transforms it into meaningful information by using an "activity aggregator". This calculates the activity for a given period of time for an individual learner or the whole community according to different ratings that each activity has (beginners have different way of counting activity from power-users).

CONTROL LAYER

In this prototype the control layer defines how the indicators adapt to the learner behaviour. There are two elemental strategies:
  • motivate learners to participate to the community activity
  • raise awareness on the personal interest profile and stimulate reflection on the learning process
Moreover, a third level control strategy uses the activity aggregator as well as the interest aggregator.

INDICATOR LAYER

This layer embeds the indicators into the user interface of the community system. The prototype is being tested by a group of PhD students now.

Glahn, Christian, Specht, Marcus, Koper, Rob (2007) Smart Indicators on Learning Interactions
http://hdl.handle.net/1820/941

Thursday, September 27, 2007

Some thoughts after SIRTEL07

Last week the SIRTEL workshop took place. The papers are found here and the slides, well, most of them, at the EC-TEL07 conference wiki. I have pretty good feeling about the workshop, it was one day long, we had about 20 people participating, some of whom chose to stay with us for the whole time, and some who were hopping between workshops. For me that is totally fine, we all are responsible for our own learning! Especially in conferences where many parallel sessions are running, I would encourage people to try to get best out of them.

For those who could not make it at all, you can soon find recording on the SIRTEL site.

The workshop had four main sessions:
  • We started with a keynote address from people who work with music recommenders. MyStrands people talked about applying social recommender systems to technology enhanced learning. It was an interesting talk that challenged all of us to think what are recommenders for learning purposes in the first place (goal) and what kind of data do we want to use to do that.

    As any good keynote, this one gave more ideas to think than answers. It nicely set the base for the further discussions during the workshop that focused on the need to define the field of Social Information Retrieval for Technology Enhanced Learning, and to establish a baseline so that we know what are we really set to do.

  • The second session was about Tagging and Visualisation. We had my presentation about the user behaviour on tagging in multiple languages; then there was a presentation from COSL that talked about "Activities of Daily Living on the Web", Brandon also showed a few demos of the widgets that can be used to rate or recommend related content. That was followed by a talk on reward structures to encourage teachers to share open educational material. Finally, we listened about Visualisation of social bookmarks, a work that leads into visualising bookmarks in an educational repository.

  • The 3rd session was on Recommender Systems. Here we first heard about some R&D work that OU NL is carrying out using the idea of learning paths to better support learning activities of students. Then, there was a study about using affiliation networks as a mechanism for collaborative filtering (understood largely). This was followed by a study on simulating recommendations based on multi-attribute ratings on learning resources by teachers. Finally, we had a system demo of Daffodil that supports collaborative information seeking.

  • The final session was what we called "Enablers and Challenges". It was a discussion session, and as we advanced, it was clear that people had a lot to say. It might even have been better to allow more time for this, but hey, you live and you learn.
I try to sum-up, but basically it illustrates the main topics that we talked about. If you look at the left side, there are the fundamental questions:
  • How to define and chart out the area of Social Information Retrieval (SIR) for learning?
  • Is this application domain different from other SIR, on micro and macro level?
  • What do we recommend?
  • In what context?
  • and based on what?
On the right hand, there are the issues related to implementation and evaluation of it. These are:
  • What are the best SIR methods for TEL?
  • And what is the data that we should use? The "data issue" was something that was heavily emphasised by the MyStrands folks, who obviously speak of experience.
  • The questions rouse also: when do we start implementing these for real or are we just over-engineering and never ready to launch?
  • Evaluation and empirical data for real evidences was on the focus a lot.











More will follow. This is quick and dirty now, hopefully I will get more input from people participating in order to get more depth on our summary.

Sunday, September 16, 2007

SIRTEL'07: la raison d'etre

"We use people to find content. We use content to find people."*

On Sept 18 our SIRTEL workshop takes place. It's gonna be "Serious Fun"! Let me just outline why:

SIRTEL'07: Raison d'etre

Recommender systems, as well as social navigation, have been around since the popularisation of WWW, that's some 15-20 years now. The idea is to help people choose the right stuff from a potentially overwhelming set of choices. To facilitate that users could be helped with information from other users, the choices made before (by themselves or similar users), the ratings or reviews other people had done, etc. (Rescnik et al., 1997)

The field of learning technologies has seen recommenders of some sort being discussed and prototyped since the late nineteens. In the review of the field in Manouselis et al (2008) we identified about 10 recommenders, and even more conceptual papers of them, but very little has matierialised so far.

Since the last few years recommenders have made a second arrival into the discussion topics of technology, or network, enhanced learning. Undoubtedly, this has been influenced by the arrival "Web 2.0" with all its ideas:

- Collaborative tagging, for example, has changed lots of ideas of how metadata should be produced and how static a metadata record should be: it's not anymore one metadata record produced by a librarian, but lots of annotational and attentional metadata by lots of users.

- Other annotations by users that express their subjective judgements have seen a huge growth too, we don't only talk about ratings or reviews in their traditional sense, but also tumbs-up or down, giving pokes to people or objects, etc.

- Social bookmarking, which allows users to create easy references to their own collections of digital resources (photos, books, links, music,..), has given a new dimension to the concept of social navigations. The link between resource-user(-tag) allows users to navigate other people's collections and thus find novel resources. Also, the same resource-user-tag link gives researchers an itch to use this information to group similar users for recommendation purposes, as well as to study the emerging networks.

- Expressing social ties between people has also brought new possibilities along. We are not only seeing networks of friends, but there are new possibilities where people can express different networks, ones for professional use, others for personal, recreational, etc purposes. Also, portability of these networks has become an issue discussed for better designs (social-network-portability group, PeopleWeb ,..).

- Something else is also happening behind the scenes. Clicksteam and user behaviour on the Web is not anymore a property of the commercial portal on which users are, but users are starting to take seriously how their "attention" is being used, who owns it, etc. Attentional metadata is a huge source of information that educationalists are also starting to take more seriously and thinking how it could be used for better serving learners and teachers (Contextual Attention Matadata, Attention Profiling Mark-up Language, Attention Trust,..). Attentional metadata can also become crucial when it comes to better understanding the intentions of a user, why are they, for example, looking for some information and for what task at hand!

- Finally, content for educational use, or rather its production, is also seeing a change. Users generate more and more of the content on the Web in general, a trend which is also seen in the e-learning. Of course, traditionally teachers have always produced lots of their own material, but now its re-use also has been facilitated (e.g. repositories/referatories). Also, the collaboration aspect is facilitated by the Web, it has become easier for people to work together on things (e.g. wikis, collaborative platforms,..). Additionally, learners produce plenty of material which also should be seen and used as educational content.

To sum-up: two main topics evolve around social context and social content. Social context is how we express the who, where and with whom, and social content are the objects or digital artefacts that are in the center of the communication, exchange and networks.

All the above has hopefully also changed how we will see the future of social information retrieval for technology enhanced learning. This workshop will all be about that! Serious Fun!

-------

N. Manouselis, R. Vuorikari, F. Van Assche, “Collaborative Filtering of Learning Objects for Online Communities: An Experimental Investigation”, accepted for publication in Computers in Human Behavior, Special Issue on ‘Advances of Knowledge Management and Semantic Web for Social Networks’, 2008.

P.Morville, 2004

Resnick P. & Varian H.R., “Recommender Systems”, Communications of the ACM, 40(3),1997

Wednesday, September 12, 2007

How do teachers network?

Just wondering...

I had an awsome chance to join this teachers' community.

View my profile on eTwinning Reunion

Thursday, August 30, 2007

More thoughts on multilinguality and tags

Lately I've been thinking more about tags and how the fact that users use them in different languages effect on the tagging system. In the case where I work (EU+education) we want to use tags in different languages as something that unifies people rather than divides them in different sections. That is why in our system we are NOT thinking of keeping multilingual tags and other annotations separated.

This type of separation along language and/or national lines can be seen in quite a few places on the Web. For example in Amazon, the reviews and ratings are not shared between the .com and .fr version. I understand that reviews and ratings, especially from experts and authoritative reviewers, are something really culturally biased, but I would think that it is interesting for readers in the US to know how a book has been received in France.

Lately in our team we've also talked about translating tags. Like if my tags, which I have added in English would be translated by someone, or a machine, into Finnish, French, etc. I don't like that idea. If I translate my tags, for example I add a tag in English and in Finnish, it's fine. But if it is done for me, I don't think it Ok. Let me explain:

I prefer to display to users only USER generated tags, not translations (neither people done nor automated ones). The key thing with tags, and the big difference compared to normal vocabularies, is that they not only describe the resource, but are associated to a user. This relation of users, tags and resources is fundamental!

In the scenario where tags are translated the following questions arises: to which user do you link a translated tag? In my opinion (say, if pushed to an edge), if a tag is translated, it ceases to be a tag and becomes just a mere keyword.

The above does not mean that tags could not have translations or equivalent terms in other languages or in its own language. Most likely in our system we will see lots of both. But the difference is that those tags all are created by other users and can be associated to users and resources. Translated keywords can be only associated to resources. This connection of users, tags and resources becomes our main asset for connecting people across the national and linguistic borders. Let's keep it that way!

If tags are translated, they could be used for other purposes than for displaying (I mean tag clouds, social navigation, etc). Translated tags could, for example, be helpful as keywords to make the search better. If a tag is translated, it should also be indicated in the metadata of the tags.

Monday, August 06, 2007

Inspirational wine tasting and the magical web2.0

This is SO funny, I was always wondering how wine-tasters come up with those weird and sometimes hallucinative descriptions of wine. Now I know, thanks to Wine library tv!



During our 3,5 weeks sailing trip we discussed (and yep, also drank) quite a lot about wine and web 2.0. Dan, the co-skipper, got really inspired about recommenders and the web 2.0 stuff. Having been on the boat for the last 3 months in a row, he'd of course missed a few things going on. So, we had enough time to catch up on lots of things and also to make a concept for a wine-soso-small-producer-mash-up.

On my return to the land of Internets, I took a few hours to see what has already been done on wine and web 2.o. Am I right, or far off, if I say that wine inspires lots of web apps (well, I guess in many sense..), maybe just up on the top there with all the sex business?

Certainly seems like that. Anyway, a cool, well-needed API was released by Wine.com related to their cataloque of wines. I think it's cool that they syndicate the information about their cataloque, but more importantly, how they syndicate all the metadata about the wine, too. I must say, though, that I agree about the comments related to naming all the elements category, instead of giving clear names like year, country, producer. Now, that would be re-usable!

Saturday, August 04, 2007

Draft paper: Analysis of User Behavior on Multilingual tagging of learning resources

This is an almost final draft of a paper that I'm currently working on. It's been accepted as a full paper to the SIRTEL workshop.

Ah, should be mentioned, maybe, that I'm also co-chairing it :) It's gonna be very cool, so try to make it there, if possible.

If not, you can always think of posting a question in YouTube, like they did in the US presidential campaign. I kind of like that, although I don't think that I get CNN to co-host it!

Anyway, comments are welcome on this on. All images are missing, I was testing Google docs for this-copy and paste from OO did not include images.

Also, the formating took some damage, sorry about that. The final, more readable version will be at the conference site in about 10 days.

Analysis of User Behavior on Multilingual Tagging of Learning resources
Riina Vuorikari1,, Xavier Ochoa2, and Erik Duval1


Abstract. Although social, collaborative classification through tagging has been the focus of recent research, the effect of multilingual tags is often overlooked. This work presents an early exploratory study of the production and consumption of multilingual tags in a European educational K-12 context. The data, produced by teachers bookmarking and tagging learning resources during three month period, was analysed. Thereafter, this information was presented in the form of metadata keywords to a focus group of teachers who evaluated its descriptiveness, usefulness and overall quality. The results of this early study suggest that users are divided about the benefits of multilingual tags, however, some tags are useful for some users, thus “hiding all but the right tags” becomes crucial for the success of a multilingual collaborative tagging system.

Keywords: Collaborative tagging, multilinguality, learning resources.

1 Introduction
The use of social, collaborative classification systems has gone through a continuous growth in the latest years [1]. An example of this is a multitude of sites that provide some type of social annotation of digital artefacts and a social navigation system (Flikr, del.icio.us , CiteULike, Last.fm, among others). Social tagging, i.e. allowing individuals to apply free text keywords to digital objects, potentially offers advantages in terms of personal knowledge management, serendipitous access to objects through tags, and enhanced possibilities to share content with emerging social networks.

Several studies have been undertaken to better understand the behaviour and evolution of social tagging systems. Early research has been conducted by Mathes [2] where the term “folksonomy” is used to compare the emerging socially generated vocabulary with the more formal ontology concept. Golder and Huberman [3] first looked at user patterns of collaborative tagging systems. Recent studies focus on the navigability of such social systems [4] and on understanding the network properties [5].

A prevailing aspect among current studies concerning tagging is that they assume that tags are represented in a common language [6], understandable by all the members of the user community. Guy suggests that it is not always the case [7], but does not offer insight on how to deal with tags in multiple languages.

Lately, multilingual tags have started emerging on popular social tagging systems as their user-base grows, and different ways to deal with multiple languages can be observed. Delicious users, for example, add tags in different languages for a bookmark (e.g. achat, shopping) and even in some occasions add language identification in tags (e.g. lang:fi) for the language of the resource. However, it does not offer any system level support, that allows users to see tags, say, only in French or Finnish. Other services, like Yahoo!'s MyWeb on the other hand, offer tags and tag clouds in different languages in their localised parts of the portal (e.g. .fr, .es, ...), thus some language identification of tags takes place on the system level. Thirdly, In LibraryThing experienced users can combine tags, where in some occasions tags in different languages have been grouped together.

Our work, still at its early stage, attempts to shed light on a community of users who shares a common educational interest to use a social tagging system across country and language borders, but does not necessarily share a common language, as the users are free to choose the language(s) in which they apply tags. This exploration takes place in the context of two European Community founded projects, CALIBRATE1 and MELT2, both focusing on sharing and re-using of digital learning resources for K-12.

European education, especially that of K-12 education, is inherently multilingual and multicultural. Offering educational resources and services in native languages is deemed important, but equally important is the exposure to other languages. One way to promote this is to make learning resources available across national and linguistic boarders. This puts constraints on semantic interoperability, i.e. how well content and its metadata can be understood by other systems and users.

Controlled vocabularies, such as multilingual LRE Thesaurus3, can be used to overcome some hurdles of semantic interoperability. However, the gap between the terms used by experts and practitioners in the field is also problematic. For that reason, the current research looks into co-existence of taxonomies and end-user generated tags.

A federation of learning resource repositories in a multilingual context needs to support multiple languages at the system level in order to support each repository and its national user-base, but at the same time, there is a need to allow people (i.e. user information and preferences), resources and tags to “travel” across national and linguistic borders.
This paper is structured as follows: first, in section 2, we analyse the early stage of the bookmarking and tagging behavior of our community in order to better understand how teachers bookmark and tag resources in a multilingual context; what types of tags are provided and in which languages. Then, in the section 3, an experiment is presented that measures the effect of multilingual tags on the descriptiveness, usefulness and overall quality of the metadata. Finally, the findings are discussed and applied to design decisions for multilingual tagging systems.

2 Analysis of tagging behaviour in multilingual context

The CALIBRATE project makes K-12 digital learning resources available to its pilot schools (78 schools ) in Hungary, Austria, Estonia, Czech Republic, Lithuania and Poland in their different curriculum areas. Schools can access material in different languages through a portal that is connected to a federation of learning resource repositories [8] in the pilot countries.
As part of the project's multilingual search interface4, a personal bookmarking and tagging tool has been available since the beginning of 2007. This tool allows a user to create personal collections of learning resources by bookmarking interesting resources found through the portal. To facilitate the management of these personal collections (also called favourites in the project), the user can also add keywords to resources to make it easier to ”keep found things found”. These keywords are free for the user to choose and can be expressed in any language. The collections and keywords are kept private to the user, and at this stage of the experiment, they cannot be shared among users.

The data for this analysis is from a period of about three months (January 24 to April 21 2007). There were 77 teachers who made 459 bookmarks with 417 multilingual tags on 320 different learning resources. It is intended to have regular analysis of this data within the projects lifespan (-2008).

2.1 Quasi-Experimental Set-up

A total of 173 subjects used the portal during the time of the experiment, however, the subjects of this dataset comprises of a group of 77 teachers who had done at least one bookmark during this time. Thus, it was a self-selected group formed based on the bookmarking behavior during the period of three months and it represents 45% of all the pilot participants. As there was no overall methodology to introduce bookmarking and tagging to the subjects, more than half of the participants had not shown interest in using this feature of the portal.

The bookmarking habits, at this very early stage, varied a lot in terms of what languages to use, how many tags to add, how to add multiple tags (with comma separated or without commas), etc. Also, hardly any of the participants had previous experience on tagging, so not one single tagging convention emerged, rather many different ways to use tags in multiple languages. As there is very little research done on the multilingual context, we think it is important to study the early stage of tagging behaviour to better anticipate the effect of multilinguality on the system to improve its design.

It is noteworthy to mention that the bookmarking and tagging system, at this stage of the pilot, offers very little social influence in what comes to choosing what to bookmark and what keywords to choose. Oftentimes in social bookmarking sites, social cues are made available (e.g. most bookmarked items, tag clouds, tags are recommended based on previous tags, etc). At the time of the experiment, the system had hardly any tags attached to resources, so teachers started from an empty plate. In the case where a resource was already tagged by another participant, the user would see the term(s) only if they were in the same language as the interface is.

2.2 Results

In the part we present the results of the analysis, which will be discussed further in conjunction with the other results in the discussion section.

When we look at the distribution of bookmarks per users, we can find that on the average, each user had 6 bookmarks (Fig.1). However, the distribution was very wide; 10% of the users had more than the average amount of bookmarks, which leaves 90% under the average. Eight of the users could be called “super users”, as they had more than 20 bookmarks, and 12 users had between 20 and 6 bookmarks. About 30% of the users seem to have only experimented with the bookmarking system, as they only have one single bookmarked item in their favorites folder.


We had recorded 418 tags in the system. During the semantic analysis of tags we found that many tags actually contained multiple terms, i.e. they were bundles of terms without comma separation. This was due to a technical feature of the tool that treated terms without comma separation as one tag. When broken down, they resulted in 585 terms. They were translated into English and a semantic analysis was performed to better understand the types of tags. We used the classification from Sen [9] that is also based on the categories of Golder et al. [3], which are Factual tags (Golder: item topics, kinds of item, category refinements); Subjective tags (Golder: item qualities) and Personal tags (Golder: item ownership, self-reference, tasks organisation)

The vast majority of the tags at this early stage (Table 1) are of the factual type. From the factual tags, 79% were put into a rough category of topic and 14% of the category refinement with richer information. The rest of the tags were subjective in their nature and could be used to describe the quality of the resources or how the person felt about them. None of the tags fell into the category of personal tags as Golder describes them (e.g. tags related to item ownership, self-reference or personal tasks organisation). When we analysed how these tags were used and re-used among users, we found that 80% of tags related to bookmarks were factual and 20% of tags subjective tags. In a MovieLens study [9], for comparison, the distribution was 63% factual, 29% subjective, 3% personal and 5% other.
Table 1. Types analysis of each tags (no re-use)
Factual
340
93%
Topic
Category refinement
288
52
79%
14%
Subjective
24
7%
Personal
0
0%

After categorising the tags, we further studied their nature. Two main trends seemed to emerge, first, many of the tags contained the same terms as in the title, i.e. user had just copied the title in the tag field. Second, about 13% of tags contain a general term, a name, place, e.g. EU, Euroopa, Euroopa, Europa, europe, geograafia, Phytagoras, etc . We hypothesise that this type of “travel well” tags, even if not translated, could be found useful for other users for their close similarity in spelling in many languages. We think it could be of interest to work towards automatically filter this type of terms from the pool of all multilingual tags, for example, by matching them against existing multilingual vocabulary lists available on the Internet.
When we looked at the number of tags that users related to bookmarks, we were able to identify some early trends. For the total of 459 bookmarked resources, we found that some of the tags were re-used, there was an average of 1.92 tags/resource. More than half (56%) of the tags were entered as a bundle of terms, i.e. most teachers had added 2 to 6 terms without a comma separation. In quite a few cases these terms were comprised of the terms in the title of the resource (Fig.2). In 28% of the cases only one term was entered as one tag.

The rest had used multiple separate tags (2-6 tags). In the latter case the terms were not necessary related to the title alone, but carried other types of information (e.g. title: Umweltkids and tags: Oekologie, Artenschutz, Regenwald, Tierschutz, Skisport).
Contrary to our expectations, the users took liberties to add tags in multiple languages and to use the portal interface in different languages than that of their mother tongue (interface was made available in the languages of the pilot and in English). This made the identification of the language of tags more difficult, as we had expected to be able to identify the language of the tag from the language of the interface that the user used when inserting the tag. In about 70% of the cases we were able to identify the language of the tag correctly using this method, which leaves us with a 30% error rate on language identification. This error in identifying the language of the tag correctly would make it hard, for example, to display tags and tag clouds in one single language, an issue that is related to the usability of the portal, and the one of which the second experience was set up to find more evidence.
We found the following scenarios for tagging, however, due to our logging, we can't give percentages for these use cases:
  • Interface and tags in mother tongue
  • Interface was used in mother tongue, but tags in other language
  • Interface was used in a language that is other than the mother tongue, but tags were entered in mother tongue
  • The tagging language was other than the interface language and the mother tongue
These scenarios were found through comparing the real language of the tags to that of identified language by using the interface language. In this early stage of the experiment it is impossible to draw firm conclusions, but it seems that users are likely to use, or at least try, the interface in different languages. We found, for example, that the tags entered through the English interface were in English only in 50% of the cases, which means that users added tags in languages within their areas of competences. On the other hand, we also found that there were many more tags in English than we expected from the choice of the interface language. These users had chosen to tag in English, even if they used the interface in some other language, most likely to be able to share tags with users from other countries.

3 Experiment with Multilingual Tags

An exploratory experiment was set up in order to measure the perceived usefulness and quality [10] of multilingual tags, traditional metadata and expert classification keywords. We were also interested in how users reacted when they were confronted with tags in multiple languages that they did not have knowledge of. The experiment subjects were shown a list of learning resources metadata with keywords in multiple languages, the list was imitating the search result list of the portal. The results of this experiment will be useful to guide design decision in the development of retrieval tools for learning objects in a multilingual environment.

3.1 Experimental Set-up

Thirteen teachers, who belong to the MELT focus group, were selected to participate in the experiment. They were confronted with metadata regarding five learning resources in different areas of primary and secondary education curricula, namely in health education, social science, physics, mathematics and biology. An online form was used for the experiment5.
Each learning resource had a metadata description, but the number of elements varied. However, they all had the following metadata: title, description, age range (all in English) and keywords. The keywords were comprised of tags and thesaurus terms, they were mixed together and displayed in an alphabetical order. The number of Thesaurus terms and tags varied for resources. Twenty of these keywords were thesaurus terms in English that an expert cataloger had used to classify the resource. The rest (39) were multilingual tags provided by pilot teachers during the three first months of the CALIBRATE pilot. These tags were both in commonly used languages and in less used languages as listed below:
  • 11 in Hungarian
  • 7 in German
  • 7 in English
  • 6 in Polish
  • 4 in Estonian
  • 1 in Finnish
The participants were asked to look at each learning resource at the time and go through the metadata related to it. Then, they were exposed to two different task related questions: first, to select the keywords that they found helped them to learn about the resource for the given learning resource (i.e. descriptiveness), and secondly, they were asked about decision support (i.e. help using the learning resources in teaching). Finally, they were also asked to rate the perceived overall quality of all the metadata displayed (traditional metadata plus keywords). This procedure was repeated for each one of the five learning resources.

Once the review of all the resources was concluded, the users were asked to identify their language competencies, and to indicate their comfort level when keywords were presented in languages that they did not understand. All these questions were mandatory to answer. The subjects commented later that in some cases they did not feel that any of the keywords was useful, but they had to choose one to conclude the web-survey. This might have skewed the results to some extend. Finally, participants had a choice to leave free comments about their experience during the experiment.

3.2Results

In this part we present the results of the experiment, which will be discussed further in in the following section.
On average, only 35% of the presented keywords, both Thesaurus and tags, were found descriptive for the learning resource. The thesaurus terms were found descriptive in 58% of the cases, while the tags only in 25% (Fig.3). When we look at the two top terms for each resource, we find that Thesaurus terms were somewhat more popular (60%) than tags (40%) (Table.2). All but one of the most popular tags were in English, which was also the most spoken language among the focus group. There were a lot of variations, by resource and by language groups, on how users perceived the keywords.

For example, for the first resource in the Fig.3, there was only one Thesaurus term and nine tags, which were in English and German, the languages widely spoken by participants. In this case two of the tags were chosen almost as often as the Thesaurus term. As for the second resources in Fig3, there was almost equal amount of tags and Thesaurus terms; two tags, both generic terms (EU, Europa) were chosen more often than Thesaurus terms.


Fig. 3. Percentage of tags and thesaurus terms found descriptive

The no:3 in the same figure represents a case of multilingual tags in less spoken languages in which the users did not have competences in. In this case two Thesaurus terms were most chosen, however, two “travel well” tags (JavaApplets, Applets) were very high on the list. As for the resource no:4, there was an equal number of tags and Thesaurus terms which was also displayed in the results, top two positions were held by both. In the last case the tags were in less spoken languages, in Hungarian and Estonian; one Hungarian tag was found useful by all with Hungarian skills.


Table 2. The two most popular keywords for each resource

From the total number of keywords, 54% were in a language within users competencies; however 87% of the keywords found descriptive were in a language that the user had skills in (Fig. 4). The remaining 13% of tags that were found useful, but not in the languages that users had competences, seem to comprise of terms of the generic type, the “travel well” tags, as described previously.

Fig. 4. Percentage of keywords in a known and unknown language that were found descriptive

When we asked about how well the keywords would help to use the resource, in average, only 27% of the presented keywords were found useful to indicate possible uses of the learning resource. Thesaurus terms were found useful 50% of the time, against only 18% for tags. In this case, when we look at the languages in which the participants had skills in, we find that in 83% of the time they mark those terms useful.
We can say that the issue of multilingual tags evokes sentiments and also splits users. From the thirteen users, two “love” being able to see multilingual tags and four found them useful, whereas six found them confusing and one “hates” to see keywords in languages that he/she does not understand (Fig. 5).


Fig. 5. Answer of the participants to the question: “What do you think when you see the keywords in many languages?”

Lastly, we were also interested in how users evaluated the overall quality of the metadata record [10]. The quality assigned to the metadata record correlate in a statistical significant way with the amount of words in the description (.909) and with how descriptive (.944) and useful (.994) the user found the keywords for that learning object. The first correlation was already found in a previous study [10].

4 Discussion on the results

The main argument that comes out of this early experimental research is that certain multilingual tags seem to be useful for some users – the challenge is how to make the other tags invisible? Moreover, the results can lead us to discuss the multilingualism of tags and indexing keywords from different perspectives; what are the user needs and requirements in a multilingual Europe, how can they be supported at the system level, what are the ramifications on their usability and how is the overall quality of the portal enhanced through multilingual tags?

In the spirit of “how to hide all but the right tags for each user”, this research has identified two topics that need further investigation: one is that of identifying “travel well” tags and the other that of how to correctly identify the language of each entered tag. After tackling these two issues, hiding all but the right tags becomes a much more manageable task.
Solving those two issues would greatly enhance the usability of the portal that offers multilingual tags: as shown in the experiment with the focus group, being exposed to tags in many languages has a dividing effect. One half of the subject expressed that they liked to see multilingual tags, whereas the other half found them rather irritating, especially when they were in languages that they did not recognise. It was also mentioned that multilingual tags make it harder and slower to pick the useful terms out of all the tags.

Two possible ways to further advance the cause could be envisaged: to automate the recognition of “travel well” tags and the identification of languages of all tags by using already existing vocabulary and dictionary lists on the Internet, or by crowd-sourcing” it to users, which is asking the end-users to identify “travel well” tags and allow them to translate and correct the language of tags. A co-existence of both could also be envisaged.
Another interesting outcome of the study is that keywords in general received a rather low appreciation rate among the subjects: 35% of the keywords were found descriptive and 27% were found helpful to the use of the resource. Overall, the Thesaurus terms performed better than the tags, however, it can be argued that tags, after all being produced with no outlay, showed an overall encouraging and potential gain in their usefulness. This needs to be investigated further and more in depth with a bigger sample size.

It could be envisaged that, in the case of sharing the accumulated knowledge regarding the actual use of resources in teaching and learning, social tagging could be in the future interesting in adding value to keywords. Thus, more design level effort is needed in guiding and encouraging users in using tags for such purposes.

5 Conclusions

This early study contributes to the understanding of tags in multiple languages. Despite the small sample size and early tagging behaviour of the participants, we can assume that tags in a multi-cultural and lingual context offer potential advantages to the collaborative tagging system and its multilingual user communities (e.g. Europe). However, there are challenges and research questions that need further attention. As it becomes clear that some tags are useful for some users, the design challenge becomes “hiding all but the right tags”. This implies for both entering and viewing the tags, e.g. what tags and in what languages to show/recommend to users when they are about to add a tag and what kind of tags to show for retrieval and social navigation.

First, it seems important that the system has a capacity to infer and identify tags entered in multiple languages, so that users can be shown or exposed to tags only in languages that they desire. Second, it appears that there are tags that “travel well”, i.e. tags that are easily understood by many users despite the lingual barriers. It appears important that those terms are identified, either automatically or by users, so that they could be better taken advantages of. The two above findings seem to further indicate that tags in different languages should not be kept as separate silos, but interaction between languages should be used for connecting like-minded people across country and linguistic borders.

The issue of multilingual tags is intriguing and offers interesting possibilities for both the learning resources repository managers and administrators, as well as for end users. In a multilingual environment such as Europe, where making learning resources available in languages others than in mother tongue is becoming more mainstream, mixing tagging with top-down expert classification system seem to offer interesting possibilities for accessing resources and for other novel educational applications that leverage the social network aspects of a given community. From this early experiment it becomes clear that further research into the topic of multilingualism is needed to better understand its complexity, but also to be able to design more adaptable applications.

Acknowledgments. We would like to thank Sylvia Hartinger from European Schoolet for making the tags available for analysis and Jim Ayre from Multimedia Ventures Europe Ltd. for valuable comments. Acknowledgment also goes to Helsingin Sanomain 100-vuotissäätiö for the research grant that made this research possible.

References
1. Marlow, C., Naaman, M., Boyd, D., Davis, M.: Position paper, tagging, taxonomy, flickr, article, toread. In: Collaborative Web Tagging Workshop at WWW2006, Edinburgh, Scotland. (2006).
2. Mathes, A.: Folksonomies-cooperative classification and communication through shared metadata. In: Computer Mediated Communication, Graduate School of Library and Information Science, University of Illinois Urbana-Champaign. (2004)
3. Golder, S.A., Huberman, B.A.: Usage patterns of collaborative tagging systems. Journal of Information Science 32(2), pp. 198—208. (2006).
4. Chi, E., Mytkowicz, T.: Understanding navigability of social tagging systems. In: Proceedings of CHI. Volume 7. (2007)
5. Catutto, C., Schmitz, C., Baldassarri, B., Servedio, V.D.P., Loreto, V., Hotho, A.,
Grahl, M., Stumme, G. Network Properties of Folksonomies. AI Communications
Journal, Special Issue on "Network Analysis in Natural Sciences and Engineering",
2007.
6. Hammond, T., Hannay, T., Lund, B., Scott, J.: Social bookmarking tools (i). D-Lib Magazine 11(4) (2005)
7. Guy, M., Tonkin, E.: Tidying up tags. D-Lib Magazine 12(1) (2006)
8. J.-N. Colin and D. Massart. LIMBS: Open source, open standards, and open content to foster learning resource exchanges. In Kinshuk, R. Koper, P. Kommers, P. Kirschner, D. Sampson, and W. Didderen, editors, Proc. of The Sixth IEEE International Conference on Advanced Learning Technologies, ICALT'06, pp. 682-686, Kerkrade, The Netherlands, July 2006.
9. Sen, S., Lam, S.K., Cosley, D., Frankowski, D., Osterhouse, J., Harper, F.M., Riedl, J.: tagging, communities, vocabulary, evolution. In: Proceedings of the 2006 20th anniversary conference on Computer supported cooperative work. pp. 181–190 (2006)
10. Ochoa, X., Duval, E.: Towards automatic evaluation of learning object metadata quality. In: Advances in Conceptual Modeling - Theory and Practice, ER 2006 Workshops BP-UML, CoMoGIS, COSS, ECDM, OIS, QoIS, SemWAT. pp. 372–381. Lecture Notes in Computer Science, Tucson, AZ, USA, Springer (November 2006)

Monday, July 23, 2007

Multilinguality of tags and etiquetas

I'm currently looking into the multilingual use of tags in different applications that allow users to taguér les favoris (=signet sociaux), music, photos, etc. and that display these etiquetas in a Nube de etiquetas = Tag-Wolke = nuage de tags. E.g. I took at quick tour on 10 collaborative tagging services to see how European multi-linguality is reflected on these services.

Lots of collaborative tagging efforts on the Web take place in English. English being the lingua-franca of the Internet, it is probably not that surprising. However, we here in Europe live in a multi-cultural and lingual environment, so traces of that should be found from here and there.

Fair enough, I was able to find French and German services (blogmarks.net, MisterWong.de, Oneview.de,..) that harbor communities of users who tag in their native language - as well as in English, too. Some people do not seem to have any difficulties in adding tags in a few different languages, e.g. talking about blogging, they add tags like "blogue", "blog" for for e-commerce "shopping" and "achat".

I also looked at some "global" services such as Yahoo!, del.icio.us, Amazon and LibraryThing.com to see how do they deal, if at all, with the issue of tags being in different languages, and in the case of Amazon and Yahoo!, who offer localised sites, how are tags managed and provided in different languages. Below I give a few examples.

Yahoo!'s MyWeb

Yahoo!'s MyWeb, which is their social bookmarking and tagging service. The service does not exist in all their localised sub-sites, but with a quick look I found it at least in Yahoo.fr, Yahoo.es, Yahoo.uk, Yahoo.de. On these respective MyWeb sites one can find popular bookmarks and tag clouds in the language of the sub-site.

In all the above mentioned sub-sites, regardless in which language I was looking at, 1116436 tags were registered. This first lead me to think that they actually have that many tags in German, Spanish and in French. Then I realised that it was a total of the tags in any languages, and they had much less tags in other languages than in English.
  • 1079 in German
  • 1052 in Spanish
  • 873 in French
What is remarkable in Yahoo!, though, is that they seem to have a way to recognise the language of the tag somehow, as they are able to display German tags in their .de service and French tags in their .fr service. This is transparent to the user, so I have only little idea (although a few guesses) how they do that. This focus on languages is rather unique on Yahoo!'s service, it is not the case with any of the other services that I looked. I'll be interested in knowing what they will come up with this!

LibraryThing.com

LibraryThing.com users, as well as developers, love tags! Users actually use them, after all, most of them are book freaks who probably hang out a lot in libraries, so tagging comes easy to these folks. But also, the developers have done a few really clever things with tags and objects to tag, like the concept of "work" (a work brings together all different copies of a book, regardless of edition, title variation, or language) and combining tags (see "concepts"). The difference here to delicious "bundled" tags is that it is done once for all users.

By combining tags, some multilinguality is also taken advantage of.

Like in this example, 19th century also includes 19. Jahrhundert and 19eme siecle.

del.icio.us

Take Delicious as an example of a different approach. It is one of the most used social bookmaking services, but there is little indication to be found of hundreds of languages that exist in the world. Well, I dug a bit harder and found two different types of indications of multilinguality existing.

Either people added tags in a few languages (achat, shopping), or they added a tag like "lang:fi" to indicate the that source site was in Finnish. Also lang:fr, lang:es, lang:pt, .. were there with various amounts of bookmarks. If you know if this was a user initiated activity or whether delicious encourages it, let me know. I could not find anything on it.

Amazon

I also looked at Amazon.com and Amazon.fr, funnily enough, the Frencheis do not even get to tag! (not that tags have been taken up in Amazon in the first place...) Moreover, they only get the reviews that are done by the users of the fr-site, not by the users of the .com-site. That sounds a bit silly, but maybe preserves a way to keep the cultural taste intact.

Different strokes for different folks

It seems to me that not many sites have paid much attention to the issue of multilinguality, however, it might also well be that it is not an important issue to them. Take for example these two German bookmarking sites, Mister Wong and Oneview.de, and compare the Top-50 tag lists.

MrWong.de (image on the left) has mostly English terms in their Top-50 and they are very Web2.0 and developer oriented.

Oneview.de has many German terms (image on the right) and they seem to cover larger area than only Web2.0, there are terms about holidays, music, etc.

So, each user group seem to share the terms that are important for them. Most likely this is also the case in Delicious, by using English tags I, as a Finn, can easily share interesting bookmarks on, say, folksonomies with everyone else in the world who is interested in it.

This to say, I also think that multilinguality has a place in tagging and we in EUN are very interested in the issue. Currently, we are running a pilot where teachers can add tags in their own language(s) to learning resources. We try our best to be able to recognise the language of the tags, as we think cool things can be done with it. But more about that another time.

Sites that I looked at

French:
  • social bookmarks =signet sociaux
  • tag, taguer

Blogmarks, http://blogmarks.net/
- most popular tags are in English
- 577366 bookmarks
- people sometimes add tags both in French and English (e.g. blogue, blog; tag cloud nuage de tags)
- Tags in French also, but they don't seem to have that many users, tried some about 10 or so


Bookmakrs, http://bookmarks.fr
- call bookmarks "favoris" and tags "tags"
- most popular tags in different languages, in 50 most popular tags: 14 in French and 1 in Russian
- maybe only about 2500 bookmarks (25x107 pages)
- some people had indicated the language of the source page in En

In German:
  • Tags as "Schlagworte", "tags"
  • "Tag-Wolke"
  • bookmarks " Lesezeichen" (Soziale Lesezeichen)

Netselektor, http://netselektor.de/

Mister Wong, http://www.mister-wong.de/?tag_type=list
- 2.142.735 Bookmarks
- very techy, lots of En tags

Oneview http://www.oneview.de/home/index.jsf
- Trendwolke http://www.oneview.de/home/discover.jsf
- more tags in German in general areas


Yahoo! MyWeb

Yahoo! Fr
- top tags http://fr.myweb2.search.yahoo.com/myweb?dg=6&sort=pop
- 1 116 436 tags in the "nuage de tags", although 873 in French.
- in top 50 tags mostly in French, however, many terms are easily understandable like web2.0, internet, google, yahoo..

Yahoo! De
- top tags http://de.myweb2.search.yahoo.com/myweb?ei=UTF-8&dg=6&dmode=vtags&sortby=count
- 1.116.436 tags in "Tag-Wolke", although only 1079 in German
- in top 50 most entirely in German, but many terms like computer, software

Yahoo! es
- toptags http://es.myweb2.search.yahoo.com/myweb?dg=6&sort=pop
- call tags "Etiquetas" and tagcloud
- 1.116.436 tags in "Nube de etiquetas", 1052 in Spanish
- the top 50 mostly in Spanish, only a few terms like web2.0, blog, yahoo



Flickr
http://www.flickr.com/photos/tags/

- interface offered in different languages, however,
- most popular tags are the same in all different interface languages, thus could think that there is no separation of languages.


Delicious

- All top 50 tags in English
- no other interface languages
- some tags do exist in different languages e.g. achat, shopping; ..
- some tags indicate the language of the source; lang:fi,..

Google

if I understand right, google does not even display the tags from their users? Please, correct me if you know anything about using their bookmarks.

Last.fm

- in LastFM all the tags are the same, although they offer different interface languages
- tags in many language exist, especially if you look for bands from different countries e.g. suomipoppia

Sunday, June 03, 2007

Cyprus and 3 things I didn't know...

Taxi drivers are a good source of information. Ok, of course, there are many types of taxi drivers; the ones that literally drive you crazy, the ones to whom you pay no attention, and then the chatty ones who take advantage of the fact that you can't escape.

On the drive to the airport in Cyprus this morning at 6am, the local taxi driver, Pavarotti, as they all call him, told me about his passionate, yet unsuccessful love life, about the recent surge of Russians on the island, and about the division of their small island. Quite an interesting hour.

All in all, my 3-day trip to Cyprus revealed a nation in the mist of turbulence, with lots of smart and warm-hearted people. Three things that I did not know about Cyprus:

1/ They were under British rule for quite some time, hence the legacy of driving on the left, using pounds, and the messed up UK style plugs and sockets. They still seem to have a rather affectionate relation with Brits and they all speak rather good English and take bride of it!

There is only one university on the island, so many study abroad. About half of them in Greece, the rest the UK and the US mostly. Surprisingly many who work in the Ministry of Ed also held MEds and PhDs from abroad! (see the pic of them put in good use: reading my fortune from a Turkish coffee cup)

2/ Since the collapse of Soviet Union, many Russians have moved to Cyprus. I was told that even Putin has his own datcha there. According to Pavarotti (taxi driver), they find it as a small paradise. (maybe it is because of the 30% of votes that the Communist party still receives - I didn't know that either!)

The standard of living is rather OK, not too expensive, the weather is good, and the Cypriot men seem to be crazy about the glamorous style of Russian women, you know, the bells and whistles on high-heels.

I met quite a few Russians there, some really smart ones working in the oil business (apparently for tax reasons some businesses have re-located there), more working as waitresses, and then also the lot earning living with what they got. People told me about the problems raising of these gals on loose - the small society is facing issues on many levels; on the family-level it's resulted in broken marriages, money spent on prostitution, and in schools one also faces issues that were not there a few years back.

Even though Cyprus has been occupied in many occasions, the topic of Russians came up directly and indirectly in quite a few discussions. Interestingly, one person said that instead of seeing influence coming from the EU, now that they are members, they only see issues stirring up with Russians and Asians. I also noticed quite a few Asians in the country. I was told that many work as maid at homes and hotels.

3/ They also have mountains in Cyprus! Almost 2km! One can ski in the winter time. A new destination to be added on my list of strange places to go skiing :)