Saturday, August 04, 2007

Draft paper: Analysis of User Behavior on Multilingual tagging of learning resources

This is an almost final draft of a paper that I'm currently working on. It's been accepted as a full paper to the SIRTEL workshop.

Ah, should be mentioned, maybe, that I'm also co-chairing it :) It's gonna be very cool, so try to make it there, if possible.

If not, you can always think of posting a question in YouTube, like they did in the US presidential campaign. I kind of like that, although I don't think that I get CNN to co-host it!

Anyway, comments are welcome on this on. All images are missing, I was testing Google docs for this-copy and paste from OO did not include images.

Also, the formating took some damage, sorry about that. The final, more readable version will be at the conference site in about 10 days.

Analysis of User Behavior on Multilingual Tagging of Learning resources
Riina Vuorikari1,, Xavier Ochoa2, and Erik Duval1


Abstract. Although social, collaborative classification through tagging has been the focus of recent research, the effect of multilingual tags is often overlooked. This work presents an early exploratory study of the production and consumption of multilingual tags in a European educational K-12 context. The data, produced by teachers bookmarking and tagging learning resources during three month period, was analysed. Thereafter, this information was presented in the form of metadata keywords to a focus group of teachers who evaluated its descriptiveness, usefulness and overall quality. The results of this early study suggest that users are divided about the benefits of multilingual tags, however, some tags are useful for some users, thus “hiding all but the right tags” becomes crucial for the success of a multilingual collaborative tagging system.

Keywords: Collaborative tagging, multilinguality, learning resources.

1 Introduction
The use of social, collaborative classification systems has gone through a continuous growth in the latest years [1]. An example of this is a multitude of sites that provide some type of social annotation of digital artefacts and a social navigation system (Flikr, del.icio.us , CiteULike, Last.fm, among others). Social tagging, i.e. allowing individuals to apply free text keywords to digital objects, potentially offers advantages in terms of personal knowledge management, serendipitous access to objects through tags, and enhanced possibilities to share content with emerging social networks.

Several studies have been undertaken to better understand the behaviour and evolution of social tagging systems. Early research has been conducted by Mathes [2] where the term “folksonomy” is used to compare the emerging socially generated vocabulary with the more formal ontology concept. Golder and Huberman [3] first looked at user patterns of collaborative tagging systems. Recent studies focus on the navigability of such social systems [4] and on understanding the network properties [5].

A prevailing aspect among current studies concerning tagging is that they assume that tags are represented in a common language [6], understandable by all the members of the user community. Guy suggests that it is not always the case [7], but does not offer insight on how to deal with tags in multiple languages.

Lately, multilingual tags have started emerging on popular social tagging systems as their user-base grows, and different ways to deal with multiple languages can be observed. Delicious users, for example, add tags in different languages for a bookmark (e.g. achat, shopping) and even in some occasions add language identification in tags (e.g. lang:fi) for the language of the resource. However, it does not offer any system level support, that allows users to see tags, say, only in French or Finnish. Other services, like Yahoo!'s MyWeb on the other hand, offer tags and tag clouds in different languages in their localised parts of the portal (e.g. .fr, .es, ...), thus some language identification of tags takes place on the system level. Thirdly, In LibraryThing experienced users can combine tags, where in some occasions tags in different languages have been grouped together.

Our work, still at its early stage, attempts to shed light on a community of users who shares a common educational interest to use a social tagging system across country and language borders, but does not necessarily share a common language, as the users are free to choose the language(s) in which they apply tags. This exploration takes place in the context of two European Community founded projects, CALIBRATE1 and MELT2, both focusing on sharing and re-using of digital learning resources for K-12.

European education, especially that of K-12 education, is inherently multilingual and multicultural. Offering educational resources and services in native languages is deemed important, but equally important is the exposure to other languages. One way to promote this is to make learning resources available across national and linguistic boarders. This puts constraints on semantic interoperability, i.e. how well content and its metadata can be understood by other systems and users.

Controlled vocabularies, such as multilingual LRE Thesaurus3, can be used to overcome some hurdles of semantic interoperability. However, the gap between the terms used by experts and practitioners in the field is also problematic. For that reason, the current research looks into co-existence of taxonomies and end-user generated tags.

A federation of learning resource repositories in a multilingual context needs to support multiple languages at the system level in order to support each repository and its national user-base, but at the same time, there is a need to allow people (i.e. user information and preferences), resources and tags to “travel” across national and linguistic borders.
This paper is structured as follows: first, in section 2, we analyse the early stage of the bookmarking and tagging behavior of our community in order to better understand how teachers bookmark and tag resources in a multilingual context; what types of tags are provided and in which languages. Then, in the section 3, an experiment is presented that measures the effect of multilingual tags on the descriptiveness, usefulness and overall quality of the metadata. Finally, the findings are discussed and applied to design decisions for multilingual tagging systems.

2 Analysis of tagging behaviour in multilingual context

The CALIBRATE project makes K-12 digital learning resources available to its pilot schools (78 schools ) in Hungary, Austria, Estonia, Czech Republic, Lithuania and Poland in their different curriculum areas. Schools can access material in different languages through a portal that is connected to a federation of learning resource repositories [8] in the pilot countries.
As part of the project's multilingual search interface4, a personal bookmarking and tagging tool has been available since the beginning of 2007. This tool allows a user to create personal collections of learning resources by bookmarking interesting resources found through the portal. To facilitate the management of these personal collections (also called favourites in the project), the user can also add keywords to resources to make it easier to ”keep found things found”. These keywords are free for the user to choose and can be expressed in any language. The collections and keywords are kept private to the user, and at this stage of the experiment, they cannot be shared among users.

The data for this analysis is from a period of about three months (January 24 to April 21 2007). There were 77 teachers who made 459 bookmarks with 417 multilingual tags on 320 different learning resources. It is intended to have regular analysis of this data within the projects lifespan (-2008).

2.1 Quasi-Experimental Set-up

A total of 173 subjects used the portal during the time of the experiment, however, the subjects of this dataset comprises of a group of 77 teachers who had done at least one bookmark during this time. Thus, it was a self-selected group formed based on the bookmarking behavior during the period of three months and it represents 45% of all the pilot participants. As there was no overall methodology to introduce bookmarking and tagging to the subjects, more than half of the participants had not shown interest in using this feature of the portal.

The bookmarking habits, at this very early stage, varied a lot in terms of what languages to use, how many tags to add, how to add multiple tags (with comma separated or without commas), etc. Also, hardly any of the participants had previous experience on tagging, so not one single tagging convention emerged, rather many different ways to use tags in multiple languages. As there is very little research done on the multilingual context, we think it is important to study the early stage of tagging behaviour to better anticipate the effect of multilinguality on the system to improve its design.

It is noteworthy to mention that the bookmarking and tagging system, at this stage of the pilot, offers very little social influence in what comes to choosing what to bookmark and what keywords to choose. Oftentimes in social bookmarking sites, social cues are made available (e.g. most bookmarked items, tag clouds, tags are recommended based on previous tags, etc). At the time of the experiment, the system had hardly any tags attached to resources, so teachers started from an empty plate. In the case where a resource was already tagged by another participant, the user would see the term(s) only if they were in the same language as the interface is.

2.2 Results

In the part we present the results of the analysis, which will be discussed further in conjunction with the other results in the discussion section.

When we look at the distribution of bookmarks per users, we can find that on the average, each user had 6 bookmarks (Fig.1). However, the distribution was very wide; 10% of the users had more than the average amount of bookmarks, which leaves 90% under the average. Eight of the users could be called “super users”, as they had more than 20 bookmarks, and 12 users had between 20 and 6 bookmarks. About 30% of the users seem to have only experimented with the bookmarking system, as they only have one single bookmarked item in their favorites folder.


We had recorded 418 tags in the system. During the semantic analysis of tags we found that many tags actually contained multiple terms, i.e. they were bundles of terms without comma separation. This was due to a technical feature of the tool that treated terms without comma separation as one tag. When broken down, they resulted in 585 terms. They were translated into English and a semantic analysis was performed to better understand the types of tags. We used the classification from Sen [9] that is also based on the categories of Golder et al. [3], which are Factual tags (Golder: item topics, kinds of item, category refinements); Subjective tags (Golder: item qualities) and Personal tags (Golder: item ownership, self-reference, tasks organisation)

The vast majority of the tags at this early stage (Table 1) are of the factual type. From the factual tags, 79% were put into a rough category of topic and 14% of the category refinement with richer information. The rest of the tags were subjective in their nature and could be used to describe the quality of the resources or how the person felt about them. None of the tags fell into the category of personal tags as Golder describes them (e.g. tags related to item ownership, self-reference or personal tasks organisation). When we analysed how these tags were used and re-used among users, we found that 80% of tags related to bookmarks were factual and 20% of tags subjective tags. In a MovieLens study [9], for comparison, the distribution was 63% factual, 29% subjective, 3% personal and 5% other.
Table 1. Types analysis of each tags (no re-use)
Factual
340
93%
Topic
Category refinement
288
52
79%
14%
Subjective
24
7%
Personal
0
0%

After categorising the tags, we further studied their nature. Two main trends seemed to emerge, first, many of the tags contained the same terms as in the title, i.e. user had just copied the title in the tag field. Second, about 13% of tags contain a general term, a name, place, e.g. EU, Euroopa, Euroopa, Europa, europe, geograafia, Phytagoras, etc . We hypothesise that this type of “travel well” tags, even if not translated, could be found useful for other users for their close similarity in spelling in many languages. We think it could be of interest to work towards automatically filter this type of terms from the pool of all multilingual tags, for example, by matching them against existing multilingual vocabulary lists available on the Internet.
When we looked at the number of tags that users related to bookmarks, we were able to identify some early trends. For the total of 459 bookmarked resources, we found that some of the tags were re-used, there was an average of 1.92 tags/resource. More than half (56%) of the tags were entered as a bundle of terms, i.e. most teachers had added 2 to 6 terms without a comma separation. In quite a few cases these terms were comprised of the terms in the title of the resource (Fig.2). In 28% of the cases only one term was entered as one tag.

The rest had used multiple separate tags (2-6 tags). In the latter case the terms were not necessary related to the title alone, but carried other types of information (e.g. title: Umweltkids and tags: Oekologie, Artenschutz, Regenwald, Tierschutz, Skisport).
Contrary to our expectations, the users took liberties to add tags in multiple languages and to use the portal interface in different languages than that of their mother tongue (interface was made available in the languages of the pilot and in English). This made the identification of the language of tags more difficult, as we had expected to be able to identify the language of the tag from the language of the interface that the user used when inserting the tag. In about 70% of the cases we were able to identify the language of the tag correctly using this method, which leaves us with a 30% error rate on language identification. This error in identifying the language of the tag correctly would make it hard, for example, to display tags and tag clouds in one single language, an issue that is related to the usability of the portal, and the one of which the second experience was set up to find more evidence.
We found the following scenarios for tagging, however, due to our logging, we can't give percentages for these use cases:
  • Interface and tags in mother tongue
  • Interface was used in mother tongue, but tags in other language
  • Interface was used in a language that is other than the mother tongue, but tags were entered in mother tongue
  • The tagging language was other than the interface language and the mother tongue
These scenarios were found through comparing the real language of the tags to that of identified language by using the interface language. In this early stage of the experiment it is impossible to draw firm conclusions, but it seems that users are likely to use, or at least try, the interface in different languages. We found, for example, that the tags entered through the English interface were in English only in 50% of the cases, which means that users added tags in languages within their areas of competences. On the other hand, we also found that there were many more tags in English than we expected from the choice of the interface language. These users had chosen to tag in English, even if they used the interface in some other language, most likely to be able to share tags with users from other countries.

3 Experiment with Multilingual Tags

An exploratory experiment was set up in order to measure the perceived usefulness and quality [10] of multilingual tags, traditional metadata and expert classification keywords. We were also interested in how users reacted when they were confronted with tags in multiple languages that they did not have knowledge of. The experiment subjects were shown a list of learning resources metadata with keywords in multiple languages, the list was imitating the search result list of the portal. The results of this experiment will be useful to guide design decision in the development of retrieval tools for learning objects in a multilingual environment.

3.1 Experimental Set-up

Thirteen teachers, who belong to the MELT focus group, were selected to participate in the experiment. They were confronted with metadata regarding five learning resources in different areas of primary and secondary education curricula, namely in health education, social science, physics, mathematics and biology. An online form was used for the experiment5.
Each learning resource had a metadata description, but the number of elements varied. However, they all had the following metadata: title, description, age range (all in English) and keywords. The keywords were comprised of tags and thesaurus terms, they were mixed together and displayed in an alphabetical order. The number of Thesaurus terms and tags varied for resources. Twenty of these keywords were thesaurus terms in English that an expert cataloger had used to classify the resource. The rest (39) were multilingual tags provided by pilot teachers during the three first months of the CALIBRATE pilot. These tags were both in commonly used languages and in less used languages as listed below:
  • 11 in Hungarian
  • 7 in German
  • 7 in English
  • 6 in Polish
  • 4 in Estonian
  • 1 in Finnish
The participants were asked to look at each learning resource at the time and go through the metadata related to it. Then, they were exposed to two different task related questions: first, to select the keywords that they found helped them to learn about the resource for the given learning resource (i.e. descriptiveness), and secondly, they were asked about decision support (i.e. help using the learning resources in teaching). Finally, they were also asked to rate the perceived overall quality of all the metadata displayed (traditional metadata plus keywords). This procedure was repeated for each one of the five learning resources.

Once the review of all the resources was concluded, the users were asked to identify their language competencies, and to indicate their comfort level when keywords were presented in languages that they did not understand. All these questions were mandatory to answer. The subjects commented later that in some cases they did not feel that any of the keywords was useful, but they had to choose one to conclude the web-survey. This might have skewed the results to some extend. Finally, participants had a choice to leave free comments about their experience during the experiment.

3.2Results

In this part we present the results of the experiment, which will be discussed further in in the following section.
On average, only 35% of the presented keywords, both Thesaurus and tags, were found descriptive for the learning resource. The thesaurus terms were found descriptive in 58% of the cases, while the tags only in 25% (Fig.3). When we look at the two top terms for each resource, we find that Thesaurus terms were somewhat more popular (60%) than tags (40%) (Table.2). All but one of the most popular tags were in English, which was also the most spoken language among the focus group. There were a lot of variations, by resource and by language groups, on how users perceived the keywords.

For example, for the first resource in the Fig.3, there was only one Thesaurus term and nine tags, which were in English and German, the languages widely spoken by participants. In this case two of the tags were chosen almost as often as the Thesaurus term. As for the second resources in Fig3, there was almost equal amount of tags and Thesaurus terms; two tags, both generic terms (EU, Europa) were chosen more often than Thesaurus terms.


Fig. 3. Percentage of tags and thesaurus terms found descriptive

The no:3 in the same figure represents a case of multilingual tags in less spoken languages in which the users did not have competences in. In this case two Thesaurus terms were most chosen, however, two “travel well” tags (JavaApplets, Applets) were very high on the list. As for the resource no:4, there was an equal number of tags and Thesaurus terms which was also displayed in the results, top two positions were held by both. In the last case the tags were in less spoken languages, in Hungarian and Estonian; one Hungarian tag was found useful by all with Hungarian skills.


Table 2. The two most popular keywords for each resource

From the total number of keywords, 54% were in a language within users competencies; however 87% of the keywords found descriptive were in a language that the user had skills in (Fig. 4). The remaining 13% of tags that were found useful, but not in the languages that users had competences, seem to comprise of terms of the generic type, the “travel well” tags, as described previously.

Fig. 4. Percentage of keywords in a known and unknown language that were found descriptive

When we asked about how well the keywords would help to use the resource, in average, only 27% of the presented keywords were found useful to indicate possible uses of the learning resource. Thesaurus terms were found useful 50% of the time, against only 18% for tags. In this case, when we look at the languages in which the participants had skills in, we find that in 83% of the time they mark those terms useful.
We can say that the issue of multilingual tags evokes sentiments and also splits users. From the thirteen users, two “love” being able to see multilingual tags and four found them useful, whereas six found them confusing and one “hates” to see keywords in languages that he/she does not understand (Fig. 5).


Fig. 5. Answer of the participants to the question: “What do you think when you see the keywords in many languages?”

Lastly, we were also interested in how users evaluated the overall quality of the metadata record [10]. The quality assigned to the metadata record correlate in a statistical significant way with the amount of words in the description (.909) and with how descriptive (.944) and useful (.994) the user found the keywords for that learning object. The first correlation was already found in a previous study [10].

4 Discussion on the results

The main argument that comes out of this early experimental research is that certain multilingual tags seem to be useful for some users – the challenge is how to make the other tags invisible? Moreover, the results can lead us to discuss the multilingualism of tags and indexing keywords from different perspectives; what are the user needs and requirements in a multilingual Europe, how can they be supported at the system level, what are the ramifications on their usability and how is the overall quality of the portal enhanced through multilingual tags?

In the spirit of “how to hide all but the right tags for each user”, this research has identified two topics that need further investigation: one is that of identifying “travel well” tags and the other that of how to correctly identify the language of each entered tag. After tackling these two issues, hiding all but the right tags becomes a much more manageable task.
Solving those two issues would greatly enhance the usability of the portal that offers multilingual tags: as shown in the experiment with the focus group, being exposed to tags in many languages has a dividing effect. One half of the subject expressed that they liked to see multilingual tags, whereas the other half found them rather irritating, especially when they were in languages that they did not recognise. It was also mentioned that multilingual tags make it harder and slower to pick the useful terms out of all the tags.

Two possible ways to further advance the cause could be envisaged: to automate the recognition of “travel well” tags and the identification of languages of all tags by using already existing vocabulary and dictionary lists on the Internet, or by crowd-sourcing” it to users, which is asking the end-users to identify “travel well” tags and allow them to translate and correct the language of tags. A co-existence of both could also be envisaged.
Another interesting outcome of the study is that keywords in general received a rather low appreciation rate among the subjects: 35% of the keywords were found descriptive and 27% were found helpful to the use of the resource. Overall, the Thesaurus terms performed better than the tags, however, it can be argued that tags, after all being produced with no outlay, showed an overall encouraging and potential gain in their usefulness. This needs to be investigated further and more in depth with a bigger sample size.

It could be envisaged that, in the case of sharing the accumulated knowledge regarding the actual use of resources in teaching and learning, social tagging could be in the future interesting in adding value to keywords. Thus, more design level effort is needed in guiding and encouraging users in using tags for such purposes.

5 Conclusions

This early study contributes to the understanding of tags in multiple languages. Despite the small sample size and early tagging behaviour of the participants, we can assume that tags in a multi-cultural and lingual context offer potential advantages to the collaborative tagging system and its multilingual user communities (e.g. Europe). However, there are challenges and research questions that need further attention. As it becomes clear that some tags are useful for some users, the design challenge becomes “hiding all but the right tags”. This implies for both entering and viewing the tags, e.g. what tags and in what languages to show/recommend to users when they are about to add a tag and what kind of tags to show for retrieval and social navigation.

First, it seems important that the system has a capacity to infer and identify tags entered in multiple languages, so that users can be shown or exposed to tags only in languages that they desire. Second, it appears that there are tags that “travel well”, i.e. tags that are easily understood by many users despite the lingual barriers. It appears important that those terms are identified, either automatically or by users, so that they could be better taken advantages of. The two above findings seem to further indicate that tags in different languages should not be kept as separate silos, but interaction between languages should be used for connecting like-minded people across country and linguistic borders.

The issue of multilingual tags is intriguing and offers interesting possibilities for both the learning resources repository managers and administrators, as well as for end users. In a multilingual environment such as Europe, where making learning resources available in languages others than in mother tongue is becoming more mainstream, mixing tagging with top-down expert classification system seem to offer interesting possibilities for accessing resources and for other novel educational applications that leverage the social network aspects of a given community. From this early experiment it becomes clear that further research into the topic of multilingualism is needed to better understand its complexity, but also to be able to design more adaptable applications.

Acknowledgments. We would like to thank Sylvia Hartinger from European Schoolet for making the tags available for analysis and Jim Ayre from Multimedia Ventures Europe Ltd. for valuable comments. Acknowledgment also goes to Helsingin Sanomain 100-vuotissäätiö for the research grant that made this research possible.

References
1. Marlow, C., Naaman, M., Boyd, D., Davis, M.: Position paper, tagging, taxonomy, flickr, article, toread. In: Collaborative Web Tagging Workshop at WWW2006, Edinburgh, Scotland. (2006).
2. Mathes, A.: Folksonomies-cooperative classification and communication through shared metadata. In: Computer Mediated Communication, Graduate School of Library and Information Science, University of Illinois Urbana-Champaign. (2004)
3. Golder, S.A., Huberman, B.A.: Usage patterns of collaborative tagging systems. Journal of Information Science 32(2), pp. 198—208. (2006).
4. Chi, E., Mytkowicz, T.: Understanding navigability of social tagging systems. In: Proceedings of CHI. Volume 7. (2007)
5. Catutto, C., Schmitz, C., Baldassarri, B., Servedio, V.D.P., Loreto, V., Hotho, A.,
Grahl, M., Stumme, G. Network Properties of Folksonomies. AI Communications
Journal, Special Issue on "Network Analysis in Natural Sciences and Engineering",
2007.
6. Hammond, T., Hannay, T., Lund, B., Scott, J.: Social bookmarking tools (i). D-Lib Magazine 11(4) (2005)
7. Guy, M., Tonkin, E.: Tidying up tags. D-Lib Magazine 12(1) (2006)
8. J.-N. Colin and D. Massart. LIMBS: Open source, open standards, and open content to foster learning resource exchanges. In Kinshuk, R. Koper, P. Kommers, P. Kirschner, D. Sampson, and W. Didderen, editors, Proc. of The Sixth IEEE International Conference on Advanced Learning Technologies, ICALT'06, pp. 682-686, Kerkrade, The Netherlands, July 2006.
9. Sen, S., Lam, S.K., Cosley, D., Frankowski, D., Osterhouse, J., Harper, F.M., Riedl, J.: tagging, communities, vocabulary, evolution. In: Proceedings of the 2006 20th anniversary conference on Computer supported cooperative work. pp. 181–190 (2006)
10. Ochoa, X., Duval, E.: Towards automatic evaluation of learning object metadata quality. In: Advances in Conceptual Modeling - Theory and Practice, ER 2006 Workshops BP-UML, CoMoGIS, COSS, ECDM, OIS, QoIS, SemWAT. pp. 372–381. Lecture Notes in Computer Science, Tucson, AZ, USA, Springer (November 2006)

Monday, July 23, 2007

Multilinguality of tags and etiquetas

I'm currently looking into the multilingual use of tags in different applications that allow users to taguér les favoris (=signet sociaux), music, photos, etc. and that display these etiquetas in a Nube de etiquetas = Tag-Wolke = nuage de tags. E.g. I took at quick tour on 10 collaborative tagging services to see how European multi-linguality is reflected on these services.

Lots of collaborative tagging efforts on the Web take place in English. English being the lingua-franca of the Internet, it is probably not that surprising. However, we here in Europe live in a multi-cultural and lingual environment, so traces of that should be found from here and there.

Fair enough, I was able to find French and German services (blogmarks.net, MisterWong.de, Oneview.de,..) that harbor communities of users who tag in their native language - as well as in English, too. Some people do not seem to have any difficulties in adding tags in a few different languages, e.g. talking about blogging, they add tags like "blogue", "blog" for for e-commerce "shopping" and "achat".

I also looked at some "global" services such as Yahoo!, del.icio.us, Amazon and LibraryThing.com to see how do they deal, if at all, with the issue of tags being in different languages, and in the case of Amazon and Yahoo!, who offer localised sites, how are tags managed and provided in different languages. Below I give a few examples.

Yahoo!'s MyWeb

Yahoo!'s MyWeb, which is their social bookmarking and tagging service. The service does not exist in all their localised sub-sites, but with a quick look I found it at least in Yahoo.fr, Yahoo.es, Yahoo.uk, Yahoo.de. On these respective MyWeb sites one can find popular bookmarks and tag clouds in the language of the sub-site.

In all the above mentioned sub-sites, regardless in which language I was looking at, 1116436 tags were registered. This first lead me to think that they actually have that many tags in German, Spanish and in French. Then I realised that it was a total of the tags in any languages, and they had much less tags in other languages than in English.
  • 1079 in German
  • 1052 in Spanish
  • 873 in French
What is remarkable in Yahoo!, though, is that they seem to have a way to recognise the language of the tag somehow, as they are able to display German tags in their .de service and French tags in their .fr service. This is transparent to the user, so I have only little idea (although a few guesses) how they do that. This focus on languages is rather unique on Yahoo!'s service, it is not the case with any of the other services that I looked. I'll be interested in knowing what they will come up with this!

LibraryThing.com

LibraryThing.com users, as well as developers, love tags! Users actually use them, after all, most of them are book freaks who probably hang out a lot in libraries, so tagging comes easy to these folks. But also, the developers have done a few really clever things with tags and objects to tag, like the concept of "work" (a work brings together all different copies of a book, regardless of edition, title variation, or language) and combining tags (see "concepts"). The difference here to delicious "bundled" tags is that it is done once for all users.

By combining tags, some multilinguality is also taken advantage of.

Like in this example, 19th century also includes 19. Jahrhundert and 19eme siecle.

del.icio.us

Take Delicious as an example of a different approach. It is one of the most used social bookmaking services, but there is little indication to be found of hundreds of languages that exist in the world. Well, I dug a bit harder and found two different types of indications of multilinguality existing.

Either people added tags in a few languages (achat, shopping), or they added a tag like "lang:fi" to indicate the that source site was in Finnish. Also lang:fr, lang:es, lang:pt, .. were there with various amounts of bookmarks. If you know if this was a user initiated activity or whether delicious encourages it, let me know. I could not find anything on it.

Amazon

I also looked at Amazon.com and Amazon.fr, funnily enough, the Frencheis do not even get to tag! (not that tags have been taken up in Amazon in the first place...) Moreover, they only get the reviews that are done by the users of the fr-site, not by the users of the .com-site. That sounds a bit silly, but maybe preserves a way to keep the cultural taste intact.

Different strokes for different folks

It seems to me that not many sites have paid much attention to the issue of multilinguality, however, it might also well be that it is not an important issue to them. Take for example these two German bookmarking sites, Mister Wong and Oneview.de, and compare the Top-50 tag lists.

MrWong.de (image on the left) has mostly English terms in their Top-50 and they are very Web2.0 and developer oriented.

Oneview.de has many German terms (image on the right) and they seem to cover larger area than only Web2.0, there are terms about holidays, music, etc.

So, each user group seem to share the terms that are important for them. Most likely this is also the case in Delicious, by using English tags I, as a Finn, can easily share interesting bookmarks on, say, folksonomies with everyone else in the world who is interested in it.

This to say, I also think that multilinguality has a place in tagging and we in EUN are very interested in the issue. Currently, we are running a pilot where teachers can add tags in their own language(s) to learning resources. We try our best to be able to recognise the language of the tags, as we think cool things can be done with it. But more about that another time.

Sites that I looked at

French:
  • social bookmarks =signet sociaux
  • tag, taguer

Blogmarks, http://blogmarks.net/
- most popular tags are in English
- 577366 bookmarks
- people sometimes add tags both in French and English (e.g. blogue, blog; tag cloud nuage de tags)
- Tags in French also, but they don't seem to have that many users, tried some about 10 or so


Bookmakrs, http://bookmarks.fr
- call bookmarks "favoris" and tags "tags"
- most popular tags in different languages, in 50 most popular tags: 14 in French and 1 in Russian
- maybe only about 2500 bookmarks (25x107 pages)
- some people had indicated the language of the source page in En

In German:
  • Tags as "Schlagworte", "tags"
  • "Tag-Wolke"
  • bookmarks " Lesezeichen" (Soziale Lesezeichen)

Netselektor, http://netselektor.de/

Mister Wong, http://www.mister-wong.de/?tag_type=list
- 2.142.735 Bookmarks
- very techy, lots of En tags

Oneview http://www.oneview.de/home/index.jsf
- Trendwolke http://www.oneview.de/home/discover.jsf
- more tags in German in general areas


Yahoo! MyWeb

Yahoo! Fr
- top tags http://fr.myweb2.search.yahoo.com/myweb?dg=6&sort=pop
- 1 116 436 tags in the "nuage de tags", although 873 in French.
- in top 50 tags mostly in French, however, many terms are easily understandable like web2.0, internet, google, yahoo..

Yahoo! De
- top tags http://de.myweb2.search.yahoo.com/myweb?ei=UTF-8&dg=6&dmode=vtags&sortby=count
- 1.116.436 tags in "Tag-Wolke", although only 1079 in German
- in top 50 most entirely in German, but many terms like computer, software

Yahoo! es
- toptags http://es.myweb2.search.yahoo.com/myweb?dg=6&sort=pop
- call tags "Etiquetas" and tagcloud
- 1.116.436 tags in "Nube de etiquetas", 1052 in Spanish
- the top 50 mostly in Spanish, only a few terms like web2.0, blog, yahoo



Flickr
http://www.flickr.com/photos/tags/

- interface offered in different languages, however,
- most popular tags are the same in all different interface languages, thus could think that there is no separation of languages.


Delicious

- All top 50 tags in English
- no other interface languages
- some tags do exist in different languages e.g. achat, shopping; ..
- some tags indicate the language of the source; lang:fi,..

Google

if I understand right, google does not even display the tags from their users? Please, correct me if you know anything about using their bookmarks.

Last.fm

- in LastFM all the tags are the same, although they offer different interface languages
- tags in many language exist, especially if you look for bands from different countries e.g. suomipoppia

Sunday, June 03, 2007

Cyprus and 3 things I didn't know...

Taxi drivers are a good source of information. Ok, of course, there are many types of taxi drivers; the ones that literally drive you crazy, the ones to whom you pay no attention, and then the chatty ones who take advantage of the fact that you can't escape.

On the drive to the airport in Cyprus this morning at 6am, the local taxi driver, Pavarotti, as they all call him, told me about his passionate, yet unsuccessful love life, about the recent surge of Russians on the island, and about the division of their small island. Quite an interesting hour.

All in all, my 3-day trip to Cyprus revealed a nation in the mist of turbulence, with lots of smart and warm-hearted people. Three things that I did not know about Cyprus:

1/ They were under British rule for quite some time, hence the legacy of driving on the left, using pounds, and the messed up UK style plugs and sockets. They still seem to have a rather affectionate relation with Brits and they all speak rather good English and take bride of it!

There is only one university on the island, so many study abroad. About half of them in Greece, the rest the UK and the US mostly. Surprisingly many who work in the Ministry of Ed also held MEds and PhDs from abroad! (see the pic of them put in good use: reading my fortune from a Turkish coffee cup)

2/ Since the collapse of Soviet Union, many Russians have moved to Cyprus. I was told that even Putin has his own datcha there. According to Pavarotti (taxi driver), they find it as a small paradise. (maybe it is because of the 30% of votes that the Communist party still receives - I didn't know that either!)

The standard of living is rather OK, not too expensive, the weather is good, and the Cypriot men seem to be crazy about the glamorous style of Russian women, you know, the bells and whistles on high-heels.

I met quite a few Russians there, some really smart ones working in the oil business (apparently for tax reasons some businesses have re-located there), more working as waitresses, and then also the lot earning living with what they got. People told me about the problems raising of these gals on loose - the small society is facing issues on many levels; on the family-level it's resulted in broken marriages, money spent on prostitution, and in schools one also faces issues that were not there a few years back.

Even though Cyprus has been occupied in many occasions, the topic of Russians came up directly and indirectly in quite a few discussions. Interestingly, one person said that instead of seeing influence coming from the EU, now that they are members, they only see issues stirring up with Russians and Asians. I also noticed quite a few Asians in the country. I was told that many work as maid at homes and hotels.

3/ They also have mountains in Cyprus! Almost 2km! One can ski in the winter time. A new destination to be added on my list of strange places to go skiing :)

Thursday, May 24, 2007

Massively multiplayer object sharing by R.Sinha

I've followed some stuff from Rashmi Sinha, and I think every once in a while she comes up with good ideas. Like I liked the stuff early on that she did on the recommenders and the focus on user-centric design. Sometimes I just don't like her stuff, it sounds very popularistic and her references are, well, not very academic. But then again, maybe she does not need to be either..

Anyway, this slideshow has cool ingredients. I like the idea of object/artefacts in the center of the social networks, that's why I'm a big fan of social bookmarking, for example. I really don't care that much about connecting to people that I don't know (mySpace) or even using LinkedIn (what's the point, you get a list of people, but no substance..), but when I can connect through items and tags to people's stuff that I find interesting, I find it useful.

In the slideshow Sinha talks about models of 2nd generation networks (the 1st g was only about people):
  • Model 1: Watercooler conversations
    (around objects e.g., Flickr, Yahoo answers)
  • Model 2: Viral sharing (passing on interesting stuff, e.g., YouTube videos)
  • Model 3: Tag-based social sharing (linked by concepts. e.g., del.icio.us)
  • Model 4: Social news creation (rating news stories, e.g., digg, Newsvine)
Then, further on, she talks about Cognitive Diversity, which I also find really important. It's related to the continuum of wisdom of crowds vs. stupidity of mobs. What I got out of the slide 29:
  • Good answers need many perspectives, thus many perspectives are needed otherwise groups become too homogenous, which might have its dangers also (stupidity of mobs, see Digg for that ;). If all the new members are too similar and like-minded, they don't bring anything new to the group (that's why we want serendipity from recommenders!). Diversity reduces groupthink (think of Digg again and how fast not favourable stuff gets buried), groupthink is bad and only way to fight that is diversity.
Moreover, she also talks about the importance of social influence condition and about Watt's study.

Lastly, some design principles:
  • Make system personally useful: For end-user system should have strong personal use; Self-expression (e.g., Newsvine);Social status: Digg
  • Don’t count on altruism: System should thrive on people’s selfishness


Wednesday, May 23, 2007

Vanderwal: Tagging Today & Tommorw

Vanderwal nicely captures the spirit of tagging both for personal use and for the social aspects of it.

Personal      Social

- capture      - share
- hook/copy   - point
- annotate     - collaborate
- refind        - filter
- privacy      - trusted groups


Tagging Today & Tommorw - New Content, on the page 32
http://www.slideshare.net/vanderwal/tagging-today-tommorw

Anyway, the slides are worth taking a look. I'm glad he mentions "i" word - interoperability :)

Social influence on the convergence of tags

I'm a big fan of finding out how the social conditions influences the tags. Well, that is why it is one important aspect of my PhD studies, so we will find out - sooner than later.

In order to design a good tool for social tagging, I did quite a lot of research on our current understanding on this. I also came up with the table below, as I was mulling over what would be the best balance between giving social cues to people while tagging (e.g. showing the tags from other users) and still keeping the tags individual, personal, and specific enough.

See, I think too much of the social condition will make people lazy to come up with their own tags, thus the tag base will become less and less descriptive and consequently it becomes harder to find items tagged with these terms. Chi et al (2007) find that out too, but they do not contribute it to social influence! Just to the size of the community.

I think social cues are important for the uptake of tagging, in the first place, and in the second, they are important in terms of using common vocabularies, e.g. seeing those broad folksonomies to emerge. But too much can be too much! Thus, we are planning an experiment on tagging interfaces, where half of the users see the previous tags (social influence condition) and the other half does not (independent condition). We will study the tags from different aspects, how they converge, their originality and some other...not sure yet.



High convergenceLow convergence
Social influence conditionPrompt tags from other users (1) ; high influence from the community (2), better uptake as users may be more motivated to add tags; (3) maybe all tags become the same.
Use of existing tags becomes a habit, not many new tags are created as the time goes by (4); tags become less and less descriptive and consequently harder to find items.
Independent condition (no guidance)Even without social influence, when many users tag popular items, usually broad folksonomies startemerging (4).Original and intuitive user generated list evolves; however, low convergence of tags (1), lower uptake of tagging (3)

Table 1: Table presents two axes that affect on tagging habits and
convergence

(1) if people see other's tags (e.g. they are proposed, are auto filled when typing, ...) while they are tagging, vocabularies are more likely to converge than if users are working on their own.

(2) Also, users, who view tags by other people before adding their first tag, are more likely to have their tags influenced by other taggers in the community. The community of other users affects a user's personal vocabulary; there is a strong influence on user's first tag, if they have been exposed to others' tags.

(3) There were overwhelmingly more non-taggers in the group that had not seen examples of tags than in the one that had seen them in their tagging interface. As stated above, pre-existing tags affect the future tagging behaviour.

(4) When looking at how the vocabularies evolve while tagging, it was found that about half of the tags used were tags that the user had previously applied; thus, it was concluded that early habit and investment influence tagging behaviour and grows stronger as users apply more tags.
However, the research shows that habit and investment aren't the only factors that contribute to vocabulary evolution.

Well, there was another point, somewhat related to this, that I talked about with a studdy-buddy of mine: what is a good tag? Firstly, it is important to note that tags do have two main functions: one being a PKM element and the other is the sharing.

So, for the first one, any tag is good, if it makes sense to the user.
For the second category, we thought of different metrics; it could be for retrieval purposes or sharing with other people.

For retrieval, for example, a tag that repeats the terms in the title is not very good, especially if the search looks at the title anyway (like in our case we do have LOM already). So one could say that a good tag has some additional information that we do not have in the metadata already. After all, if we talk about social bookmarking of websites or research papers, there is some metadata already available. A tag that is a synonym of the title, however, can be useful for retrieval purposes, if any automated way for understanding synonyms are used (like a thesaurus that knows the relations of different terms).

A good tag could also hint something in the use of that content, for example in digital content, it could be something that would hint how that content could be used.

Chi, E. H. and Mytkowicz, T. Understanding Navigability of Social Tagging Systems. In Proceedings of CHI'07, February, 2007.

Sen, S., Shyong K., L., Cosley, D., et al. (2006). tagging, community, vocabulary, evolution. Proceedings of CSCW 2006. Retrieved from http://www.grouplens.org/papers/pdf/sen-cscw2006.pdf.

Vander Wal, T. (2005). Explaining and Showing Broad and Narrow Folksonomies :: Personal InfoCloud. Blog posting. Retrieved November 13, 2006, from http://www.personalinfocloud.com/2005/02/explaining_and_.html.

Tuesday, May 22, 2007

tag: deserted island

What if you were going to be out of the reach of the Internet, email and phones for 3 weeks, what books would you want to read? Oh, and let's add that you would not be able to pack many books with you, only a few. What would you take with you? Tag it with , please!

I was sure that I can just turn to LibraryThing or del.icio.us to find books that people really, really like, the kinds that they would take with them in case they were to be stranded on a deserted island for 3 weeks. According to my logic, you would tag that book with tag desearted island, desert island or something like that.

To my surprise I found only a few things for books! Some people had done lists of music and films that they would take with them, but I want books. After all, there is my iPod. (hmm..audio books anyone??)

So please help me! Tag a few books that you would take with you on a deserted island using the tag of , and that will help me to my choice.

Oh yeah, and did I mention that I will spend about 3 weeks on a sailing boat in the South Pacific? Apart from sailing, we'll be diving and hanging out in numerous atolls situated between Cooks Island and American Samoa. It'll be with D&D on Confetti. I'm getting pretty exited, I must say :)
Countless travellers' tales, books, plays and films have created a vision of an archetype of heaven in the South Seas -- massed coconut palms, jungle-clad peaks, the boom of combers smashing on the reef, the crimson flamboyant trees and the beat of the drum dance. Amazingly enough it is all true. Word66

That'll be the from June 10 to about July 5. So don't expect me to answer any emails. Try a message in a bottle?

Monday, May 21, 2007

Interpersonal networks in finding information

A hugely interesting study on patterns on information seeking about culture. I wonder how much this would match with what teachers do? Are they also inclined first to turn to their interpersonal ties, e.g. human network of colleagues, friends and families, to find information, before turning to the Internet, text book publishers, educational portals and such?


When searching for information about culture, the participants in this study look first to their families and social networks, specialized governmental and non-governmental organizations (such as Heritage Canada or the Danny Grossman Dance Company), and published and broadcast sources (Toronto Globe and Mail; People magazine; CBC radio and TV). It is only after they have a recommendation or suggestion—from their interpersonal ties or from elsewhere—that they turn to the Web for information. Then, they usually seek specific information, such as upcoming performances by a favorite band, book reviews, or hotel prices for a summer vacation. This suggests that for many people, the Web tends to satisfy curiosity rather than inspire it.

Yep, seems like supporting social information retrieval thorough Web is like a killer-ap!

Kayahara, J., and Wellman, B. (2007). Searching for culture—high and low. Journal of Computer-Mediated Communication, 12(3), article 4. http://jcmc.indiana.edu/vol12/issue3/kayahara.html

Sunday, May 20, 2007

A kid-Ceo designs an educational game

A kid designing a game for other kids, that's not SO new. But this kid designs an educational game for other kids and plans to make a million out of it by middle school, that is in a year's time! This kid is pretty amazing! I have no doubt that he won't do it. After all, he's already in the Wally.



Makes me think if we got it all wrong with all the e-learning things and going digital..

Saturday, May 19, 2007

Are digital bohemians bohemians among digital bohemians?

Yo! The BlogWalkEleven took place yesterday in Amsterdam. Theme: Digital bohemians. It was inspired by some German book on the topic of young people making their ends meet by earning a € here and there while leading a digital, networked and somewhat vagabond lifestyle. Of course, part of it is glamour, some are real addicts to the lifestyle, and the down side is the exploitation of young talents with short term contracts, issues with pension fees and how to live on a shoe string while trying to find a next freelance contract.

The idea of BlogWalk is inspired by concept of Open Scpace. Heard people talking of un-conferences? That's the same idea. no PowerPoints, no formal schedule, everyone contributes and cross-pollinates the conversation.
In Open Space meetings, events and organizations, participants create and manage their own agenda of parallel working sessions around a central theme of strategic importance..

Is digital bohemian a trait of character, a mindset, or a definition of an individual in relation to her/his surroundings?


So, the central theme was Digital Bohemians, which I found somewhat detached from in the first place. Towards the end, I'm sorry to say, but I left with an impression that we did not really "walk the talk".

The fun part of all was that the ~30 people around were all pretty inspiring and fun to talk to. They all had clearly anticipated an event with lots of interaction, so they were eager and ready to ask questions, talk to you about your interests and theirs, and just hang out. The first part of the day was fun, we all hurdled around the "window wiki" with post-its to write down our keywords, lines of thoughts and concerns on the topic. There was some geniousity on those post-its, I'm looking forward for Ton to sum'em up.

We had a lunch and - yeah, finally we physically walked in the city as a group! See, the whole idea of the thing is that you are able to find or initiate a discussion that you are interested in. If not, tant pis, walk on or do something. In this thing it's not up to the organisers to entertain and court you, but yourself! Walking is an excellent exercise for that.

Walking part was fun, but I must say that a sit-down lunch was not my idea of this type of organisation. I love cocktail parties for the part of being excused on the fly to change the group. Once you sit down, the "sofa-magnet" starts sucking you in and it's hard to find an excuse to move about 5 people from the same bench to allow you to change the scene just because you are bored with them.

So, that is exactly what happened in the afternoon. People sat down to start the second session, and nothing moved on from that point onwards. I was sad not to have my laptop with me (not many did, and get this, there was no wifi around! radical!), it's a perfect escape route of boredom and has become somewhat acceptable, too.

Anyway, my 2 €cents on digital bohemians: Are bohemians bohemians among bohemians? Is digital bohemian a trait of character, a mindset or a definition of an individual in relation to her/his surroundings?

I think digital bohemianism can best be defined by the relationship to the surrounding behaviour and conventions of practices. It's easy to point a finger to someone saying, see that one is a digital bohemian, if they do something that most people are not doing or do not want to do, e.g. live their life out of their laptop/PDA, have non conventional ways of getting their bills paid, know how to navigate in different spaces (physical and digital), are networked around the globe, etc.

However, if you have a flock of digital bohemians together, they cease to be bohemians among themselves, as they all pretty much sing the same cord. Of course, still, in the relation to others surrounding them, they would remain digital bohemians.

So, finally, maybe rather than defining and classifying digital bohemians, we should just attach tags to it and allow its folksonomic, non-exclusive base of terms to flourish just like bohemians do. Why tags are great is that they allow clustering "the thing" with many other things too, rather than having it sitting in one place in the catalogue or classification scheme, like we used to know them from library. There, I finally was able to tie it up with folksonomies, I'm getting really good at this :)

Anyway, I was glad to be part of this social experiment of BlogWalkEleven and I'm super glad to have met all these people! Let spaces be open in the future too!

Wednesday, May 16, 2007

Tomorrow BlogWalking in Amsterdam

Pretty exiting, high expectations - how could I describe the anticipation better than that? I've known Sebski for a few years now, and what brought us close in the first place, was the dislike of seminars. Then, back in the day, we were stuck in Turkey.

Sebastian told me about this undefined group of people who would piggyback any conference or other happening to get together and just walk around and talk about things. Things that matter and are important to this group of people who happen and choose to be there at that time. Things that they care about and feel passionate about. Related to networked technologies and such. I took all that in like a little gal (yes, I'm still a little girl) - and felt a bit jealous not never have been part of it.

Well, you live and you learn - tomorrow I'll be wondering around Amsterdam with BlogWalkEleven.

No rules but one - you are responsible for keeping your self interested, no-one else will do it for you. This is so cool. Will most likely let you know more about it - if worth..



http://blogwalk.interdependent.biz/wikka.php?wakka=BlogWalkEleven

Tuesday, May 15, 2007

More semantics to tags - emerging trends in tagging

Sometimes a word is not enough, so many more words are needed to explain what is it that is meant by that word.

That seems to be the case with tags. Uumh, they are ambiguous, did you know that?

First I think it was Technorati, they came up with the concept of WTF .
WTFs are short blurbs that explain the buzz around people, things, or events—why the hot topics are so hot—and you can vote the best ones to the top.

Now I saw it in del.icio.us, you can create a tag description to explain to others what is it that you really mean by your tag (it's really getting from being personal to being public). Very interesting. You now add metadata to your tags (e.g. meta-metadata) to describe the meaning by an author of that tag.

Well, if this trend keeps going, soon we can expect to fill-in all the LOM metadata fields ;) No, seriously, Sir Berners-Lee must be exited about this turn!

Monday, May 14, 2007

Notes on Everything is Miscellaneous interview

A pretty good interview on S.Weinberger's new book Everything is Miscellaneous with a Yahoo! guy Bradley Horowitz. It's actually way too long, he's a verbose guy, for sure, he's elevator pitch in the beginning would need an elevator ride in a skyscraper and back, and that would still not be enough! Some interesting bits to dive in or for fast-forwarding.
  • An interesting discourse over a unit of knowledge; he talks about group knowledge and how, for example, Wikipedia discussion pages or some mailing lists themselves are a unit of knowledge build through discussion, agreement and disagreement, and sometime arguments, too. Non of those individuals would not be able to construct that alone "the knowledge is not in anyone's head, it's literally in the discussion in the mailing list (28mins)
  • It was funny when they talked about Justin TV, how this guy carries a camera 24/7 and records his life. The interviewer comes up with something like what's the point, "you don't get a second life with which to review the first one (about at 31 min). This part is also related to gathering metadata about everything, in MIT they record individual's heartbeat throughout the day, so that you can later check when the heart rate was high, remember the moment and relive it! Dudes, get out!

  • Towards the end the whole thing gets more interesting. In about 45 min they talk about social filtering and Mr. Weinberger is very pro, he argues that it is hard for us to know what we are interested in. He goes "...the serendipity thing: we don't know what we are interested in. There are things that we can predict we are interested in, but largely not. The world is way more interesting than our interests, which is why social filtering is so important." I like that :)

  • Soon after they talk about the definition of discussion (48min) which makes me think of a lunch discussion with Teemu and his group in Medialab on how computer science has banalised terms like dialogue (a pop-up "yes" or "no"), interactive, etc..

  • At the very end (54min) they talk about folksonomies, and he says about how silly it would be to replace a taxonomy with a folksonomy. His argument was that "you don't want just one folksonomy", but many of them so that you can cluster and gather things so that it is relevant to what we care and interested in. That's good too!

Friday, May 11, 2007

1st Workshop on Social Information Retrieval for Technology-Enhanced Learning

I hereby proudly present the call for the first ever workshop on Social Information Retrieval for Technology-Enhanced Learning!

The complete call can be found from here: http://ariadne.cs.kuleuven.be/sirtel

A few words on the raison d'être of this workshop, what are the drivers for it?

Everyone in the field of e-learning has their ears full of talks of Communities of Practice (CoP) and networks of users, but not very often do we see how they actually are leveraged in practical terms. This workshop focuses on one part of the process, namely on retrieval of useful resources, either learning resources or human resources, for that matter. The tag line could be as P.Morville said "We use people to find content. We use content to find people."

Take that a step further and think of using digital traces to find people, and also leaving digital traces so that you can be found by other people. In this workshop we are interested in both; social navigation systems and recommenders for retrieving resources to enhance learning and teaching.

Social information retrieval (SIR) refers to a family of techniques that assist users in obtaining information to meet their information needs by harnessing the knowledge or experience of other users. Examples of SIR techniques include sharing of queries, collaborative filtering, social network analysis, social navigation, social bookmarking and the use of subjective relevance judgements such as tags, annotations, ratings and evaluations.

SIR methods, techniques and systems open an interesting new approach to facilitate and support learning and teaching. There are plenty a resource available on the Web, both in terms of digital learning content and people resources (e.g. other learners, experts, tutors) that can be used to facilitate teaching and learning tasks. The remaining challenge is to develop, deploy and evaluate systems that provide learners and teachers with guidance to help identify suitable learning resources from a potentially overwhelming variety of choices.

Several questions are being researched around the application of SIR methods in Technology-Enhanced Learning (TEL) settings. The aim of the SIRTEL'07 Workshop is to bring together researchers and practitioners who are working on topics related to the application of SIR methods, techniques and systems in educational settings, as well as to present the current status of research in this area to interested researchers and practitioners. It aims to serve as a discussion forum where researchers will present the results of their work, and also establish liaisons between different groups that are exploring related subjects. In addition, it aims to outline the rich potential of emerging SIR methods, techniques and systems in order to better build TEL systems and services.

Feel free to involve yourself, submit a contribution, blog about this, social bookmark the call (tag sirtel07) and talk about this to your pals!

See you in Crete in Spetember!

Monday, May 07, 2007

Workshop on Social Information Retrieval in Technology-Enhanced Learning (SIRTEL07)

Good news! The workshop proposal for EC-TEL 07 was accepted, so I will be co-organising my first workshop on social information retrieval techniques in support of learning and teaching later this September.

The tag line will be "We use people to find content, we use content to find people" by Morville. On the other hand, maybe it should be "We use digital traces to find people, and we leave digital traces to be found"..

Two main focuses: Recommender systems and Social navigation

The list of topics will be LONG, but I put it in here as an appetiser:

  • Defining the scope, purpose and objects of social information retrieval in TEL
  • Recommender systems and collaborative filtering in educational settings
  • Novel ways of generating input information for recommenders in the area of learning and teaching
  • Ranking of search results to support individualised learning needs
  • Folksonomies, tagging and other collaboration-based information retrieval systems
  • Social navigation processes and metaphors for searching information related to teaching and learning
  • Analysing social interactions in learning communities and social networks on the Web to facilitate information sharing and retrieval
  • Approaches to TEL metadata that reflect social ties and collaborative experiences in the field of education
  • Interoperability of SIR systems for TEL
  • Integrating SIR services in existing learning management systems
  • Visualisation techniques to support social navigation in learning and teaching
  • Semantic annotation and tagging for social information retrieval purposes
  • Evaluating the performance of SIR systems in educational applications
  • Measuring the effectiveness of SIR systems in supporting learning and teaching
  • Evaluation the user satisfaction with SIR systems in supporting learning and teaching

The idea is that as this is the first European workshop on the topic, we will try to scout out who are there to work on this topic and set the ground for better future collaboration . Of course we wish to run the workshop again, not as a pre-workshop , but really as a part of the main show.

Voila, more info to come shortly and the website for the call!

Saturday, May 05, 2007

The LibraryThing Recommender

I knew that LibraryThing.com had plans to work on a recommender for books, and seems like its out now. It's called LibrarySuggester; you can type a name of any book that you own or have read and the systems spills out suggestions in different categories:
  • People with this book also have...(v 1)
  • Special sauce recommendations!
  • Books with similar tags
  • Books with similar library subjects and classifications..
  • Amazon recommendations
  • People with this book also have...(v2)
LibraryThing Suggester analyses the more than thirteen million books and sixteen million tags LibraryThing members have added, and comes back with reading suggestions. Amazon suggestions come from Amazon.com, not LibraryThing.

Crowdsourcing

13 million books and 16 million tags, holy cow! That's some serious amount of data that people have free-willingly entered into the system! Just imagine trying to do the same before the day when the Web was crowdsourced. It would have taken an enormous amount of man-hours to enter people's likes and dislikes in books into a recommeder system as input to compute a list of recommendations, let alone the ratings, evaluations and discussions people have added too.

This is exactly the same way we want to go down with learning resources; first create a tool for teachers to create their favourite collections of learning resources and then use those to better serve them in terms of recommendations.


















Transparency

When I look at the recommendations from LibrarySuggester, what I like is that they are clearly classified in different classes of recommendations and on what those are based on. It is nice, as a user, to get the reasoning behind, e.g. ah, I was recommended this book because other "people with this book also have.." or I know that it is based on similar tags, etc.

This kind of practice of being transparent about the recommendations has also been argued about in previous research in the field, and it seems to be something that people appreciate, as opposed to a "black box" recommendations where the user has no idea on what the recommendations are based upon (Swearingen, 2001; Rafaeli 2005).

List of recommendations

Also, what I like is that LibrarySuggester offers a list of recommendations, as opposed to one or a few to choose from. However, in my list there were 74 recommendations all together, which I find way too much!

There are also some really evident ones, like books from the same author, which is not really a salient recommendation. McNee, et al. (2006) talk about a "similarity hole" that item-item collaborative filtering algorithm can trap users into by only giving similar recommendations. They argue that the old-skool accuracy metrics should be taken with a caution, as they only are designed to judge the accuracy of individual items and not the list of items. Thus, "the recommendation list should be judged for its usefulness as a complete entity, not just as a collection of individual items."

Moreover, within the same framework, which is called Human-Recommender-Interaction, these folks talk about three aspects that should be improved in recommendations. They are similarity (discussed above), recommendation serendipity, and the importance of user needs and expectations in a recommender.

Serendipity

Take the list of "Special sauce recommendations" for Dune by F.Herbert. On the list of 20 books you can find on the top 2 of his other books, and 2 by B.Herbert, his son. This sounds rather dull and not really anything surprising, you could find that easily from a bookstore too. By serendipity, the authors mean how unexpected the recommendation is for the user and how novel it is. For me personally this is a very important factor and why I like the idea of recommenders, as opposed to just content-based retrieval of resources.

I won't discuss the importance of user needs and expectations in a recommender, as in this case it is pretty clear. In some other cases, though, like for learning resources, this comes very important, as teachers do have different tasks at hand when they are looking for learning resources. This is something I've blogged before about and will keep exploring in my context of research.



McNee, S.M. , Riedl, J. , and Konstan, J.A. (2006) "Being Accurate is Not Enough: How Accuracy Metrics have hurt Recommender Systems". In the Extended Abstracts of the 2006 ACM Conference on Human Factors in Computing Systems (CHI 2006) [to appear], Montreal, Canada, April 2006

Rafaeli S., Dan-Gur Y., Barak M. (2005), “Social Recommender Systems: Recommendations in
Support of E-Learning”, Journal of Distance Education Technologies, 3(2), 29-45,
April - June 2005.

Swearingen K., Sinha R. (2001). , “Beyond algorithms: An HCI perspective on recommender
systems”, ACM SIGIR 2001 Workshop on Recommender Systems, 2001.

Thursday, April 19, 2007

D.Watts on Social influence and popular songs

This study is pretty interesting: there were 14,000 participants who were asked to listen and rate songs by bands they had never heard of. The point was to study the social influence, i.e. how seeing cues from other people, like Top10 downloads, no of downloads, etc. would influence on people's choice.

The set-up of this study is pretty neat, the participants were sliced into eight parallel “worlds” so that participants could see the prior downloads of people only in their own world. Everyone started from the same line, zero downloads — but because the "worlds" were kept separate, they subsequently evolved independently of one another.
What we found....In all the social-influence worlds, the most popular songs were much more popular (and the least popular songs were less popular) than in the independent condition. At the same time, however, the particular songs that became hits were different in different worlds, just as cumulative-advantage theory would predict. Introducing social influence into human decision making, in other words, didn’t just make the hits bigger; it also made them more unpredictable.

Our experimental design has three advantages over both theoretical models and observational studies. (i) The popularity of a song in the independent condition (measured by market share or market rank) provides a natural measure of the song's quality, capturing both its innate characteristics and the existing preferences of the participant population. (ii) By comparing outcomes in the independent and social influence conditions, we can directly observe the effects of social influence both at the individual and collective level. (iii) We can explicitly create multiple, parallel histories, each of which can evolve independently. By studying a range of possible outcomes rather than just one, we can measure inherent unpredictability: the extent to which two worlds with identical songs, identical initial conditions, and indistinguishable populations generate different outcomes. In the presence of inherent unpredictability, no measure of quality can precisely predict success in any particular realization of the process.

This makes me want to test and set up experiments, too. In the project that I'm part of, and where I will get my data, we are planning some experiences on the input part of the tags to see how social influence in terms of seeing other users' tags when inserting own ones, will effect on the nature of tags, their number, their convergence, etc.

But, on the retrieval side of things this would be very interesting too! To have two different interfaces to see the search result list, where on the one there would be all the social cues for social influence (no of downloads, no of bookmarks, others' tags), and on the other one there would be nothing. The experiment would test whether the users, in this case teachers, would be viewing the metadata of similar resources and what would they actually download, bookmark and rate, if they did any.

Well, actually the latter is the situation as it is now. So maybe I can just compare the data from this year and the year after, when we actually start implementing the social navigation part.


Link:
In NYTimes

Science 10 February 2006:
Vol. 311. no. 5762, pp. 854 - 856
DOI: 10.1126/science.1121066
http://www.sciencemag.org/cgi/content/full/311/5762/854

Supporting material:
http://www.sciencemag.org/cgi/content/full/311/5762/854/DC1

Thursday, April 12, 2007

What tasks teachers have on a learning resources portal?

The attitude of "if we build it, they will come" has resulted in national and regional learning resources portals where the offer, no matter how many learning objects or assests, does not necessary match the need of teachers. Why is that? What is it that teachers look for?

Curriculum coverage

I first started by looking at 29 different learning resources portals that national and regional educational authorities offer for K-12 teachers in Europe. My task was to find out how many of them offer curriculum related material, i.e. so that a teacher, knowing that tomorrow he has to teach an area that covers a certain goals of the national curriculum, can just go to the portal and pick a resource that actually goes through this particular area.

This type of standards-based curriculum material seem to have been on the offer in the US for some time. Two examples could be the DLESE http://www.dlese.org, focusing on Earth Science and IDEAS, http://ideas.wisconsin.edu a repository held by Wisconsin educators. In both teachers can search for a curriculum coverage.

I found that in about 10 out of 29 examples in Europe teachers can explicitly look for curriculum coverage on learning resources, however, it was not always very clearly indicated which goals or skills a resource aims to attain.

On rest of the portals teachers were able to find resources that were categorised by the school level and subject, thus, by using teachers internal knowledge of curriculum, they would, with little poking around, find the matching material. But the problem with this case is that there is nothing that makes the teachers' tacit knowledge externalised for other users, no one else can take advantage of the fact that this teacher knows how well this piece of learning material covers certain areas and goals of the local curriculum.

In MELT, we are trying to get to the core of this problem by encouraging teachers to add tags to resources that they find in the repository. We'll be eager to see whether tags can externalise any of that tacit knowledge that teachers have, and that could help other teachers in finding and using the material.

In Calibrate, another project that I have a minor involvement, we have an opposite approach to the problem. A few topic areas have been selected where 3 EU-countries share a similar curriculum. We try to map the different curriculum through its goals and skills, and match that with the material available. A complex task!

Tasks at hand

Then, I started thinking whether learning material that clearly covers the standard-curriculum is what teachers want? One could think that a busy teacher would really appreciate it, however, there might be other reasons why do they come to a learning portal. So I made a little poll with some tasks that I thought teachers could have, and asked some teachers to vote. I only got 36 votes so far (you can still vote here), but it seems that, at least for these teachers, the curriculum coverage was not what they preferred!









The majority of the respondent teachers (I don't know who they are) seem to be leisurely window-shopping while at the learning resources portal (I browse around to see what is available) or looking for material that could support the lesson that they are planning. Only one teacher is looking for learning resources with an exact curriculum coverage!

This rises two questions in my mind: either teachers have given up to look for curriculum covered material, because a) it never was on the offer, b) it was never successfully on the offer or c) they have better sources for that, like school books, etc.

Secondly, maybe teachers don't really know what to expect from a learning resources portal or a repository. Maybe it was never clear for them what was the intended goal of a learning resources portal and they just come there to see what's up, what's in there and maybe they return if something good is found.

Seriously thinking, do we even know for what tasks and goals the resources repositories are build? If we think of libraries, we know they have a clear goal, or a school book, yes, it has a clear goal. But a LOR (for learning object repository), isn't it something in-between, kind of pretends to be one or the other, without being either of them.

I just looked at the sneak-preview for Yahoo!'s new service for teachers. Pretty neat. They clearly aim it to be to create learning resources, and re-use the ones from other teachers. They also offer neat peer-networking possibilities. It's about using and re-using material that is in the center, not searching for it! A very different focus.

So, I think what should be build in on learning portals is a better support for the tasks that teachers and learners have at hand. So, when they come to a portal,
  • if they clearly are just window-shopping, let's provide them with first-grade Champs-Elysees shop-windows. Make nice pre-views available of resources that are there with added value lesson plans and case studies how others have used them in their lessons. Allow browsing other users' collections of learning resources with annotations and comments. These other users can be from any part of the continent, as the goal is to inspire and show how things can be done. If we know anything about the user, let's try to match them with like-minded peers!

  • if they look for material with curriculum coverage, let's lead them to an area where they can either search for material with curriculum coverage, or browse bookmarks and tagged resources for cues from other teachers on attaining certain skills and goals. For the latter, it could be useful to first show "traces", e.g. bookmarks, tags, pedaggical annotations, from teachers who come from the same area, like Yahoo! peer-network allows getting close to teachers from the same area (=same curriculum).

  • etc.
The point that I'm making is to be clear about the task, and the information seeking patterns in general that teachers have, and then match it with the best way providing search, social navigation, recommendations, shop windows, etc., but always thinking what would yield the benefits for the user and task at hand.

That's something to study deeper, and don't worry, I'm on it ;) We don't know yet if the best benefits can stem from using underlying social networks (location) or more implicit ones (profile, tags, similar bookmarks), or from using a search or what?

Wednesday, April 11, 2007

Social Navigation as seen a decade ago

I'm always fascinated when I find "old" papers that still resonate today. Well, I'm not talking about the manuscripts from the Library of Alexandria, but I just came across a paper on "Design Principles for Social Navigation Tools", written in 1998. The principals still seem rather relevant!

That makes me think; what was social navigation about a decade ago and how differently we perceive it today, after all, there was no social bookmarking nor tags out there, like we know them today.

Defining different flavours of social navigation

Social navigation can happen in many different forms... One may distinguish between direct and indirect social navigation [Dieberger, 1998, Svensson, 1998]. In direct social navigation, we talk directly to other users. In indirect social navigation, we can see the traces of where people have gone through the space, as for example in the Footprints system [Wexelblat and Maes, 1998].

In my context of use, direct social navigation could be seen to happen through networks of friends and colleagues, that the users of a social tagging system have established. Or, it can also be sort of "ask the expert", or ask your colleague type of thing. Direct social navigation could most likely take place in sharing pedagogical practices, for example.

Indirect social navigation, on the other hand, would be like following other users collections of bookmarks, browsing them through " xx other user have this in their collection" or through common tags, for example in the personal or common tag cloud.

Furthermore, social navigation may be intended or unintended by the advice-giver.. distinction can be made for when the advice-giver is one particular person, known to us, or when it is just a group of anonymous users that have happened to navigate through the same space as us. In-between these two extremes, we may have groups of users that are similar to the navigator in terms of interests, profession, knowledge or task. The advice-giver may also be an agent [Foner, 1993]1.

The idea of having intended or unintended advice-giver is intriguing, many times in collaborative environments for learning, we are very occupied in setting up advisors whose intend it to give advices. But in the real world, people are somewhat shy asking a specific "advisor" for hints or help, peers seem to work better, or many people look for "traces" for that purpose. So, it seem to me that it is very important to design a lot of unintended advice-givers to help social navigation in learning and teaching contexts, they can be anonymous crowds, agents, traces, annotations, hints of task orientation, shared interests, or what ever we can think of.

I wonder if this table makes sense like it is now: (I'm not talking about blanc spaces...)






Unintended advice-giverIntended advice-giver
Indirect social navigationTag cloud guiding the navigationA recommendation (person, agent,recommender,..)
Direct social navigationBookmarks from network of friends to navigateA friend, expert, peer gives advice or recommendation


The 6 design principles

The principles according to Forsberg et al. (1998) are Integration, Trust, Presence, Privacy, Appropriateness and Personalisation. The examples are pretty hilarious, really like from 10 years ago, but nevertheless, the principles stand like they should!

Integration:
It is emphasised that social navigation should be part of everyday tools to make best use of it. This is what we see nowadays a lot, more and more things are integrated in our workflow, for example the browser with all the add-ons and blug-ins has become a central tool.

To make the best use of social bookmarking, it is of utmost importance to make the process easy and part of what one was doing right at that moment. The delicious-bookmarklet added in the browser toolbar is a good example, it is so easy to use that you hardly even have to stop what you are doing at that moment to bookmark. Thus, you leave more traces for others, besides arranging your own information space.

The idea of Attention Metadata is another one, if you use Slogger to generate Attention Metadata on what all you do on your web-browser, you don't even have to think of doing it.

Presence:
This refers to how do you make the presence of others shown to users, how do they know that someone has been here before. Annotations (tags, comments, evaluations, opinions, ratings...) are a perfect way to do that (Kilroy was here!), you know that someone passed through that space.

Or showing how many other users are online at the moment, think of how much fun is it to log into Skype at 1am and see that despite the quietness of your work room, some other people are out there still slaving away.

Trust:
In order to take a note of someone else's doings, it helps if you know whether you can trust them, are they a reliable source, do they like the same things as you do, etc. When reading film critics, one quickly picks up the critic who is like-minded, and disregards the other one who always seem to have too much of a mainstream thinking. Similarly, to trust the source for online social navigation can become crucial, if there are many ways to go forward.

Appropriateness:
In some situations one design choice is more appropriate than the other one, for example indirect social navigation can be suitable for finding inspiring information, whereas you might rather turn to someone to talk to when you have a specific question at hand. This is in my opinion related to finding a means that fits the purpose of the information seeking task at hand, something that I've been mulling around a lot and is related to the Human-Computer Interaction framework.

Privacy:
This concerns the issues of making users aware of the traces that they leave, data that is logged and used for personalising their searches, etc. This is also related to some codes-of-conducts that some systems have. Very important for things like Attention Metadata, Google personalized, etc.

Personalising Navigation:
"Social navigation provides excellent opportunities for tailoring navigational advice to individual user's task, knowledge or abilities". How I see this, it is related again to the information seeking tasks, interests, etc.



Forsberg, Mattias and Höök, Kristina and Svensson, Martin (1998) Design Principals of Social Navigation. In: 4th ERCIM Workshop on User Interfaces for All, Stockholm, Sweden.

@InProceedings{sicsprint95,
author = {Mattias Forsberg and Kristina Höök and Martin Svensson},
year = {1998},
title = {Design Principals of Social Navigation},
address = {Stockholm, Sweden},
url = {http://eprints.sics.se/95/01/designprincip.pdf},
booktitle = {Proceedings of 4th ERCIM Workshop on User Interfaces for All}