Friday, January 25, 2008

Visualising tags from LORs

This tagcloud is comprised of all most used tags in 3 repositories.



Whereas this one is cleaned from tags that are shared with a specific community of users within one repository. Much better?



Or, should the tags be displayed by the biggest number of users, not how many times they have been applied?

Tuesday, January 15, 2008

You gotta be careful what you wish for: data portability group

I just read about the Data Portability group, or rather just watched the video (funky music!) on Read/Write web. Philosophy, like explained on http://dataportability.org/

Philosophy As users, our identity, photos, videos and other forms of personal data should be discoverable by, and shared between our chosen (and trusted) tools or vendors. We need a DHCP for Identity. A distributed File System for data. The technologies already exist, we simply need a complete reference design to put the pieces together.

Mission To put all existing technologies and initiatives in context to create a reference design for end-to-end Data Portability. To promote that design to the developer, vendor and end-user community.


A year ago, I posted about Fighting read/write web fatique. So, that's what all the traffic on my blog was about ;) (as if...)

Microsoft-free life is over for me?

When I was younger I always hated it when older folks could not remember when thinks happened. They would pause in a middle of a story and start "well, can't remember if it was '65 or '67...". I always thought that I would never do that, I always remember when things happened - or well, I used to.

So I had to go back to some old files on my old computer to find out when was it that I actually started my life without Microsoft. It was around the end of 2003 when the operating system on my ThinkPad crashed and did not boot anymore. Luckily I had a Mandrake distro installed on a separate partition, so I was able to access my files and go on working.

That incident gave me a good kick to keep experimenting with Linux on a desktop. It was quite a struggle in the beginning, as I refuse to use the command line and want to manage everything using a GUI. Installing new software was always a struggle, some dependencies were always missing, and I could never even get the installing thing working in Ubuntu. Yack.

Anyway, the point of the experiment was to know if an average Bob or Mary could use Linux on a desktop. At the time I was writing a lot in EUN about the use of open source software in education (see some reports here, I still like this one). I truly think they could. Especially when in many schools teachers and pupils cannot even have admin rights to the computer that they use, non of the installing issues would come up for the end-user. Using OO, Firefox and such is all jeu d'enfant.

Back to the title: a week ago I installed MS Office on my Mac. That's it, they won, after all these years! That's more than 4 years without Microsoft (i.e. Linux and Mac), which has been a real source of joy for me. I cannot articulate all points why I so dislike Microsoft, I think deep down it's the idea having one big guy on the playground (especially if it's not me ;).

Anyway, I'm giving it a try for a short while and see if it enhances my PhD writing. Maybe, now that I'll have all the fancy words in the Thesaurus (must admit OO's not the best), and I can start sounding much more intelligent and I'll get my PhD just with a little spin!

So far I'm not too thrilled. It crashes all the time, at least 5 to 10 times a day. It also uses all my RAM (I have 2 Gb now) and I cannot have all my usual apps open at the same time with Word and Excell without running real slowly. Oh yeah, and the user interface is god dam busy, I much prefer the serenity of OO.

The reason I gave up was that my OO, after the OS X upgrade, did not work properly anymore. I did not recognise the changes in the virtual keyboard (I switch between the Fi and Us all the time), and the spreadsheet did not recognise the control click (right click). Oh, and it was slow tooo! I later heard that my study-buddy has no problems with NeoOffice, but that'll be next on my list.

In Three Fearless Predictions (The Economist) talks nicely about "openness" in 2008 in terms of open source software on desktops (linux), iPhone being forced to open up in Germany and about how SCO seems to be gone ,once for all, with their silly lawsuits against companies developing Linux software. Maybe I'll get some of that openness too...

Sunday, January 06, 2008

My Photo in a travel guide: innovative use of Creative Commons

We often times hear about "innovative use of new technologies" or how "new licensing schemes offer innovative ways to use content". Unfortunately, too little of that is seen in the real life and too little of it comes to your way. This time I was happily surprised, though.

One of my travel photos on Flickr, which I always put a Creative Commons license on, was selected as a runner-up for an online travel guide. They contacted me through Flickr account and asked if I wanted to submit my photo into their Travel guide. Hell yes, I though, it would be fun to have one of my pics up.

My photo now is the main picture of Cafe Imperial, a cool Art Nouveau style cafe in Prague. I found the cafe a few years ago when in Prague, Matt and I had passed by it during the day and decided to go there for dinner. I had a really good Stroganoff, the one that has a true Russian flavour to it. It was a great dinner in a great place, and now I have a memory of it in a travel guide. Funny how things go!

Monday, December 31, 2007

End of the year account: travels

I decided to write an end of the year summary for myself about 2007. This is a personal account, so if you come across this, don't bother boring yourself with it.

It will include important things like my travels, studies and worklife. I'll do it so that I can look back and say, what the hell of a year it was. I bet there will be years that are not that gentle with me, so it'll be fun to dig this year up and read how things were once upon a time.

Start with travels, most important ;)

2007 was a good travel year. Like any wanna-be globetrotter, I managed to put my foot on a new continent this year. I also got to discover new pretty rare places and visited the usual ones, so it's a good count for 2007.

To start off, 2007 was going to be my EPIC ski year. I had a season pass to Winter Park in Colorado and I planned to spent some two months in Wyoming. After all, this was my last year on scholarship, so I had to make a good use of that freedom ;)

It turned out not to be an epic ski season. Not for the shake of snow this time, but for the fact that Matt broke his leg, pretty seriously, during our second weekend of skiing in Utah. Well then, I got some seven days in, though, and busted my dear snowboard after many faithful years of use. So, the best part of Q1 was spent in Wyoming and I also got to visit Utah, especially the LDS Memorial Hospital in Salt Lake City (Matt had his surgery there).

Beginning of June was the start of the great big sailing adventure (pics and blog). Oh my, that's gotta rank as one of the best trips ever. Matt and I took a flight to Rarotonga, Cook Islands, where we met with Dan and Danielle, who had sailed there from San Francisco via Mexico, French Polynesia and such.

We boarded on Confetti, a beautiful 44-footer, a hand-grafted wooden boat (Farr), to sail for about 1200 sea miles. We spent some 3 weeks on and off board exploring Rarotonga, Beverage reef (wow, even Wikipedia does not have an entry for it!!), Nuie and American Samoa, all small tiny places in the vast South Pacific.

We put in a total of seven-eight days of off-shore sailing (in 3 legs), which is by far the longest time I've ever been out on the open. And I really mean out on the open, hell, those "trade wind" areas, there is no one out there! During the first leg from Raro to Beveragereef and Niue, which took about 10 days, we saw no one, no boats, no planes, nothing. Just us in Confetti with our Ham-radio connection to weather forecasts and some necessary emails. That was some experience that I'll keep cherishing for a long time.

Another curiosity of this trip was the type of "civilization" that we encountered when we stayed on those islands. First, I gotta admit that they are something to dream about; beautiful white beaches, palm trees, thick green vegetation and so on. Just what you would expect. Sadly, though, the life that people lead on these "paradise" islands is far from my notion of paradise. It seems that "civilization" from missionaries, and early and nowadays merchandisers and explorers has dramatically altered the way of life, and globalization with the help of satellite TV and subsidized food from New Zealand has taken over any tiny bit that was left of the original way of living.

It's hard to put words on what we saw without sounding like a "conservationist". I don't want to say that people should still live like they did when Captain Cook and other discoverers first met them. What I want to say is that it hurts seeing such level of obesity, unemployment and apathy against your own nature and environment on some of the most beautiful spots on earth that I've ever seen.

Apart from the South Pacific, this year's new discoveries include Cyprus. I finally got to see Barcelona, it's as great as people say it is!

I also got two trips on my miles, which is good.

Wednesday, December 12, 2007

I'm not a happy bunny with Leopard, grrr

Oh my, just gotta say this out loud! I'm not happy to have the new Leopard running on my Mac, really not :(

It's not that I'm very picky about things not working smoothly, after all, before switching back to Mac I ran only Linux on my desktop for about a year and a half. It wasn't always easy for a non techky user like me (I still get pimples if I have to use command line...), but I wanted to give it a try. After all, at that time I talked a lot about using FLOSS in education. So to know what to talk about, I needed to do it myself.

About Leopard. Last drop today was that I can not get one of my favourite widgets working, namely Screenshot plus. They claim they became Leopard compatible, but bullocks, it still does not work. And that is not the only thing!

My Mail crashes about five times a day, luckily still saving the email drafts. And that is a Mac native application, what the hec?? I also use OpenOffice, probably the only individual in the world to do it on Mac. But hey, after all I bought a Mac 'cause I did not want to use MS programmes.

So, OO became REALLY slow and it does not support all the keyboard shortcuts or virtual keyboard language selections (I use Finnish and US keyboard and switch a lot). This makes my work really sucky.

It's not actually only OO that became slow, it the whole computer. I constantly use all my virtual memory, 1.5Gb, which does not seem to be sufficient for this hungry animal. All the searches within the computer are really slow too.

Yes, also the Java support is inexistent, and I have no idea how to deal with that.

Well, I feel much better now. Thanks god there are blogs to vent out... whew..

Monday, November 19, 2007

eTwinning/bookmarks and social networks

Excellent write up here about SNA on social networking sites. This makes me think of eTwinning, or my social bookmarks, and the SNA analysis there to better support users.

Kumar, Novak and Tomkins (200&) saw that network activity is of three types:
  • “Singletons,” who have no connections and are least central
  • The “giant component,” which is the largest group of nodes tightly connected to the central nodes and to each other
  • The “middle region,” which represents isolated groups which interact amongst themselves but not with the rest of the network, forming isolated stars. These groups grow one user at a time. Over time they merge with the giant component.




















The node analysis of these networks showed that more than half of a social network is outside the giant component where the greatest centrality lies. They used the “control” definition of centrality to determine this. The research also highlighted a prevalence of “stars” in the middle region which are mini social networks, typically driven by one dynamic member who serves as the point of centrality with others serving as satellite nodes – connected to the dynamic member but not to each other. In Kumar, Novak and Tomkins’ analysis the middle region represented one-third of users on Flickr and about ten percent of users on Yahoo! 360.

Also keep in mind that the most growth happens in the middle region where dynamic members influence others to join their network. These sub-networks can gradually join the giant component over time. Once they do, the importance of the dynamic member diminishes. Even if that dynamic member were to leave the network, the others would stay in the network.
So, what is needed is to support the "stars" in their growth so that they become independent of that one dynamic member and are able to continue even without that person.

Wednesday, November 14, 2007

Such a cool way to send a message, thanks B.Dylan!

Tagging in different e-learning environments

In the last days we've had a few discussions about tagging in e-learning environments. My environment, where the tagging takes place, is a portal for learning resources.

Today I came across this nice graph that displays "power law of participation". Now, haven't looked at the scientific background of it yet, so no comments on that. Anyway, it kinda rang the bell with what I'm doing when looking into levels of user engagement on the portal.

According to this graph, adding things to favourites (e.g. bookmarking) and tagging them represents a pretty low threshold to participate in the activities of that given community.




















I'm also looking at leMill environment, which on the other hand, demands a pretty hight level of user engagement, as it is about collaborative authoring of digital learning resources. About a year down with users, there is somewhat little collaborative authoring that actually takes place, Hans told me yesterday.

Maybe tagging in some way could help the participant to take the first steps? well, they can already tag and favourite things in LeMill, so maybe the issue is rather to see if similar levels of engagement appear in that community.

So, along with that, I am interested in looking at the tags in LeMill from the same point of view that I'm doing for tags in our learning resources portal. The difference is that our case is clearly what is called broad folksonomies, whereas leMill should be a rather classical narrow folksonomy. Or is it? Maybe once we start looking at those tags as a triple {user, resource, (tags)} with a timestamp on them, it appears that participants first start by bookmarking and tagging resources from other users, before the user takes a step to create her own resources and finally collaboratively work on other's resources.

Update:

So, some data to back-up was found:





















Social Technographics®
Mapping Participation In Activities Forms The Foundation Of A Social Strategy
by Charlene Li
http://www.forrester.com/Research/Document/Excerpt/0,7211,42057,00.html
with Josh Bernoff, Remy Fiorentino, Sarah Glass

This is a document excerpt EXECUTIVE SUMMARY
Many companies approach Social Computing as a list of technologies to be deployed as needed — a blog here, a podcast there — to achieve a marketing goal. But a more coherent approach is to start with your target audience and determine what kind of relationship you want to build with them, based on what they are ready for. Forrester categorizes Social Computing behaviors into a ladder with six levels of participation; we use the term Social Technographics® to describe a population according to its participation in these levels. Brands, Web sites, and any other companies pursuing social technologies should analyze their customers' Social Technographics first and then create a social strategy based on this profile.

Wednesday, November 07, 2007

notes on "Context, (e)Learning, and Knowledge Discovery for Web User Modeling: Common Research Themes and Challenges"

"Context, (e)Learning, and Knowledge Discovery for Web User Modeling: Common Research Themes and Challenges" by B.Berendt

This paper is about context and how to define it or how it is defined differently. The following is related to "Context in Web usage mining and eLearning"

2.1 Context as data and as metadata

"In order to evaluate whether intended and actual usage coincide or not, and in order to obtain a more fine-grained picture of actual usage, it is of course interesting to measure aspects of actual usage. "

- This is also one thing that we are interested to find out in MELT, and partly also in my PhD. As we have very little access to "actual use" we try to infer this type of information from usage logs. E.g. We have a teacher who has said in his profile that he teachers students from 12 to 13 year olds. If he bookmarks LOs that have intended audience of 14-18, we can maybe infer that this LO can also be used for younger students. Especially, if we start seeing this taking place a lot, we might want to update the LOM on intended audience: instead of 14-18 we could say 12-18.

- My interest is also to see if tags can give us any hints of this.


2.2 Context and model parts

"context representations can form and/or enrich (a) user models, (b) material/environment models, or (c) interaction models."

- EUN uses a) in one search to rank resources, but we are still only implementing it and we don't know how users react to it. That is related to my own PhD, as are how different search methods are used. In general, we do way too little with user modeling (I guess bigger issues are still more imminent)


2.3 Context: parameters of the (inter)action

- For my PhD I'm looking into user logs to create "levels of user interaction", e.g. what does it mean if a user views a page vs. makes a bookmark on it. We want to use this as an input for a recommendation system, for example.

- I'm also interested in the type of search that the user has chosen and its relation to the search task that the user has at hand.

- Tags were mentioned in this context, that is also a huge part of what I am looking. There are different questions around them, one most interesting related to search is how they can be used for discovering resources.

Need to look into these papers:

- B. Berendt, G. Stumme, and A. Hotho. Usage mining for and on the semantic web. In H. Kargupta, A. Joshi, K. Sivakumar, and Y. Yesha, editors, Data Mining: Next Generation Challenges and Future Directions, pages 461–480. AAAI/MIT Press, 2004.

- Claus-Peter Klas, Hanne Albrechtsen, Norbert Fuhr, Preben Hansen, Sarantos Kapidakis, L aszl o Kov acs, Sascha Kriewel, Andr as Micsik, Christos Papatheodorou, Giannis Tsakonas, and Elin Jacob. A logging scheme for comparative digital library evaluation. In Julio Gonzalo, Costantino Thanos, M. Felisa Verdejo, and Rafael C. Carrasco, editors, ECDL, volume 4172 of Lecture Notes in Computer Science, pages 267–278. Springer, 2006.

- Totally agree with his observation, not the method: Tanimoto [53] emphasizes that may be difficult to conclude, from a mere clicking event, that there was indeed attention paid to (specific) content of the requested page.


2.4 Context: background knowledge

Tags, tags, tags. multiple views.


2.5 Context: Activity structure
" This metadatum can provide important information about a visitor’s intention or expectation (e.g., whether they followed a prescribed link from a course page, or whether they found a material by actively searching with a very detailed search phrase)."

For me this is important, I guess using terms from this paper, I'm interested in user's intentions and expectations and finding out the ways the users choose to access or discover resources in our portal. I'm also interested in seeing whether one method is more useful to a given task, e.g. if people like browsing to find inspirational material and some other method (social information retrieval vs. information retrieval) for another task. If we know what kind of method is useful for a given task, I think we can help our users a lot.

2.6 example

An example is given using the three aspects of context; activity structure, parameters of the (inter)action and background knowledge. The type of analysis allows answering questions like: which search options are popular and are there differences between users? Which content areas were frequented, and how did people navigate between then; did they go back to the search options, or did they use the inter-content links? Did certain content areas become hugs for navigation and thus served to organise the domain and the presentation of the domain? On the other hand, questions like; were there differences between users with high verbal and users with high visuo-spatial competencies; did certain textual or pictoral material become hub?

These are also questions that I am looking at in lre portal and am getting a good idea of them. However, I have not been able to link them with the task at hand yet, which is something that I'm interested in.

Schooling for Tomorrow scenarios and Science-Fiction

Last autumn I heard a few e-learning keynotes with a heavy science-fiction emphasis. That was wild, I totally loved it. I can not agree more on the idea that working with future scenarios, or science-fiction for that matter, is actually helping the future to come along. It is about helping to shape the future after having peeked in to the future with a positive or a negative outlook. And, I'd like to say that looking or creating scenarios in many perspectives is useful too, as it helps you to see whether this is where you want to end-up or not.

There's been some work going on since the beginning of the millennium regarding scenarios for the future of schooling, we also in EUN worked on that. The OECD report Schooling for Tomorrow, is out too. It's way less exiting than talking about 2048 in a science-fiction scenario, I'm afraid, but still aiming at the same goal - see how technology and network enhanced learning could be used in the days to come.

Of course I picked upon the scenario called "Learning in Networks replacing schools"













This scenario imagines the disappearance of schools per se, replaced by learning networks operating within a highly developed “network society”.

Networks based on diverse cultural, religious and community interests lead to a multitude of diverse formal, non-formal and informal learning settings, with intensive use of ICTs.

How about that for science-fiction?

Well, if it is up to me, I would like to see learning networks in schools even if schools per se are not facing the extinction. You know, before we need to go to a dinosaur museum to see a replica of a teacher.

I don't mind the idea of "networked society", but somehow, when wearing those gray classes, I'm thinking of the efficiency of terrorist cells workings, and how religious and local interest groups could manipulate their own learning interests on me, while at the same time monitoring with whom do I want to learn on the international scale. Yak!

The Schooling for Tomorrow scenarios by OECD are a real tool set for policy-makers, and why not others, to work on. I like practical things like these are. Lately, I've been toying with the idea of making people, who work with education, technology and networks, to write science-fiction short stories of learning in the future. I also think that this would be a very helpful exercise for other PhD students to open up their thinking and not be stuck with what we got now.

How to get started with your own science-fiction story on Schooling for Tomorrow

It's important to understand that it's the journey that is important, not the destination. To say it in other words, I think the whole thinking process to come up with a plot for your story is what counts! Don't worry about picking up the right publisher now ;)

This is what I found out about writing science-fiction, how to get started
  • idea: the premise, the basic thought around which the story will turn
  • a setting: worl for your story to take place in; it can be familiar, or wholly new
  • Characters: two or three, at least, to people your story
  • Aliens: (optional) strange and mysterious beings for your characters to encounter
  • Problem: something your characters want, need, must escape from, etc.
Hmm, I really like to think of Aliens and education...

Next, when you have those sorted out, think about the five dimensional framework from Schooling for Tomorrow. What are:
  • “attitudes, expectations, and political support”,
  • “goals and functions” of education systems
  • “organisations and structures”,
  • the “geo-political aspects”
  • “the teaching force”.
Then, just let it flow!

My own attempt on this is in a wiki. I've already had quite a few engaging and hilarious dinner discussions with my friends about what should the plot be. We never got very far, but it's always been a lot of fun! I only wish we could somehow get the discussion transcripts on the wiki....

Tuesday, November 06, 2007

Notes on "Collaborative tagging and Semiotic Dynamics"

By Gattuto, C., L.Vittorio and L.Pietronero (2006).

Firstly, I must say that I was glad to read this paper. Lately, I've been seeing many papers talking about the properties of folksonomies, like co-occurrence, etc., which have intrigued me quite a lot. This paper explains the process pretty well and underlines an important point - they factor out the users and only deal with streams of tagging events and their statistical properties!

I must admit that this makes the whole area of Semiotic Dynamics less attractive to me. I think it is important to study tags and their properties, but not in isolation from the user. I see (barely) the point to explain tagging activity and the growth of tags in separation from the users. But fair enough.

Problem statement: Uncovering the mechanisms governing the emergence of shared categorisatioins or vocabularies in absence of global coordination is a key problem with significant scientific and technological potential. Collaborative tagging provides a precious opportunity to both analyze the emergence of shared conventions and inspire the design of large agent systems.


Semiotic Dynamics study how populations of humans or agents can establish and share semiotic systems, typically driven by their use in communication. The author argue that the emergence of a folksonomy exhibits dynamical aspects also observed in human languages, such as the crystallisation of naming conventions, competition between terms, takeovers by neologisms, and more.

  • Users interact with a collaborative tagging system by using tags or adding new resources to system
  • Basic unit of information in collaborative tagging systems is a (user, resources, {tags}) triple, which they refer as post in this paper. Tagging event is a tri-partite graph (with partitions corresponding to users, resources and tags, respectively) and can be used as a navigation aid in browsing tagged information
    • Comment: I like the tri-partite graph as navigation aid, yes!, but as the authors mention just above, they don't think of other users and those networks as navigational aid. In contrary, they omit the users just to study the properties, which strikes bizzarre to me.
The authors cite the "rich get richer" model (Yule-Simon's stochastic model) and propose to enhance it with a "fat-tailed memory kernel". This original model is related to the construction of text from scratch:
At each discrete time step one word is appended to the text: with probability p the appended work is a new workd, never occurred before, while with probability 1-p one work is copied from the existing text, choosing it with a proability proportional to its current frequency of occurrence. This simple process ields frequency-rank distribution that display a power öaw tail with exponent alpha = 1-p, lower than the exponent we observe in actual data. This happends because the Yule-Simon process has no notion of "aging", i.e., all positions within the text are regarded as identical ..
This all leads to a model of users' behaviour: the process by which users of a collaborative tagging system associate tags to resources can be regarded as the construction of a "text", build one step at a time by adding "words" (tags) to a text initially comprised of n 0 words. There is also that same Yule-Simon model with long-term memory (about inventing new tags or using existing ones), but recent tags are used more often than old ones.

Also, "in our model,.., the average user is exposed to a few roughly equivalent top-ranked tags and is translated to mathematically into a low -rank cutoff of the power law, i..e., the observed low-rank flattening".

Conclusion: It seems that users of collaborative tagging system share universal behaviour which, despite the intricacies of personal categorisation, tagging procedures and user interactions, appear to follow simple activity pattern.

There is also something about the co-occurrence between high-rank and low-rank tags: it says: "This suggest that high-frequency tags partition - or "categorize" - the resources marked by tags of lower frequency. "
Comment: This all sounds interesting and important, but will need to look into that later.


Monday, November 05, 2007

My PhD research

Notes on "Aspects on Broad Folksonomies"

Aspects on Broad Folksonomies by M.Lux and M. Granizer (2007)

This paper continues the trend in studying and analysing the underlying statistical properties of broad folksonomies that aims to identify laws and characteristics which allow inferring those properties. A few notes on what I found interesting related to the emerging notion of quality of tags, something that I've also spared a few thoughts on.

First, though, on some other issues. The paper talks about the emergence of power law distribution in folksonomies. They describe which approach they took to fit the sample to a power law, which was something that I've sometimes contemplated on the how-part of things. The paper aims at analysing whether one can find similar term distribution in folksonomies as in classical term retrieval (e.g. Zipf. note: Zipf's law with an exponent between 1 and 2). The dataset is that of delicious (uh, with about 800 000 bookmarks and about 27 000 users- I got a way to go with my MELT bookmarks).

Tag co-occurrence
They are able to show that "for around 80% of the tags of a folksonomy the co-occurring tags follow a power law distribution, which approves Cattuto's assumption. We found that for about 90% of the estimated power law exponent B xxx [-1.5, -0.5], which shows that for most tags co-occurrence follows a model with similar parameters. "
Resource and user based tagging characteristics
Secondly, they looked into frequently used tags (more than 30 users).
  • For resources statistics they (frequency of users tagging the resource with a tag) found that around 18,4% of resources followed a power law distribution.
    • assigned by lot of users to few resources (head) and to a lot of different resources by a few users (tail)
  • For user statistics (frequency of resources tagged with a tag), around 13% are following a power law.
    • few users tag a lot, whereas lot of users tag a few
  • i.e. the characteristics of the user statics are similar to the characteristics of the resource statics.
  • They argue that those tags, which follow a power law w.r.t users and resources are high quality tags (i.e. tags describing resources with high accuracy [no misspellings and meaningful tags] ) for most of the users involved in the investigated social bookmarking system.
  • A small fraction of tags have overlapping user groups, which points towards sub communities (user groups sharing the same link selection and tagging behavoiur) in the tail of the power law distribution.
    • this was found through splitting resources in 3 (high, mid and low rank resources)
They also looked at the big chunk of tags that were not following the power law.
  • Unique assignments. More than half (57%) of less frequently tags are used only once. They think that they can be seen as "shortcuts" for a user to a resource or a misspellings. They argue that these tags are useless from retrieval point of view (hmm..).
  • Personal vocabulary. especially in less frequently used tags (19%) of tags were only used by one user but assigned to many resources. They are useful for personal retrieval but useless for the rest of the community.
  • Unpopular vocabularies. between 1/5 and 2/5 of tags are assigned to different resources by different users only once. Unpopular vocs used by a small fraction of users.
  • they conclude that from retrieval point of view (e.g. inverted indices, TF*IDF) a large fraction of tags are good for single or sub-communities, and only the power law distributed tags are good for that.
    • They don't say anything about how to include the large fraction of tag not distributed by power law into IR methods.
Retrieval Aspects
Q: Do tags add information to further to description and title for retrieval purposes? This is a lot along the lines that I am also interested in, although I will look more into the networks of users. They say that for retrieval tags can be seen as an additional resource. Moreover, about 50% of available description contain information similar to the information described by tags, whereas the remaining 50% can be seen as orthogonal information.

Comment. This all is treating tags only as additional keywords that can be useful for conventional retrieval purposes. I think the connection tag-resource-user is more interesting. Just the fact that even if the tag is misspelled or hooks to a small user community is less important to me, because I know that the fact that this resource was tagged shows that the user has an interest to this resources, thus it is a vote. This aspect has an immense potential for retrieval (recommender point of view), but is seldom regarded in papers with very conventional retrieval approach.

Open social and education

I wonder who is going to come up with the first OpenSocial app or widget for educational use? We certainly are talking about it, for example for our eTwinning platform. It could be cool to be able to use information about teachers collaborative networks to allow, say, better retrieval of learning resources relevant for the project, purpose or task that teachers are undertaking; link with some other sources that teachers are working on through cool widgets, etc.

I never thought that Facebook, which has lately become really popular among my friends (not early adapters), would be the seul app that would "take it all". I was glad to read this:
"The market has already decided that there's going to be a long tail of social networks, and that people are going to belong to more than one. As soon as you belong to more than one, this kind of interoperability is critical," Dash says. "Open standards win every time." wired

Hurray for open standards!

Tuesday, October 30, 2007

Multilingual tags and the language of LO

I've looked into tagging in different languages before. An interesting thing came out of our little pilot: teachers, non of whom mother tongue was English, still had about 20-30% of tags in English. We had two different thoughts on this,
  • either tag is in English because teacher wanted to share these tags with other teachers, or
  • tag was in English because it is related to the language of the learning resources that was bookmarked
I was now interested in the second possibility, and took a look at a sample of 136 bookmarks with tags in multiple languages related to them.
  • The LOs were in English, Hungarian, Polish and Estonian.
  • The users (43) were Hungarian, Polish, Estonian and Lithuanian
Fair enough, all the English tags were related to the English resources! In close to 30% bookmarks (39 out of 136) this was the case (which also means that 30% of LOs were in English).

Moreover, it seems that for about 1/3 of the times the language of the LO was the same as that of the tag, whereas 2/3 of the cases it varies according to the language of the user. In about 3% of bookmarks one was able to observe multi-lingual tags.

Thursday, October 25, 2007

Radioheads making €, good for them

Cutting the record lable out of the equation seems to be a good deal. I'm glad to see the bold decision to go directly to fans has not only shown a great example, but also proved profitable! I've always liked the idea of Magnatune.com, although never bought anything..
En trois jours, Radiohead avait vendu 1,3 millions d’albums. Le prix moyen aurait été de 6 euros - un chiffre qui semble être tombé à 4 euros après que les premiers fans aient passé commande. Avec l’élargissement de l’audience a un plus grand public, la moyenne du prix d’achat s’est tassé : on estime entre 1/4 et 1/3 le nombre d’internautes qui auraient choisis de ne rien débourser. Wired estime néanmoins que le groupe aurait déjà pu récolter entre 4 et 8 millions de d’euros.

Même avec une moyenne basse de près de 3 euros par album vendu souligne Guillaume Champeau sur Ratiatum, c’est près de 4 millions d’euros que le groupe aurait gagné en quelques jours. A une dizaine de pourcent de rémunération par album, “dans les circuits classiques, Radiohead aurait du vendre 2,5 millions d’albums pour gagner l’équivalent”. Internet actu

Friday, October 19, 2007

what every PhD should know: dinner discussion with a google guy

Just barely hanging out there. Today was lots of serious fun and intellectual challenges at the RecSys 2007 Doctoral Consortium . Interestingly, all the participants came from lots of different backgrounds from computer science, information retrieval to me from education. I think the diversity of backgrounds and focuses of studies represent the growth of the Field of Recommenders, it's not only about the best algorithm anymore, but a plethora of questions around.

Anyway, being somewhat a newbie here (yeah, I do not know all the people or study areas here, very eye opening!), it makes me think that all the PhD students should be exposed to the question " If you were to have dinner next to a main researcher in Google/Yahoo/or any other big name, what would you want to talk about?".

Well, as it happens to be, I never thought of that before. Neither was I prepped for that by my study programme. Nevertheless, I just spent my dinner next to Krishna Bharat, you know, the guy who greated Google News, nothing less, nothing more. Probably tomorrow I'll have like ten things I want to ask from him with no chance to get his attention anymore.

Bottom line: it is not only about the 1 minute elevator pitch, but about the life and such in general.

Tuesday, October 16, 2007

New acquitance: Semiotic Dynamics

Pretty exiting, I came across this new area of Semiotic Dynamics, which is described as "a new field that studies how semiotic relations can originate, spread, and evolve over time in populations, by combining recent advances in linguistics and cognitive science with methodological and theoretical tools from complex systems and computer science." One topic of this study field is folksonomies, which draw my attention. The stuff can look like this.

Everyone nowadays repeat the same mantra of web 2.0, but somehow this project managed to say things sets it apart:

..users are no longer limited to consuming or creating online content, they also provide the semantic scaffolding holding together such content, thus taking on an active role in shaping the architecture of online information. The collaborative character underlying many Web 2.0 applications puts them in the spotlight of complex systems science,..

"Semantic scaffolding holding together .. content", that's a pretty awesome way to put it!

The paper "Vocabulary growth in collaborative tagging systems" investigates the temporal evolution of a tagging vocabulary size (of delicious) both on a
  • global level (the number of distinct tags in the entire system) and
  • local level (the growth of the number of distinct tags used in the context of a given resource or user).
It asks questions like how does the number of tags grow?; what is the rate of invention of new tags? is the asymptotic number of tags finite (uugh, a nice way to say it)? etc...

The paper finds out that the growth behaviours are remarkably regular throughout the entire history of the system with power-law behaviours with exponent smaller than one (non of that "fat head and long thing tail"!) and across very different resources being bookmarked.

Moreover, they find that there are some intrinsic characteristics of the system which do not depend strongly on the size of the dataset, like that the average number of tags is about 3.4 (local level). If I get it all right, they conclude on this that on the local scale (resource or user) "all curves tend to lie along a "universal" growth curve with an exponent close to 2/3".

The authors of this paper also highlight that the tools and concepts from complex system science may prove valuable for understanding the structure and dynamics of folksonomies.

Some interesting papers towards this direction: http://www.furl.net/members/vuorikari/semiotic_dynamics

Wednesday, October 10, 2007

Google goes micro-blogging

All the roads lead to ...Google. I guess we could re-phrase the old saying. Just received a notification from Jaiku that they are joining Google. In the other words, Google acquired them. Good for those guys, I hope. I wonder how many Finnish SMEs have become part of Google in the past?

I kinda enjoy Jaiku even if I don't micro-blog from my phone. I like it as an aggregator of feeds and to check what my pals are doing. Unfortunately the Facebook app. does not work that well, but hey, maybe Google will fix this one?