Monday, December 31, 2007

End of the year account: travels

I decided to write an end of the year summary for myself about 2007. This is a personal account, so if you come across this, don't bother boring yourself with it.

It will include important things like my travels, studies and worklife. I'll do it so that I can look back and say, what the hell of a year it was. I bet there will be years that are not that gentle with me, so it'll be fun to dig this year up and read how things were once upon a time.

Start with travels, most important ;)

2007 was a good travel year. Like any wanna-be globetrotter, I managed to put my foot on a new continent this year. I also got to discover new pretty rare places and visited the usual ones, so it's a good count for 2007.

To start off, 2007 was going to be my EPIC ski year. I had a season pass to Winter Park in Colorado and I planned to spent some two months in Wyoming. After all, this was my last year on scholarship, so I had to make a good use of that freedom ;)

It turned out not to be an epic ski season. Not for the shake of snow this time, but for the fact that Matt broke his leg, pretty seriously, during our second weekend of skiing in Utah. Well then, I got some seven days in, though, and busted my dear snowboard after many faithful years of use. So, the best part of Q1 was spent in Wyoming and I also got to visit Utah, especially the LDS Memorial Hospital in Salt Lake City (Matt had his surgery there).

Beginning of June was the start of the great big sailing adventure (pics and blog). Oh my, that's gotta rank as one of the best trips ever. Matt and I took a flight to Rarotonga, Cook Islands, where we met with Dan and Danielle, who had sailed there from San Francisco via Mexico, French Polynesia and such.

We boarded on Confetti, a beautiful 44-footer, a hand-grafted wooden boat (Farr), to sail for about 1200 sea miles. We spent some 3 weeks on and off board exploring Rarotonga, Beverage reef (wow, even Wikipedia does not have an entry for it!!), Nuie and American Samoa, all small tiny places in the vast South Pacific.

We put in a total of seven-eight days of off-shore sailing (in 3 legs), which is by far the longest time I've ever been out on the open. And I really mean out on the open, hell, those "trade wind" areas, there is no one out there! During the first leg from Raro to Beveragereef and Niue, which took about 10 days, we saw no one, no boats, no planes, nothing. Just us in Confetti with our Ham-radio connection to weather forecasts and some necessary emails. That was some experience that I'll keep cherishing for a long time.

Another curiosity of this trip was the type of "civilization" that we encountered when we stayed on those islands. First, I gotta admit that they are something to dream about; beautiful white beaches, palm trees, thick green vegetation and so on. Just what you would expect. Sadly, though, the life that people lead on these "paradise" islands is far from my notion of paradise. It seems that "civilization" from missionaries, and early and nowadays merchandisers and explorers has dramatically altered the way of life, and globalization with the help of satellite TV and subsidized food from New Zealand has taken over any tiny bit that was left of the original way of living.

It's hard to put words on what we saw without sounding like a "conservationist". I don't want to say that people should still live like they did when Captain Cook and other discoverers first met them. What I want to say is that it hurts seeing such level of obesity, unemployment and apathy against your own nature and environment on some of the most beautiful spots on earth that I've ever seen.

Apart from the South Pacific, this year's new discoveries include Cyprus. I finally got to see Barcelona, it's as great as people say it is!

I also got two trips on my miles, which is good.

Wednesday, December 12, 2007

I'm not a happy bunny with Leopard, grrr

Oh my, just gotta say this out loud! I'm not happy to have the new Leopard running on my Mac, really not :(

It's not that I'm very picky about things not working smoothly, after all, before switching back to Mac I ran only Linux on my desktop for about a year and a half. It wasn't always easy for a non techky user like me (I still get pimples if I have to use command line...), but I wanted to give it a try. After all, at that time I talked a lot about using FLOSS in education. So to know what to talk about, I needed to do it myself.

About Leopard. Last drop today was that I can not get one of my favourite widgets working, namely Screenshot plus. They claim they became Leopard compatible, but bullocks, it still does not work. And that is not the only thing!

My Mail crashes about five times a day, luckily still saving the email drafts. And that is a Mac native application, what the hec?? I also use OpenOffice, probably the only individual in the world to do it on Mac. But hey, after all I bought a Mac 'cause I did not want to use MS programmes.

So, OO became REALLY slow and it does not support all the keyboard shortcuts or virtual keyboard language selections (I use Finnish and US keyboard and switch a lot). This makes my work really sucky.

It's not actually only OO that became slow, it the whole computer. I constantly use all my virtual memory, 1.5Gb, which does not seem to be sufficient for this hungry animal. All the searches within the computer are really slow too.

Yes, also the Java support is inexistent, and I have no idea how to deal with that.

Well, I feel much better now. Thanks god there are blogs to vent out... whew..

Monday, November 19, 2007

eTwinning/bookmarks and social networks

Excellent write up here about SNA on social networking sites. This makes me think of eTwinning, or my social bookmarks, and the SNA analysis there to better support users.

Kumar, Novak and Tomkins (200&) saw that network activity is of three types:
  • “Singletons,” who have no connections and are least central
  • The “giant component,” which is the largest group of nodes tightly connected to the central nodes and to each other
  • The “middle region,” which represents isolated groups which interact amongst themselves but not with the rest of the network, forming isolated stars. These groups grow one user at a time. Over time they merge with the giant component.




















The node analysis of these networks showed that more than half of a social network is outside the giant component where the greatest centrality lies. They used the “control” definition of centrality to determine this. The research also highlighted a prevalence of “stars” in the middle region which are mini social networks, typically driven by one dynamic member who serves as the point of centrality with others serving as satellite nodes – connected to the dynamic member but not to each other. In Kumar, Novak and Tomkins’ analysis the middle region represented one-third of users on Flickr and about ten percent of users on Yahoo! 360.

Also keep in mind that the most growth happens in the middle region where dynamic members influence others to join their network. These sub-networks can gradually join the giant component over time. Once they do, the importance of the dynamic member diminishes. Even if that dynamic member were to leave the network, the others would stay in the network.
So, what is needed is to support the "stars" in their growth so that they become independent of that one dynamic member and are able to continue even without that person.

Wednesday, November 14, 2007

Such a cool way to send a message, thanks B.Dylan!

Tagging in different e-learning environments

In the last days we've had a few discussions about tagging in e-learning environments. My environment, where the tagging takes place, is a portal for learning resources.

Today I came across this nice graph that displays "power law of participation". Now, haven't looked at the scientific background of it yet, so no comments on that. Anyway, it kinda rang the bell with what I'm doing when looking into levels of user engagement on the portal.

According to this graph, adding things to favourites (e.g. bookmarking) and tagging them represents a pretty low threshold to participate in the activities of that given community.




















I'm also looking at leMill environment, which on the other hand, demands a pretty hight level of user engagement, as it is about collaborative authoring of digital learning resources. About a year down with users, there is somewhat little collaborative authoring that actually takes place, Hans told me yesterday.

Maybe tagging in some way could help the participant to take the first steps? well, they can already tag and favourite things in LeMill, so maybe the issue is rather to see if similar levels of engagement appear in that community.

So, along with that, I am interested in looking at the tags in LeMill from the same point of view that I'm doing for tags in our learning resources portal. The difference is that our case is clearly what is called broad folksonomies, whereas leMill should be a rather classical narrow folksonomy. Or is it? Maybe once we start looking at those tags as a triple {user, resource, (tags)} with a timestamp on them, it appears that participants first start by bookmarking and tagging resources from other users, before the user takes a step to create her own resources and finally collaboratively work on other's resources.

Update:

So, some data to back-up was found:





















Social Technographics®
Mapping Participation In Activities Forms The Foundation Of A Social Strategy
by Charlene Li
http://www.forrester.com/Research/Document/Excerpt/0,7211,42057,00.html
with Josh Bernoff, Remy Fiorentino, Sarah Glass

This is a document excerpt EXECUTIVE SUMMARY
Many companies approach Social Computing as a list of technologies to be deployed as needed — a blog here, a podcast there — to achieve a marketing goal. But a more coherent approach is to start with your target audience and determine what kind of relationship you want to build with them, based on what they are ready for. Forrester categorizes Social Computing behaviors into a ladder with six levels of participation; we use the term Social Technographics® to describe a population according to its participation in these levels. Brands, Web sites, and any other companies pursuing social technologies should analyze their customers' Social Technographics first and then create a social strategy based on this profile.

Wednesday, November 07, 2007

notes on "Context, (e)Learning, and Knowledge Discovery for Web User Modeling: Common Research Themes and Challenges"

"Context, (e)Learning, and Knowledge Discovery for Web User Modeling: Common Research Themes and Challenges" by B.Berendt

This paper is about context and how to define it or how it is defined differently. The following is related to "Context in Web usage mining and eLearning"

2.1 Context as data and as metadata

"In order to evaluate whether intended and actual usage coincide or not, and in order to obtain a more fine-grained picture of actual usage, it is of course interesting to measure aspects of actual usage. "

- This is also one thing that we are interested to find out in MELT, and partly also in my PhD. As we have very little access to "actual use" we try to infer this type of information from usage logs. E.g. We have a teacher who has said in his profile that he teachers students from 12 to 13 year olds. If he bookmarks LOs that have intended audience of 14-18, we can maybe infer that this LO can also be used for younger students. Especially, if we start seeing this taking place a lot, we might want to update the LOM on intended audience: instead of 14-18 we could say 12-18.

- My interest is also to see if tags can give us any hints of this.


2.2 Context and model parts

"context representations can form and/or enrich (a) user models, (b) material/environment models, or (c) interaction models."

- EUN uses a) in one search to rank resources, but we are still only implementing it and we don't know how users react to it. That is related to my own PhD, as are how different search methods are used. In general, we do way too little with user modeling (I guess bigger issues are still more imminent)


2.3 Context: parameters of the (inter)action

- For my PhD I'm looking into user logs to create "levels of user interaction", e.g. what does it mean if a user views a page vs. makes a bookmark on it. We want to use this as an input for a recommendation system, for example.

- I'm also interested in the type of search that the user has chosen and its relation to the search task that the user has at hand.

- Tags were mentioned in this context, that is also a huge part of what I am looking. There are different questions around them, one most interesting related to search is how they can be used for discovering resources.

Need to look into these papers:

- B. Berendt, G. Stumme, and A. Hotho. Usage mining for and on the semantic web. In H. Kargupta, A. Joshi, K. Sivakumar, and Y. Yesha, editors, Data Mining: Next Generation Challenges and Future Directions, pages 461–480. AAAI/MIT Press, 2004.

- Claus-Peter Klas, Hanne Albrechtsen, Norbert Fuhr, Preben Hansen, Sarantos Kapidakis, L aszl o Kov acs, Sascha Kriewel, Andr as Micsik, Christos Papatheodorou, Giannis Tsakonas, and Elin Jacob. A logging scheme for comparative digital library evaluation. In Julio Gonzalo, Costantino Thanos, M. Felisa Verdejo, and Rafael C. Carrasco, editors, ECDL, volume 4172 of Lecture Notes in Computer Science, pages 267–278. Springer, 2006.

- Totally agree with his observation, not the method: Tanimoto [53] emphasizes that may be difficult to conclude, from a mere clicking event, that there was indeed attention paid to (specific) content of the requested page.


2.4 Context: background knowledge

Tags, tags, tags. multiple views.


2.5 Context: Activity structure
" This metadatum can provide important information about a visitor’s intention or expectation (e.g., whether they followed a prescribed link from a course page, or whether they found a material by actively searching with a very detailed search phrase)."

For me this is important, I guess using terms from this paper, I'm interested in user's intentions and expectations and finding out the ways the users choose to access or discover resources in our portal. I'm also interested in seeing whether one method is more useful to a given task, e.g. if people like browsing to find inspirational material and some other method (social information retrieval vs. information retrieval) for another task. If we know what kind of method is useful for a given task, I think we can help our users a lot.

2.6 example

An example is given using the three aspects of context; activity structure, parameters of the (inter)action and background knowledge. The type of analysis allows answering questions like: which search options are popular and are there differences between users? Which content areas were frequented, and how did people navigate between then; did they go back to the search options, or did they use the inter-content links? Did certain content areas become hugs for navigation and thus served to organise the domain and the presentation of the domain? On the other hand, questions like; were there differences between users with high verbal and users with high visuo-spatial competencies; did certain textual or pictoral material become hub?

These are also questions that I am looking at in lre portal and am getting a good idea of them. However, I have not been able to link them with the task at hand yet, which is something that I'm interested in.

Schooling for Tomorrow scenarios and Science-Fiction

Last autumn I heard a few e-learning keynotes with a heavy science-fiction emphasis. That was wild, I totally loved it. I can not agree more on the idea that working with future scenarios, or science-fiction for that matter, is actually helping the future to come along. It is about helping to shape the future after having peeked in to the future with a positive or a negative outlook. And, I'd like to say that looking or creating scenarios in many perspectives is useful too, as it helps you to see whether this is where you want to end-up or not.

There's been some work going on since the beginning of the millennium regarding scenarios for the future of schooling, we also in EUN worked on that. The OECD report Schooling for Tomorrow, is out too. It's way less exiting than talking about 2048 in a science-fiction scenario, I'm afraid, but still aiming at the same goal - see how technology and network enhanced learning could be used in the days to come.

Of course I picked upon the scenario called "Learning in Networks replacing schools"













This scenario imagines the disappearance of schools per se, replaced by learning networks operating within a highly developed “network society”.

Networks based on diverse cultural, religious and community interests lead to a multitude of diverse formal, non-formal and informal learning settings, with intensive use of ICTs.

How about that for science-fiction?

Well, if it is up to me, I would like to see learning networks in schools even if schools per se are not facing the extinction. You know, before we need to go to a dinosaur museum to see a replica of a teacher.

I don't mind the idea of "networked society", but somehow, when wearing those gray classes, I'm thinking of the efficiency of terrorist cells workings, and how religious and local interest groups could manipulate their own learning interests on me, while at the same time monitoring with whom do I want to learn on the international scale. Yak!

The Schooling for Tomorrow scenarios by OECD are a real tool set for policy-makers, and why not others, to work on. I like practical things like these are. Lately, I've been toying with the idea of making people, who work with education, technology and networks, to write science-fiction short stories of learning in the future. I also think that this would be a very helpful exercise for other PhD students to open up their thinking and not be stuck with what we got now.

How to get started with your own science-fiction story on Schooling for Tomorrow

It's important to understand that it's the journey that is important, not the destination. To say it in other words, I think the whole thinking process to come up with a plot for your story is what counts! Don't worry about picking up the right publisher now ;)

This is what I found out about writing science-fiction, how to get started
  • idea: the premise, the basic thought around which the story will turn
  • a setting: worl for your story to take place in; it can be familiar, or wholly new
  • Characters: two or three, at least, to people your story
  • Aliens: (optional) strange and mysterious beings for your characters to encounter
  • Problem: something your characters want, need, must escape from, etc.
Hmm, I really like to think of Aliens and education...

Next, when you have those sorted out, think about the five dimensional framework from Schooling for Tomorrow. What are:
  • “attitudes, expectations, and political support”,
  • “goals and functions” of education systems
  • “organisations and structures”,
  • the “geo-political aspects”
  • “the teaching force”.
Then, just let it flow!

My own attempt on this is in a wiki. I've already had quite a few engaging and hilarious dinner discussions with my friends about what should the plot be. We never got very far, but it's always been a lot of fun! I only wish we could somehow get the discussion transcripts on the wiki....

Tuesday, November 06, 2007

Notes on "Collaborative tagging and Semiotic Dynamics"

By Gattuto, C., L.Vittorio and L.Pietronero (2006).

Firstly, I must say that I was glad to read this paper. Lately, I've been seeing many papers talking about the properties of folksonomies, like co-occurrence, etc., which have intrigued me quite a lot. This paper explains the process pretty well and underlines an important point - they factor out the users and only deal with streams of tagging events and their statistical properties!

I must admit that this makes the whole area of Semiotic Dynamics less attractive to me. I think it is important to study tags and their properties, but not in isolation from the user. I see (barely) the point to explain tagging activity and the growth of tags in separation from the users. But fair enough.

Problem statement: Uncovering the mechanisms governing the emergence of shared categorisatioins or vocabularies in absence of global coordination is a key problem with significant scientific and technological potential. Collaborative tagging provides a precious opportunity to both analyze the emergence of shared conventions and inspire the design of large agent systems.


Semiotic Dynamics study how populations of humans or agents can establish and share semiotic systems, typically driven by their use in communication. The author argue that the emergence of a folksonomy exhibits dynamical aspects also observed in human languages, such as the crystallisation of naming conventions, competition between terms, takeovers by neologisms, and more.

  • Users interact with a collaborative tagging system by using tags or adding new resources to system
  • Basic unit of information in collaborative tagging systems is a (user, resources, {tags}) triple, which they refer as post in this paper. Tagging event is a tri-partite graph (with partitions corresponding to users, resources and tags, respectively) and can be used as a navigation aid in browsing tagged information
    • Comment: I like the tri-partite graph as navigation aid, yes!, but as the authors mention just above, they don't think of other users and those networks as navigational aid. In contrary, they omit the users just to study the properties, which strikes bizzarre to me.
The authors cite the "rich get richer" model (Yule-Simon's stochastic model) and propose to enhance it with a "fat-tailed memory kernel". This original model is related to the construction of text from scratch:
At each discrete time step one word is appended to the text: with probability p the appended work is a new workd, never occurred before, while with probability 1-p one work is copied from the existing text, choosing it with a proability proportional to its current frequency of occurrence. This simple process ields frequency-rank distribution that display a power öaw tail with exponent alpha = 1-p, lower than the exponent we observe in actual data. This happends because the Yule-Simon process has no notion of "aging", i.e., all positions within the text are regarded as identical ..
This all leads to a model of users' behaviour: the process by which users of a collaborative tagging system associate tags to resources can be regarded as the construction of a "text", build one step at a time by adding "words" (tags) to a text initially comprised of n 0 words. There is also that same Yule-Simon model with long-term memory (about inventing new tags or using existing ones), but recent tags are used more often than old ones.

Also, "in our model,.., the average user is exposed to a few roughly equivalent top-ranked tags and is translated to mathematically into a low -rank cutoff of the power law, i..e., the observed low-rank flattening".

Conclusion: It seems that users of collaborative tagging system share universal behaviour which, despite the intricacies of personal categorisation, tagging procedures and user interactions, appear to follow simple activity pattern.

There is also something about the co-occurrence between high-rank and low-rank tags: it says: "This suggest that high-frequency tags partition - or "categorize" - the resources marked by tags of lower frequency. "
Comment: This all sounds interesting and important, but will need to look into that later.


Monday, November 05, 2007

My PhD research

Notes on "Aspects on Broad Folksonomies"

Aspects on Broad Folksonomies by M.Lux and M. Granizer (2007)

This paper continues the trend in studying and analysing the underlying statistical properties of broad folksonomies that aims to identify laws and characteristics which allow inferring those properties. A few notes on what I found interesting related to the emerging notion of quality of tags, something that I've also spared a few thoughts on.

First, though, on some other issues. The paper talks about the emergence of power law distribution in folksonomies. They describe which approach they took to fit the sample to a power law, which was something that I've sometimes contemplated on the how-part of things. The paper aims at analysing whether one can find similar term distribution in folksonomies as in classical term retrieval (e.g. Zipf. note: Zipf's law with an exponent between 1 and 2). The dataset is that of delicious (uh, with about 800 000 bookmarks and about 27 000 users- I got a way to go with my MELT bookmarks).

Tag co-occurrence
They are able to show that "for around 80% of the tags of a folksonomy the co-occurring tags follow a power law distribution, which approves Cattuto's assumption. We found that for about 90% of the estimated power law exponent B xxx [-1.5, -0.5], which shows that for most tags co-occurrence follows a model with similar parameters. "
Resource and user based tagging characteristics
Secondly, they looked into frequently used tags (more than 30 users).
  • For resources statistics they (frequency of users tagging the resource with a tag) found that around 18,4% of resources followed a power law distribution.
    • assigned by lot of users to few resources (head) and to a lot of different resources by a few users (tail)
  • For user statistics (frequency of resources tagged with a tag), around 13% are following a power law.
    • few users tag a lot, whereas lot of users tag a few
  • i.e. the characteristics of the user statics are similar to the characteristics of the resource statics.
  • They argue that those tags, which follow a power law w.r.t users and resources are high quality tags (i.e. tags describing resources with high accuracy [no misspellings and meaningful tags] ) for most of the users involved in the investigated social bookmarking system.
  • A small fraction of tags have overlapping user groups, which points towards sub communities (user groups sharing the same link selection and tagging behavoiur) in the tail of the power law distribution.
    • this was found through splitting resources in 3 (high, mid and low rank resources)
They also looked at the big chunk of tags that were not following the power law.
  • Unique assignments. More than half (57%) of less frequently tags are used only once. They think that they can be seen as "shortcuts" for a user to a resource or a misspellings. They argue that these tags are useless from retrieval point of view (hmm..).
  • Personal vocabulary. especially in less frequently used tags (19%) of tags were only used by one user but assigned to many resources. They are useful for personal retrieval but useless for the rest of the community.
  • Unpopular vocabularies. between 1/5 and 2/5 of tags are assigned to different resources by different users only once. Unpopular vocs used by a small fraction of users.
  • they conclude that from retrieval point of view (e.g. inverted indices, TF*IDF) a large fraction of tags are good for single or sub-communities, and only the power law distributed tags are good for that.
    • They don't say anything about how to include the large fraction of tag not distributed by power law into IR methods.
Retrieval Aspects
Q: Do tags add information to further to description and title for retrieval purposes? This is a lot along the lines that I am also interested in, although I will look more into the networks of users. They say that for retrieval tags can be seen as an additional resource. Moreover, about 50% of available description contain information similar to the information described by tags, whereas the remaining 50% can be seen as orthogonal information.

Comment. This all is treating tags only as additional keywords that can be useful for conventional retrieval purposes. I think the connection tag-resource-user is more interesting. Just the fact that even if the tag is misspelled or hooks to a small user community is less important to me, because I know that the fact that this resource was tagged shows that the user has an interest to this resources, thus it is a vote. This aspect has an immense potential for retrieval (recommender point of view), but is seldom regarded in papers with very conventional retrieval approach.

Open social and education

I wonder who is going to come up with the first OpenSocial app or widget for educational use? We certainly are talking about it, for example for our eTwinning platform. It could be cool to be able to use information about teachers collaborative networks to allow, say, better retrieval of learning resources relevant for the project, purpose or task that teachers are undertaking; link with some other sources that teachers are working on through cool widgets, etc.

I never thought that Facebook, which has lately become really popular among my friends (not early adapters), would be the seul app that would "take it all". I was glad to read this:
"The market has already decided that there's going to be a long tail of social networks, and that people are going to belong to more than one. As soon as you belong to more than one, this kind of interoperability is critical," Dash says. "Open standards win every time." wired

Hurray for open standards!

Tuesday, October 30, 2007

Multilingual tags and the language of LO

I've looked into tagging in different languages before. An interesting thing came out of our little pilot: teachers, non of whom mother tongue was English, still had about 20-30% of tags in English. We had two different thoughts on this,
  • either tag is in English because teacher wanted to share these tags with other teachers, or
  • tag was in English because it is related to the language of the learning resources that was bookmarked
I was now interested in the second possibility, and took a look at a sample of 136 bookmarks with tags in multiple languages related to them.
  • The LOs were in English, Hungarian, Polish and Estonian.
  • The users (43) were Hungarian, Polish, Estonian and Lithuanian
Fair enough, all the English tags were related to the English resources! In close to 30% bookmarks (39 out of 136) this was the case (which also means that 30% of LOs were in English).

Moreover, it seems that for about 1/3 of the times the language of the LO was the same as that of the tag, whereas 2/3 of the cases it varies according to the language of the user. In about 3% of bookmarks one was able to observe multi-lingual tags.

Thursday, October 25, 2007

Radioheads making €, good for them

Cutting the record lable out of the equation seems to be a good deal. I'm glad to see the bold decision to go directly to fans has not only shown a great example, but also proved profitable! I've always liked the idea of Magnatune.com, although never bought anything..
En trois jours, Radiohead avait vendu 1,3 millions d’albums. Le prix moyen aurait été de 6 euros - un chiffre qui semble être tombé à 4 euros après que les premiers fans aient passé commande. Avec l’élargissement de l’audience a un plus grand public, la moyenne du prix d’achat s’est tassé : on estime entre 1/4 et 1/3 le nombre d’internautes qui auraient choisis de ne rien débourser. Wired estime néanmoins que le groupe aurait déjà pu récolter entre 4 et 8 millions de d’euros.

Même avec une moyenne basse de près de 3 euros par album vendu souligne Guillaume Champeau sur Ratiatum, c’est près de 4 millions d’euros que le groupe aurait gagné en quelques jours. A une dizaine de pourcent de rémunération par album, “dans les circuits classiques, Radiohead aurait du vendre 2,5 millions d’albums pour gagner l’équivalent”. Internet actu

Friday, October 19, 2007

what every PhD should know: dinner discussion with a google guy

Just barely hanging out there. Today was lots of serious fun and intellectual challenges at the RecSys 2007 Doctoral Consortium . Interestingly, all the participants came from lots of different backgrounds from computer science, information retrieval to me from education. I think the diversity of backgrounds and focuses of studies represent the growth of the Field of Recommenders, it's not only about the best algorithm anymore, but a plethora of questions around.

Anyway, being somewhat a newbie here (yeah, I do not know all the people or study areas here, very eye opening!), it makes me think that all the PhD students should be exposed to the question " If you were to have dinner next to a main researcher in Google/Yahoo/or any other big name, what would you want to talk about?".

Well, as it happens to be, I never thought of that before. Neither was I prepped for that by my study programme. Nevertheless, I just spent my dinner next to Krishna Bharat, you know, the guy who greated Google News, nothing less, nothing more. Probably tomorrow I'll have like ten things I want to ask from him with no chance to get his attention anymore.

Bottom line: it is not only about the 1 minute elevator pitch, but about the life and such in general.

Tuesday, October 16, 2007

New acquitance: Semiotic Dynamics

Pretty exiting, I came across this new area of Semiotic Dynamics, which is described as "a new field that studies how semiotic relations can originate, spread, and evolve over time in populations, by combining recent advances in linguistics and cognitive science with methodological and theoretical tools from complex systems and computer science." One topic of this study field is folksonomies, which draw my attention. The stuff can look like this.

Everyone nowadays repeat the same mantra of web 2.0, but somehow this project managed to say things sets it apart:

..users are no longer limited to consuming or creating online content, they also provide the semantic scaffolding holding together such content, thus taking on an active role in shaping the architecture of online information. The collaborative character underlying many Web 2.0 applications puts them in the spotlight of complex systems science,..

"Semantic scaffolding holding together .. content", that's a pretty awesome way to put it!

The paper "Vocabulary growth in collaborative tagging systems" investigates the temporal evolution of a tagging vocabulary size (of delicious) both on a
  • global level (the number of distinct tags in the entire system) and
  • local level (the growth of the number of distinct tags used in the context of a given resource or user).
It asks questions like how does the number of tags grow?; what is the rate of invention of new tags? is the asymptotic number of tags finite (uugh, a nice way to say it)? etc...

The paper finds out that the growth behaviours are remarkably regular throughout the entire history of the system with power-law behaviours with exponent smaller than one (non of that "fat head and long thing tail"!) and across very different resources being bookmarked.

Moreover, they find that there are some intrinsic characteristics of the system which do not depend strongly on the size of the dataset, like that the average number of tags is about 3.4 (local level). If I get it all right, they conclude on this that on the local scale (resource or user) "all curves tend to lie along a "universal" growth curve with an exponent close to 2/3".

The authors of this paper also highlight that the tools and concepts from complex system science may prove valuable for understanding the structure and dynamics of folksonomies.

Some interesting papers towards this direction: http://www.furl.net/members/vuorikari/semiotic_dynamics

Wednesday, October 10, 2007

Google goes micro-blogging

All the roads lead to ...Google. I guess we could re-phrase the old saying. Just received a notification from Jaiku that they are joining Google. In the other words, Google acquired them. Good for those guys, I hope. I wonder how many Finnish SMEs have become part of Google in the past?

I kinda enjoy Jaiku even if I don't micro-blog from my phone. I like it as an aggregator of feeds and to check what my pals are doing. Unfortunately the Facebook app. does not work that well, but hey, maybe Google will fix this one?

Saturday, September 29, 2007

Notes on Smart Indicators on Learning Interactions

Smart Indicators on Learning Interactions by Clahn et al. (2007) discusses how indicators can be used to help learners, or groups of learners, to organise, orientate and navigate through learning environments by providing contextual information that is relevant for performing learning tasks. Indicators are part of the interaction between a learner and a system (social or technical).

Indicator system is defined as a system that informs a user on a status, on past activities or on events that have occurred in a context; and helps the user to orientate, orgaise or navigate in that context without recommending specific actions.

So, it is not:
  • a feedback system (analyse user interactions to inform learners on thier performance on a task and to guide the learners though it) or
  • a recommender system (analyses interactions in order to recommend suitable follow-up activities),
  • instead it provides information about past actions or the current state of the learning process.
  • Moreover, smart indicator systems adapt their approach of information aggregation and indication according to a learner's situation and context.

The paper draws heavily on the notion of social navigation, interaction history and footprints, and offers a good review of this literature (ToRead).

The paper offers an architecture of smart indicators, where different layers are defined to support user modeling (first two) and helping the system to adapt to better decision making process (last two). Four layers:
  • sensor layer
  • semantic layer
  • control layer, where a strategy defines the conditions according to learner's context
  • indicator layer, presents aggregated information to the learner.

This approach of smart indicators adapts the strategies on the control layer (as opposed to semantic layer) to meet the changing needs of a learner.

SENSOR AND SEMANTIC LAYER

The paper further presents the information aggregates of sensor and semantic layers. The idea is to classify and organise the user's engagement (interaction foot prints) with the system, e.g. contributions, tagging activities. In the sensor layer, there is a division between "learner interaction" and "contextual sensors", e.g. location tracker, tagging activities (in my case this is considered direct) and contributions of peer-learners.

I am doing the same with my research data, and I call it the "user engagement" following the Yahoo!'s idea on STAR-metadata (kind of attentional and explicit metadata about users actions).

I tried to apply the classes of Chlan's prototype to my research data (learning repositories) that I collect using our CAM framework. Our focus being somewhat different, it did not really work out that well. The attempt below, though:

Direct: accessing resources through browsing, tag cloud, search result list, other user's favourites (implicit interest)
  • user views metadata
  • user views tags
  • user views resource ("entry selection sensor")
  • (timestamp on everything)

Direct: higher level interaction with a resource (explicit interest)
  • user adds a resource to favourites and tags it ("entry contribution sensor", "tag selection sensor", "tagging sensor" or"tag tracing sensor", hard to say in my case)
  • user rates the resource ("entry contribution sensor")
  • user comments on the resource ("entry contribution sensor")
  • shares resource with network ("entry contribution sensor")
  • (timestamp on everything)

Contextual sensors could be (here I'm blending them with user information):
  • context of a project within which the user access resources
  • the information about the country and school from where the user is from

SEMANTIC LAYER

The semantic layer users the information from Sensor layer and transforms it into meaningful information by using an "activity aggregator". This calculates the activity for a given period of time for an individual learner or the whole community according to different ratings that each activity has (beginners have different way of counting activity from power-users).

CONTROL LAYER

In this prototype the control layer defines how the indicators adapt to the learner behaviour. There are two elemental strategies:
  • motivate learners to participate to the community activity
  • raise awareness on the personal interest profile and stimulate reflection on the learning process
Moreover, a third level control strategy uses the activity aggregator as well as the interest aggregator.

INDICATOR LAYER

This layer embeds the indicators into the user interface of the community system. The prototype is being tested by a group of PhD students now.

Glahn, Christian, Specht, Marcus, Koper, Rob (2007) Smart Indicators on Learning Interactions
http://hdl.handle.net/1820/941

Thursday, September 27, 2007

Some thoughts after SIRTEL07

Last week the SIRTEL workshop took place. The papers are found here and the slides, well, most of them, at the EC-TEL07 conference wiki. I have pretty good feeling about the workshop, it was one day long, we had about 20 people participating, some of whom chose to stay with us for the whole time, and some who were hopping between workshops. For me that is totally fine, we all are responsible for our own learning! Especially in conferences where many parallel sessions are running, I would encourage people to try to get best out of them.

For those who could not make it at all, you can soon find recording on the SIRTEL site.

The workshop had four main sessions:
  • We started with a keynote address from people who work with music recommenders. MyStrands people talked about applying social recommender systems to technology enhanced learning. It was an interesting talk that challenged all of us to think what are recommenders for learning purposes in the first place (goal) and what kind of data do we want to use to do that.

    As any good keynote, this one gave more ideas to think than answers. It nicely set the base for the further discussions during the workshop that focused on the need to define the field of Social Information Retrieval for Technology Enhanced Learning, and to establish a baseline so that we know what are we really set to do.

  • The second session was about Tagging and Visualisation. We had my presentation about the user behaviour on tagging in multiple languages; then there was a presentation from COSL that talked about "Activities of Daily Living on the Web", Brandon also showed a few demos of the widgets that can be used to rate or recommend related content. That was followed by a talk on reward structures to encourage teachers to share open educational material. Finally, we listened about Visualisation of social bookmarks, a work that leads into visualising bookmarks in an educational repository.

  • The 3rd session was on Recommender Systems. Here we first heard about some R&D work that OU NL is carrying out using the idea of learning paths to better support learning activities of students. Then, there was a study about using affiliation networks as a mechanism for collaborative filtering (understood largely). This was followed by a study on simulating recommendations based on multi-attribute ratings on learning resources by teachers. Finally, we had a system demo of Daffodil that supports collaborative information seeking.

  • The final session was what we called "Enablers and Challenges". It was a discussion session, and as we advanced, it was clear that people had a lot to say. It might even have been better to allow more time for this, but hey, you live and you learn.
I try to sum-up, but basically it illustrates the main topics that we talked about. If you look at the left side, there are the fundamental questions:
  • How to define and chart out the area of Social Information Retrieval (SIR) for learning?
  • Is this application domain different from other SIR, on micro and macro level?
  • What do we recommend?
  • In what context?
  • and based on what?
On the right hand, there are the issues related to implementation and evaluation of it. These are:
  • What are the best SIR methods for TEL?
  • And what is the data that we should use? The "data issue" was something that was heavily emphasised by the MyStrands folks, who obviously speak of experience.
  • The questions rouse also: when do we start implementing these for real or are we just over-engineering and never ready to launch?
  • Evaluation and empirical data for real evidences was on the focus a lot.











More will follow. This is quick and dirty now, hopefully I will get more input from people participating in order to get more depth on our summary.

Sunday, September 16, 2007

SIRTEL'07: la raison d'etre

"We use people to find content. We use content to find people."*

On Sept 18 our SIRTEL workshop takes place. It's gonna be "Serious Fun"! Let me just outline why:

SIRTEL'07: Raison d'etre

Recommender systems, as well as social navigation, have been around since the popularisation of WWW, that's some 15-20 years now. The idea is to help people choose the right stuff from a potentially overwhelming set of choices. To facilitate that users could be helped with information from other users, the choices made before (by themselves or similar users), the ratings or reviews other people had done, etc. (Rescnik et al., 1997)

The field of learning technologies has seen recommenders of some sort being discussed and prototyped since the late nineteens. In the review of the field in Manouselis et al (2008) we identified about 10 recommenders, and even more conceptual papers of them, but very little has matierialised so far.

Since the last few years recommenders have made a second arrival into the discussion topics of technology, or network, enhanced learning. Undoubtedly, this has been influenced by the arrival "Web 2.0" with all its ideas:

- Collaborative tagging, for example, has changed lots of ideas of how metadata should be produced and how static a metadata record should be: it's not anymore one metadata record produced by a librarian, but lots of annotational and attentional metadata by lots of users.

- Other annotations by users that express their subjective judgements have seen a huge growth too, we don't only talk about ratings or reviews in their traditional sense, but also tumbs-up or down, giving pokes to people or objects, etc.

- Social bookmarking, which allows users to create easy references to their own collections of digital resources (photos, books, links, music,..), has given a new dimension to the concept of social navigations. The link between resource-user(-tag) allows users to navigate other people's collections and thus find novel resources. Also, the same resource-user-tag link gives researchers an itch to use this information to group similar users for recommendation purposes, as well as to study the emerging networks.

- Expressing social ties between people has also brought new possibilities along. We are not only seeing networks of friends, but there are new possibilities where people can express different networks, ones for professional use, others for personal, recreational, etc purposes. Also, portability of these networks has become an issue discussed for better designs (social-network-portability group, PeopleWeb ,..).

- Something else is also happening behind the scenes. Clicksteam and user behaviour on the Web is not anymore a property of the commercial portal on which users are, but users are starting to take seriously how their "attention" is being used, who owns it, etc. Attentional metadata is a huge source of information that educationalists are also starting to take more seriously and thinking how it could be used for better serving learners and teachers (Contextual Attention Matadata, Attention Profiling Mark-up Language, Attention Trust,..). Attentional metadata can also become crucial when it comes to better understanding the intentions of a user, why are they, for example, looking for some information and for what task at hand!

- Finally, content for educational use, or rather its production, is also seeing a change. Users generate more and more of the content on the Web in general, a trend which is also seen in the e-learning. Of course, traditionally teachers have always produced lots of their own material, but now its re-use also has been facilitated (e.g. repositories/referatories). Also, the collaboration aspect is facilitated by the Web, it has become easier for people to work together on things (e.g. wikis, collaborative platforms,..). Additionally, learners produce plenty of material which also should be seen and used as educational content.

To sum-up: two main topics evolve around social context and social content. Social context is how we express the who, where and with whom, and social content are the objects or digital artefacts that are in the center of the communication, exchange and networks.

All the above has hopefully also changed how we will see the future of social information retrieval for technology enhanced learning. This workshop will all be about that! Serious Fun!

-------

N. Manouselis, R. Vuorikari, F. Van Assche, “Collaborative Filtering of Learning Objects for Online Communities: An Experimental Investigation”, accepted for publication in Computers in Human Behavior, Special Issue on ‘Advances of Knowledge Management and Semantic Web for Social Networks’, 2008.

P.Morville, 2004

Resnick P. & Varian H.R., “Recommender Systems”, Communications of the ACM, 40(3),1997

Wednesday, September 12, 2007

How do teachers network?

Just wondering...

I had an awsome chance to join this teachers' community.

View my profile on eTwinning Reunion