Showing posts with label social navigation. Show all posts
Showing posts with label social navigation. Show all posts

Friday, July 17, 2009

Personalisation vs. Social

I've been thinking of this personalisation-thing a lot lately. I quite cannot get my head around it, so this blog post is just to mull over the ideas.

By personalisation it is meant that, for example on a learning resource portal, the offer is tailor made for one of the users of the system. There are three commonly known ways of doing this:
  1. Based on self-proclaimed profile (e.g. you say you teach math, so your services are personlised towards math)
  2. Based on collaborative filtering e.g. ratings (like-minded users, ppl who agree in the past tend to agree in the future) or content-based filtering
  3. Based on behaviour (e.g. other teachers who used this resource, also used xx)
In the cases 1 and 2, the idea is that there is a profile for you, whereas no: 3 can be used for any user (e.g. this is what Amazon does for any user regardless if they are logged in or not). With the first two cases, we can state that this type of personalisation is often cumbersome and labour-intensive, the problem is how to get that information from users (which creates cold start problem for users, and is also related to cold start problem of resources).

Also, not all the users are always interested in doing the work of filling in a profile or rating resources. In one of my studies (Vuorikari, Sillaots, Panzavolta, Koper, 2009) we found very different types of users behaviour, in this case related to how ppl tag (which could be also used for collaborative filtering).

About 33% of ppl tagged content (the arrows going away from the user group in the image), 32% used tags for searching but did not tag themselves (the arrows going towards a user), whereas 35% of ppl did not tag nor used tags at all (in the image the arrow going from LOM towards the group indicating that they used LOM based search methods only).

In this case, I'm interested in the 32% group, who clearly got benefit from tagging that other folks did, but did not do any work themselves (kinda freeriders, if you wish, but I don't mean to be negative).

If we were to use the tagging and bookmarking information to construct a user profile of the users and further use it for personalisation (no: 1 and 2 in the list above), we would be only able to do it for the group of taggers, i.e. 33%. The benefits could only be reaped by that group too, since for the rest of them, we have no profiling information to be used for personalisation.

However, a social tagging system is more about creating "personalisation" for all the users of the system, regardless if we know anything about them and they are willing to put in the time to create a profile and feed it in. It's about making social navigation trails visible to every user of the system, instead of going for "personalisation".

What I think is really cool is that we've shown that we more than doubled the amount of ppl who took advantage of contributions, i.e. from 27% who tag and use tags for navigation to 59% (that is adding the group of 32% who use tags but don't tag).

My assumption is that the rest would not even care about recommendations, etc., as they seem to formulate their searches in a rather acknowledgble way (40% formulate advanced searches; 38% only browse categories and 30% do both).

---------------------------------------------------
Some other thoughts about personalisation:

  • There are things that I dig, like Amazon, when it tells me "ppl who bought this book also bought xx". Thing thing is, though, that is the stuff that generally makes any user's life better on Amazon. It is not that they personalise the thing for me only, the unique Riina, the one and only, but it is something that makes any users experience on Amazon better.

  • the problem with personalisation often is that there needs to be a detailed profile of you that is based on detailed user model that is based on some abstract model that some obscure committee came out with in the 70's. Ok, that's maybe a bit exaggerated, but you get the point - there is a model where you are fitted.

  • What if I don't want to be personalised? What if I want to do the same thing as my buddies do; listen to same music as they do, study the same stuff as they do and go shopping with them? I want to share my life and experiences with other people around me because, guess what, through that type of sharing and doing stuff together, I feel related to them, I have things to talk with them and we form a community together. And it is really important for me to be part of that community, because it's part of who I am and helps me to reflect on what's out there.

  • Is peronalisation really personalised? It's not actually. By making things personalised to me, what is actually happening is "un-personlisation" of me. My taste is guided to the direction of all the other users, so I am actually being socialised! My personal music recommendations are actually very similar to other listeners, and eventually, it's all going to be the same taste!! Of course, unless there is randomness which offers serendipity.

  • Lately with all these micro-messaging things where ppl post their "mood" or what they are doing online ( e.g. I should be tweeting right now: "I'm writing a blog post" and simultaneously have it up on my Facebook), it's kinda funny that they feel this urge to yell out to all what they are doing in their über-personlised world.

Monday, June 22, 2009

Wiley calls it “dirty secret” of OER

Just picked up a fresh PhD study by S. M. Duncan from USU, a student of D.Wiley's. The study is called Patterns of Learning Object Reuse in the Connexions Repository. The punch line is that there is very little reuse of LOs among the repository studied.

What new? Similar findings have been discovered here in Europe (end elsewhere) for a while now. Ochoa (2008), for example, found in his PhD dissertation that reuse in general remains low, about 20%, across all sizes of collections. This was interesting not only for how low the reuse is (20%, common!), but also because since forever folks have been saying that resources with smaller granularity are more reusable, as they lack context, etc (insert here the infamous graph of "modular content hierarchy", the most used LO). Well, according to Ochoa (2008), this was not the case.

I also looked at the reuse on 2 different platforms: LeMill and Calibrate from European Schoolnet. My twist was to study the cross-boundary use and reuse, i.e. teachers reusing learning resources that are in a language other than their mother tongue and originate from different countries than they do. I used the same reuse definition as Ochoa (2008), which basically is the same as in Duncan's study.

The finding was that the general reuse was around 20%, but NOT across all collections. For example, in LeMill, "Multimedia material" was used more often, but in Calibrate, the smaller granularity was seldom added to Collections. The cross-boundary reuse was notably less (37% to 55% of it). Moreover, in some of the collections only around 10% of resources were ever added to a collections, which makes you really think hard about the efficiency of this all..

Anyway, the good news in Duncan's study is this:

There was a common author in 3,722 module uses, while there were only 1,013 module uses where there was no common author. This means that modules were included in collections 3.67 times more often when there was at least one person in common with both the module and the collection.p.32


So if people know each other, they are more likely to reuse material from each other! This shows that social is important when we are talking about the use and reuse of learning resources! This is similar to what I am saying in my PhD thesis, which hopefully will come out one day soon. My twist of course is that tags can make those social connections between people, and by taking advantage of these underlying social connections, we can make the learning resource discovery much better - and hopefully also more useful for teachers.

Vuorikari, R., Koper, R. Evidence of cross-boundary use and reuse of digital educational resources. Link to a revised version of the paper, not reviewed yet!

Tuesday, May 05, 2009

Challenges and lessons learned from Social tagging in MELT

Social tagging in MELT, how do we want to take the social tagging work forward

Social tagging of educational resources potentially offers new ways for:
  • Individuals to
    1.1) better manage their digital learning resources that reside in different repositories and platforms, and

    1.2) discover and access new resources from different contexts (e.g. different language, educational system) through tags and other users.

  • LOR managers to
    2.1) get third party metadata on learning resources (either the ones that already reside on their repository, or the possible new ones to be added to collections,

    2.2) create affinities (e.g. link structure) between separate pieces of resources (either on their own repository, or the ones that reside on other repositories on the federation or on the Web) that were not cross-referenced before.

  • In the MELT project so far, we have only been able to see the peak of these potentials emerging. We list issues that we see important for future work in the field, for the clarity, we only list one of the main issues for each topic:

    • 1.1 To fully support users in their knowledge management task on digital learning resources, the bookmarks (including title, url and tags) should be exportable in standard Webfeed formats. This would allow users to access and manage their MELT resources as part of their other resources collections, whereas now users need to be logged on to the MELT portal to do this.

    • 1.2 Pivotal browsing of social bookmarks takes advantage of the affinities between the user, resource and tags. In the MELT context, more metadata could also be added to support pivotal browsing, such as the country of the user, interest topics; resource metadata such as multilingual indexing keywords. This would allow novel ways to access resources that other users have already discovered within the federation, and thus build on users’ social interactions and co-construction of knowledge.

    • 2.1 Tags by end-users on the MELT portal have been shown to be of good quality as additional metadata descriptors of resources. We have enumerated possibilities of metadata ecology that the use of multilingual Thesaurus can offer to a federation such as LRE. Apart from working on ways to automatically generate LOM from tags, we urge on using the hierarchical structure and multilingual features to leverage user-generated tags.

    • 2.2 Why not do Google for learning resources? Using PageRank-like algorithms on a learning resource repository or federation has been impossible for a number of reasons, the most important is the lack of a link-structure that cross-references resources. Tags, creating underlying connections between seemingly random pieces of content in different languages, on repositories in different countries and other platforms on the Web, rely on humans’ subjective idea of its importance for a given information seeking task. Using this new, emerging link-structure with tags as “anchor texts” offers totally new ways to “organise the world's learning resources and make them universally accessible and useful”. A new tag line could be “From teachers to teachers”.

    Monday, January 12, 2009

    Discovering cross-boundary resources on the portal

    This is continuation for the previous post, the same data-

    The question now is how and where do users discover cross-border resources? By this I mean that the user and the content come from different countries, and/or that the content is in a language other than the user’s mother tongue.

    One challenge is discovering resources in general (previous post), another one is to discover resources that are in a different language and come from different countries. This is the case on our portal, so I'm interested in how to facilitate the discovery process that involves crossing those boundaries (language and national, which might have implications to the educational content of the resource), and hopefully make it more efficient.

    My take is social information: making it readily available to all to leave cues from other users. I bet that this should facilitate the discovery process and thus also make it more efficient (more resources and faster).

    So I looked at the previous data and the cross-boundary bookmarks from users. This is relative to the user, of course, so with every bookmark I compare whether the resource and the user come from the same country and/or is the users mother tongue different from the resource language.

    41 out of 48 active users had bookmarked cross-boundary resources. They had added
    • 299 distinct resources into their collections 350 times. Out of these resources
    • 163 were cross-boundary resources, which had been added to their collections 198 times.
    • This means that 55% of distinct resources obtained during the period of 1.5 months were cross-boundary resources.
    So it was interesting to see how many of these resources had Social information added to them?

    Table 1


    Table 1 shows that about a third of the bookmarked cross-boundary resources had Social Information on them, they were either bookmarked by previous users or existed in the "travel well" list. This is cool! Although we cannot say that the users only discovered these resources because of Social Information, it is important to know that it has had added value for users to discover them. I also found out that about 10% of bookmarks on resources with SI had been previously bookmarked by someone from the same country as the resource was. The fact that they were bookmarked, although still almost dismissible small (10%), is still good news for SI and social navigation based on it.

    I then looked at where the resources are discovered: 62% of cross-boundary discoveries were done in Search Result List (SRL) whereas 37% took place in Community searches, most of them in the tag cloud (30%) and 7.5% on the Travel well and Most bookmarked lists (only one case in the latter :(

    As comparison for the resource discoveries that did not cross any borders (e.g. German teacher found German resources), 90% of them took place in SRL. So it seems that for discovering cross-boundary resources the Social Information is important, as it allows users to do Community searches to discover these resources. We do see, though, that cross-boundary discovery is efficient also within the resources that cannot leverage the previous user experiences, as 63% of cross-boundary resources do not have any SI.

    Interestingly, when we look at Table 1, we see that for the resource discovery that does not imply any cross-border action, users do not seem to care much for SI. Actually, more than 90% of these bookmarked resources had no SI. This is cool, as it seems that we need users to discover and annotate resources among their comfort zone (e.g. national and regional educational material in their own language) in order to make it more readily available to others.

    Another thing that I've looked at are these measures for
    • usage coverage within the repository,
    • how many resources are shared among collections (e.g. Favourites) and
    • what I call the pick-up rate, this is how many of distinct used resources are reused. This could happen when someone discovers a resource that has SI related to it (e.g. in the travel well list or from other user's favourites, or just picks it up from SRL based on someone else's annotations). These were discussed in more details in the last paper.
    Table 2



    Table 2 shows these measure in LeMill, Calibrate and delicious, additionally, the gray column indicates this 1.5 month trial in MELT and the last column has all the data from MELT, which includes the pilot teachers, staff, partners, etc.

    We can see that as the initial amount of resources is so big in MELT, we still cover only a minimal amount of resources, from 1 to 2 %. This does not even include the assets, which more than doubles the amount. Used here means that the resource is added to Favourites once (reuse more than once).

    We can see, though, that even if the resources coverage is not that high, there is still quite a lot of sharing among used resources. This figure still remains low with 1.5month trial (about 15%, same as in Calibrate which did not make the SI available!), but if we look at all the usage so far, the sharing is at 43%. This is somewhat artificial, though, as there is a lot of staff use, but still I hope it indicates that making SI available on the long run helps sharing resources (or, I need to look better solutions on the portal for sharing, which is also planned but super delayed because of all the other dev programme).

    We can already see that the pick-up rate is higher in MELT trail than in the 3 other platforms that I've looked previously. This is an indication that SI works, I hope.

    Btw, I could not find any correlation between the act of putting the resources in Favourites and whether they have Social information related to them (I got some lousy 0.18 even if I removed 2 outliers who outperformed anyone else). Also, in some previous test I had hard time finding significant changes, so I might need to seddle with increased reuse rate, which has previously been shown to be abotu 20% across collections. I found that it was about half of this (or the same as general reuse) in my previous paper.

    Tuesday, January 06, 2009

    Users on the portal: consuming and contributing

    I continued the previous study with more data: this time I took all the logs from Oct 1 to December 18 2008. The idea, again, was to see how much the previous annotations (ratings and bookmarks) would "guide" the choice of new users.

    The datasets:

    1. All bookmarks and ratings in MELT until Oct 1 2008. This comprises of 565 distinct resources. We call these resources with Social Information (SI), as it is something that is shared and made public by users.
    • 88 of these resources were on the list called "Travel well" resources. They are made available directly from the portal front page. They were recorded to the system by a user called "EUNRecommender" suggesting that these resources should be of interest and good quality. Additionally, annotations from other users were made available.

    • 477 of these resources were annotated (ratings and bookmarks) by previous users. These were either the pilot teachers, project partners and staff in the office. The top bookmarked and rated once would appear on the "Most bookmarked list", also accessible from the front page. Similarly, annotations from other users were made available.
    2. I took server-side logs from the period of Oct 1 to Dec 18 2008. At the end of this period the portal had 340 users registered, out of which some people never used the portal when logged in and a number was project related staff. We excluded those people from the logs and were left with 168 users, out of which 82 had clicked on a resource on the portal at least once. Additionally, users who do not log in are recorded, so our click-though includes many more users. This group had:
    • clicked ("played") on 1711 distinct resources (all users);
    • out of which 974 were clicked ("played") by logged-in users (82);
    • bookmarked 294 distinct resources 351 times;
    • rated 323 distinct resources 385 times;
      = 394 distinct resources annotated.
    The method for this study is a manual log-file analyses using my own defined logging scheme (see here).

    Questions: How and where on the portal do users discover resources, do they rather discover new resources or re-discover once that have been annotated by others? Are the suggestion lists like "Travel well" resources or "Most boomarked" effective, i.e. are many resources discovered through them?

    We split our understanding of "discover resources" meaning two different things and look at the separately. Discover meaning:
    1. finding the resource on the portal, clicking on it (i.e. play). This is also called implicit interest indicator;
    2. finding a resource and creating an explicit marker on the resource in terms of rating or bookmarking with tags (explicit interest indicator).
    The first one is important to know how big coverage the resources that have been "touched upon" or "hit" cover from all resources in the repository (use). The second one, is important as it creates explicit maker which can be used for cues and social navigation purposes. Moreover, the second one can be used as a proxy for reuse. For reuse we use the following definition: " resource is integrated in a new context with other components, and when this occurred more than once, we consider the resource reused".


    1. Who discovers resources?

    Out of 82 active users, 63 had annotated a resource at least once, 75% had annotated more than once. Average was 11.6 annotations, median 4. The top two users annotated 120 and 108 times, the following users 50 times. There were 36 users who bookmarked and rated, and 14 users who only bookmarked and 13 who only rated.


    2. What kind of "actions" take place?

    By actions, in this case, I mean playing the resource (click on the link), bookmarking and/or rating it, but also viewing evaluations related to this resource or checking the "Favourites" from the user who had discovered an interesting resource.

    Table 2 presents data on how the users (described in 2. above) interacted with resources. The left column explains where on the portal this action took place, whereas the top row indicates the action. There were total of 3542 actions related to this.

    In general, this group annotated 75% resources in SRL, 15% in tag cloud, 4% on the Travel well list, 3% in their own favourites (ratings) and only 1% on Most bookmarked list.

    Table 2


    Let's first look at what happens in the Search Result List (SRL). Table 2 shows that most things happen when users are here: resources are played more than 2000 times (2/3 of all plays), most rating (70%) and bookmarking (73%) of "virgin" resources also take place here. We can also see that some small amount of resources with SI are rated and bookmarked from SRL (less than 10%). This indicates that some users paid attention to cues made available by previous SI and found them useful.

    Mostly, though, it's "virgin" resources. This is a good news, in a way, as I was worried that users might be lazy and not discover resources that other users have not annotated before. On the other hand, most users click on resources and never annotate them, so we are left to guess whether they liked them or not... only 13% of these clicks lead to rating the resource (268), for example.On the average, 71% of virgin resources "hit" get annotated, and 29% of resources with SI get annotated.

    Table 2 also shows that the tag cloud is a good catcher of ratings (15%) and taggings (15%), whereas the "Travel well" list is a real hit catcher. 13% of hits on resources with SI are played here. This list was comprised of 88 resources, out of which only 54 got played. The average was 6.8 plays/resource (median 3.5), but in reality, some got played a lot more than others. 30% received more than average hits. The highest were 30, 29 and 28 (great, just checked the top dog, and it's a dead link :(.

    The story looks worse for actually bookmarking and rating resources from the Travel well list, only 18 of them got hitched (20%). 2 got bookmarked twise and two rated 5 times (no overlap). So, I suspect that we did not succeed that well in creating appealing "recommendations". I'll get back to that point at a later stage. A rather effective features of TW list was the use of "view evaluations" and "view other users", about 40% of both were generated here, so they helped users to social navigate the portal.

    Compared to "Travel well" list, the "Most bookmarked" resources did not seem to have the same effect of chatching eyeballs. This is bizarre, as both these are available on the font page, however, "Most bookmarked" are a click away on the tab. As 84% of these hits come from users who are not logged in, I think it might be lots of clicks from our testing period, but also it's possible that this is only what users experience on the portal and then fly away (forever..?).

    Lastly, table 2 shows that some ratings take place in Favourites. That's good, since we intended favourites to be the place for that. I imagined that teachers first want to use the resource, so they put it on thier favourites and then come back later to rate it. Well, users know better, they seem to take about a second to view the resource and bookmark it on the SRL. I will have to look if there is any qualitative difference between ratings of the ones in SRL and Favourites?

    3. Where do these actions take place?

    I was also interested in seeing what kind of search methods do users choose to use. Possibilities offered on the portal are the following:
    • Explicit search: using traditional search box with text or advanced search options. This results in resources on a Search Result List (SRL) where users can view the metadata about the resources as well as the annotations by others. Searches within results in SRL can also be refined, or ranked either by popularity or ratings. Also "Browse by category" results in this.

    • Community browsing: These include browsing the tag cloud, and examining bookmarks created by the community. In our case these are lists of most boomarked items; travel well; tags; other people Favourites.

    • Personal search: Looking for bookmarks from one's own personal collecction of bookmarks (Favourites)
    Table 3.


    Of all searches 76% were Explicit searches, where 21% Community searches and 2% personal searches.

    What comes to spearing actions on other things on the portal, we can observe some differences between the logged-in and not logged-in users. Whereas not logged-in users spend most of their actions on searching (60% + 17%) and playing resources (23%), the logged in users have a more variety.

    The logged-in users search differently, they use far less the Explicit search function (28%) and also spend less time on the Community search (only 8%), but additionally they have the Personal search function available (3,5%). The difference here is that these users can interact with the resources that they have found earlier ("keep found things found"). There is quite a huge difference between how much actions are spent on searching between these groups, 77% vs. 40%. Logged-in users also play less resources (18%), but still "outsmart" not logged-in users in terms of searches returning plays:
    • the first group spears 3.4 searches to play one resource, whereas
    • the logged-in users spear 1,98 searches to play one resource.
    This apparent inefficiency is partly due to the fact that we are still testing the server and probably many tests are done when the user is not logged in. However, this is something to remark for now and keep the eye on (check: can I omit the searches from the office using the data from Analytics).

    A major difference between the two groups is what I call contributing actions. We already have seen that 27% on resource with SI get played and that on the average 13.3% of searches take advantage of Community browsing. We also saw that some 6% of annotations on the SRL are on resources that have Social Information related to them. So where does this social information come from?

    In Table 3 we see that contributing actions cover about 40% of all the actions by logged in users. This is about 16% of all actions during this period! I will come back to this point later trying to make a picture of what the input is by a group of logged-in users and how it can be taken advantage of by other users of the system.

    4. What resources do get played?


    Users can access 30116 learning resources through the portal, and many more assets. During the trial period, 1547 distinct resources got played 2828 times by all users. That make an average of 1.83 plays/resources, but in reality some get a few hits and a few gets many hits (median: 1). 27% get more than one hit, most hits were 35 ( strangely, this resource was not even bookmarked "123216875").

    I was interested in what happened with resources that had Social Information vs. the ones that were "virgins", not yet annotated by previous users.

    Of the total plays (2828 times) 73% were on resources without SI and 27% on resource with SI. If we look at it from all resources that were made available, out of 565 resources with SI 34,7% were played at least ones, whereas from all other resources (29551), 4.57% got played at least once. Also, some of the newly annotated resources got hits right away, we have 54 resources that got played 70 times, most likely thanks to their new annotations. Some of these newly annotated resources (32 cases) prompted 99 further annotations from the new users. These 99 annotations were made both in SRL (69%) and in social navigation areas like tagcloud and Favourites (31%). This shows that these new annotations became useful to other users right away, actually more so than the previsouly annotated resources, out of which only less than 10% were found of use by these users (31% vs. 10%, see Table 4).

    It's quite interesting, though, that only about one third of resources with SI got played. I would have thought that it is more. This might be due that our search result list is not ranked to start with. It is possible for the user to re-organise the results by popularity or ratings, but actually we do not know whether this happens (currently have not found a way to log it). I guess the other side of the coin that surprises me is that users clicked on so many resources that did not have any annotations on them. I guess it's a good sign of curiosity :)

    Interestingly, we find that users who are logged-in discovered less resources than the others. Out of all these resources (n=1547), 68% were played by users who were not logged in and 44% by logged-in users (last row in Table 1). The same can be observed for resources with SI and not.

    Table 1.



    5. How many annotations per resource?

    Of the total amount of annotations received during this period, we can count that 394 distinct resources received 734 annotations, they were half and half ratings and bookmarkings.
    • 322 resources received 383 ratings. 13% received more than one rating, average 1.18. Top amont of ratings were 5 ratings, which was only on 1 resources.
    • 294 resources received 350 bookmarkings. 14% received more than one bookmark, average 1.19. Top amont of bookmarks were 5, which was only on 2 resources.
    58.5% of annotations were on new resources, and 41.5% on old ones. They were distributed rather differently, the "old" resources received on the average more annotations than the new ones.
    • 38 Travel well resources received annotations 99. 63 of them received more than one annotation, average is 2.6 each. The top one received 12 annotations.
    • 97 previously annotated resources (this includes 32 resources that were discovered during this period and annotated later by other new users) received 209 annotations. Average is 2.15 annotations per resource.
    6. Story of the Travel well list

    Barely about 15% of the 88 resources on the list were re-discovered by the new users! Previously we saw that these resources received a good amount of "hits" and eyeballs, but not so many of them actually resulted in ratings, and when they did, they were not equally distributed among the resources (Table 4). Only 18 of the annotations took place on the TW-list, otherwise, an additional 20 resources from this list were annotated on the SRL and tagcloud. So the success-rate of these recommendations was less than 50%! Can barely call these recommendations. I will look later which ones were "thumbed up" and which ones downed. As with the previously bookmarked resources, ahem, maybe the taste differs between the two sets of users.

    Table 4



    The resources on the "Travel well" list received many more hits (average 7.12) than just "any previously annotated resource" (average: 0.68).

    Sunday, January 04, 2009

    New users on the portal and resource discovery

    I looked at 18 new users on the portal, and studied resources that they bookmarked. I wondered how many of these resources had previous annotations by users? By annotations I mean that previous users had added ratings on them and Favourited these resources. If this is the case, it's clearly shown on the portal.

    But are these annotations persuasive? Do they help users make their decision better or faster?



    I had two sets of data:
    • last 3.5 months (Aug to Nov 17 2008)
    • Nov 18 to Dec 18. This are my 18 users who had bookmarked 114 resources.

    Out of these results, it seems that 2/3 of the resources that these new users bookmarked had no previous annotations on them! I have to verify this finding, because currently I lack data from March to July to see what was bookmarked then.

    Table 1


    Anyhow, let's see what the current mini-study holds. Out of the third of resources that were discovered by this group, about 25% had previous annotations on them. They were mostly done by users before this group got initiated on the portal, however, some were also discovered thanks to the bookmarking by this group.

    Only about 8% of resources were discovered through a special list called "Travel well" resources. These resources have been added there by "EUNRecommender" which currently is hand operated, but mainly based on picks by other users from at least two different countries. I find this figure rather surprising, as this "Travel well" list is the first thing that the user sees when they come to the portal.

    Anyhow, I find it cool that 25% of bookmarks by this group were resources that had previous annotations. What we cannot say, though, is whether these users could have found these resources without these annotations displayed publicly on the search result list. However, I think that my measures like "pick-up rate" and "overlap" among Favourites will help me sort that out.

    Here is a visualisation of this. The huge node is "sos-rec" which means that these resources were previously annotated (rate, bookmark). What is called "EUN Recommender" are resources from the "Travel well list". Other than that, the new users are the nodes which are connected by edges to resources that they have bookmarked.

    Right from the bat, we can see that 5 users had taken their totally own trails and bookmarked resources that no one had bookmarked before - quite cool!

    Additionally, there are few users on the lower right hand corner who are only connected to the whole graph through one resource (user: 192682) and (user: 217391).

    Friday, September 12, 2008

    Cross-boundary ranking of learning resources

    Based on the idea of Interest Indicators, like social bookmarks and ratings, I've looked at the data so that we can make the cross-boundary resources better available on the MELT portal.
    The aim is that we can, based on previous users' behaviour :

    a) make separate "travel well" lists of resources that have a potential to cross-borders better,

    b) use this information to rank resources better in the normal search result list,

    c) allow users search for resources that have a good "travel well" value (e.g. give me resources in math that can cross-borders)

    This is the data that I'm using (table below) and this is how I've defined cross-boundary (e.g. cross-country and language) learning resources before. Using that definition I have manually verified the number of cross-country resources. In the dataset, about 82% of resources were cross-country.



    Now, we have a problem, though. On our MELT portal we do not have information about the country where the resource originates from. Dah!

    This is a big blunder (in my opinion) in our Application Profile, we have not defined the country where the resource originates. We do define the provider, and the country information could be inferred from the provider, but it does not always work.

    For example, one of our providers has frequently metadata about resources that do not originate from the same country!

    I've experimented with the data using the information that we have on the portal, which is LOM about the resource including the language of the resource. As we also know the mother tongue of the registered users, this gives us a kick.

    In the table below we can see the coverage of cross-boundary actions that we can get on resources without using any manual labour or verification of the country or language. As a base-line, with manual verification I found that 82% of the actions concerned cross-border rating or bookmarking of a resource.



    The first row represents the cross-language resources (i.e. user's mother tongue is different from the resource language). Just using this information, we get about 65% of resources right, as opposed using manual checking (82%). I think it's pretty good, I'd settle for that! (although I have to look what kind of material was left out!)

    The two other comparisons in the table are based only using information about users' previous behaviour. These would be:
    • rating > 2
    • bookmark
    Only using information regarding bookmarked and rated resources results in a lousy coverage of around 20%. The problem is that 25% of resources bookmarked or rated are on more than one resource, the data still is very sparse.

    Anyway, I want to use that information to "cross-boundary rank" the resources. As we do not know the country where the resource comes from, my work-around is based on countries where these users come from.

    Here is a visualisation about resources that have been bookmarked or rated by users (see also ManyEyes link below). We can see the orange node in the middle, a learning resource called "Five Days in New York..". We see 3 edges leading out to Finland, Belgium and Hungary. This means at least one user from each of these countries has bookmarked the resource!

    So, even if we do not know the origin of the resource, we know that it has users from 3 different countries. I can infer that it is a cross-boundary resource.

    As most likely one of these 3 users come from the same country than the resource comes from, I will minus one country out of the total of countries: (number of countries -1)

    My cross-boundary rank will be the following:
    • Count the number of ratings grater than 2 and/or bookmarks for a resource (actions). Give each action one point
    • Count the number of these users and give each user one point
    • Count the number of user countries of origin. Give each country one point and then minus 1
    • Compare the mother tongue of each of these users to the language of resource. If they differ, give one point/mismatch.
    Then, count the following:
    number of users + number of actions + number of cross-language x (number of countries -1)
    Let's take the above resource "Five Days in New York.." as an example
    • Count the number of ratings grater than 2 (3) and/or bookmarks for a resource (5). Give each action one point. (8)
    • Count the number of these users and give each user one point (5)
    • Count the number of user countries of origin (Hungary, Finland, Belgium). Give each country one point and then minus 1 ( 3-1=2)
    • Count the number of user mother tongue (hu, nl, fi). Compare the mother tongue of each of these users to the language of resource (en). If they differ, give one point/mismatch (3).

    • number of users (5) + number of actions (8) + number of cross-language (3) x (number of countries -1) (2) = 32 Travel well value
    This way you can count a value of "travel well" for each resource that users have previously interacted with on the portal. The value will always be an integer, which is important from the technical implementation point of view (in Lucine index it apparently needs to be an integer).

    The down side is that we'll have a huge cold-start problem. As I said, our data is very sparse. To seed the system, I actually still manually check the new resources that users have interacted with and make a fake bookmark on them so that it looks like it has at least two users from 2 different countries. This way the resource gets a "travel well" value counted and appears on the "travel well" list and is better ranked, etc.

    Of course, at the end I will evaluate how this treatment affects on users, do we, for example, see a big amount of bookmarks on these resources that I have been able to count a travel well value?

    You can see a visualisation here. This is based on on user's country of origin.

    Tuesday, September 02, 2008

    Emerging search patterns on learning resources

    I am hugely inspired by the stuff from J.Feinberg and D.Millen, especially by the studies that they've done on doger, the IBM internal social bookmarking service. I must admit, though, that I had missed on it a bit, I cannot believe! Anyway..

    This paper is really interesting, Social bookmarking and exploratory search (2007), not least for the reason that it offers a very interesting, almost similar study design that I am planning on my log files and search pattern analysis on the MELT portal. (Great minds think a like ;p yeah, right..)

    The study design is a field study of a social bookmakring service in a large corporate (IBM) with quantitative data (click level analyses of log files and boomarking data) and qualitative data like interviews.

    And, this is the data that I collect for my study (Table 1). Quite simlar!

    They use a categorisation of search that I will adapt to my usage:
    • Community browsing (Examining bookmarks created by the community. In my case these could be lists of most boomarked items, travel well; tags; other people Favourites)
    • Personal search (Looking for bookmarks from one's own personal collecction of bookmarks)
    • Explicit search (Explicit search using traditional search box)
    A few days ago I looked at the first logs from Melt. Here is the run-down. I have not used the same types of searches, but I will explore them for the later usage. But basically what you can see:

    There were 512 search events:
    • 41% Explicit searches (adv. search)
    • 34% Community browsing (tagcloud)
    • 25% browsing categories
    What was called "click-through" in this paper is when the particular navigation path resulted in a page view. This is what I call view resources. We can see that there were 538 page views, which implies that most likely users have clicked on more than one resource as a result of a search.
    • 74% of resources were viewed in search result list (srl), this means that they were results from advanced search and browsing categories

    • 20% of resources were viewed as a result of community browsing (e.g. tagcloud, lists)

    • 5% of of resources were viewed as a result of personal search (e.g. in Favourites)

    How many searches resulted in viewing resources? I have to verify this
    • 85% of explicit searches and browsing categories
    • 65% of Community browsing (tag cloud and lists)
    Millen et al. 2007 speculate that in their study, the a higher click-through rate indicates a more purposeful searching, whereas the community browsing was used more as an exploratrory search activity. I will need to keep my attention on this and whether I can make similar conclusions.

    Most likely, anyway, I do not use click-through as an indication of intenet of using a learning resources, what I find most interesting in my study will be how many of the search activities result in a bookmark and/or rating. That is much cooler in my mind than viewing the page. Actually, I've noticed that in our system users view resources a lot, but they do not necessarily show any further interest on them.

    That is why I use both implicit and explicit Interest Indicators. By Interest Indicators I mean Explicit Interest Indicators like ratings as a subjective relevance judgment on a given learning content, Marking Interest Indicators like bookmarks and tags on educational content, and Navigational Interest Indicators such as time spent on evaluating the metadata of educational resources, as well as Repetition Interest Indicators as categorised by Claypool, et al., 2001.

    In this small study we can see that 19% of viewed resources ended up in users' Favourites. Again, ratings were much less, only about 6% of viewed resources ended up being rated. In the study from IBM system, they had 34-39% as high click-through. Will be keeping my eye on that.

    My advisor asked me whether I could find search patterns in my logs. I think I could. In this paper (Millen et al. 2007) they do that :) and here is how: "Looking for Patterns using Cluster Analysis"
    To better see the patterns of use, we performed a cluster analysis (K-means) for the different types of search activities. We first normalized the use data for each end-user by computing the percentage of each search type (i.e., community, person, topic, personal, and explicit search). The K- means cluster analysis is then performed, which finds related groups of users with
    similar distributions of the five search activities. The resulting cluster solution, shown in Table 4, shows the average percentage of item type for each of four easily interpretable clusters.
    Should not be too hard :)

    Millen, D., Yang, M., Whittaker, S., Feinberg, J., Social bookmarking and exploratory search (2007). In L. Bannon, I. Wagner, C. Gutwin, R. Harper, and K. Schmidt (eds.).
    ECSCW’07: Proceedings of the Tenth European Conference on Computer Supported Cooperative. Work, 24-28 September 2007, Limerick, Ireland

    Friday, May 30, 2008

    Visualising networks of learning resources

    I'm looking at the first dataset of bookmarks from MELT portal. Here you can see some of the first descriptions created by using Many Eyes. Click on the interact button in the pic and it loads. This is a treemap visualisation of the bookmarks that users so far have found.

    What do you see here? You first see boxes in different colours. They are "boxed" by the user IDs. The bigger one is, the more learning resources this person has bookmarked. If you hoover your mouse over the boxes, you can see the ID of resources. These, of course, do not mean nothing to you now, but imagine if they were links to resources?

    Next you can explore the data a bit further. Drag the mother tongue box on the top of the graph to the first place. Now, the boxes are displayed by the languages spoken by users. You'll see that Hungarian speakers have been busy on the portal, they have the most bookmarks.

    Third, you can explore further by dragging the obj_lang to the first place. This shows the languages in which the bookmarked resources are. Interestingly, it turns out, most of these resources are in English. However, the diversity is there to be observed: users have found resources in many different languages useful.

    Let's go further. The next one is a network diagram. If you click on "click to interact" you can also zoom into the visualisation.

    What do you see here? It's a network that consist of: user mother tongue and the learning resource that those users bookmarked on the portal. You see 4 quite big vertices, which are the mother tongues of the users.
    ..network consists of a set of objects called vertices connected by edges. The visualization of the network is optimized to keep strongly related items in close proximity to each other. In this way, the overall arrangement of vertices in the network is very telling of the structure of the connections between vertices (vertices that are far away are weakly related to each other).In this visualization, the size of a vertex is proportional to the number of edges emanating from it.
    Take the Hungarian speakers, for example. They are the ones who user the portal most, and have actually bookmarked a fair amount of resource. At the end of each edge you can see an ID number. Those are the ID of learning resources that these teachers have bookmarked. The same goes for Finnish speakers, Dutch speakers, etc.

    Interestingly, we can see from this visualisation that not many resources are shared among the users from different language groups. A few are, though: take, for example, the LeMill resource that is visualised in orange in the image here. It has edges linking it to Finnish, German and Hungarian speakers. I counted 14 resources in this small dataset that were shared by users from different countries, that's about 13% of resources.

    This type of resources are what we call "travel well" resources, as they can cross borders. In this case those borders are lingual. The resource also acts as a bridge between these different language communities. If you look at the resource in question, you'll find that it is to teach English (as
    foreign language) and it is in English. Thus, it is not that surprising that it is well accepted in many language communities.

    Finally, I also visualised the languages of learning resources instead of the resource ID. You can find it here. As you see from the image on the right, I have highlighted the languages of resources from Dutch speaking users. They have been pretty busy finding resources in all kinds of languages!

    Tuesday, May 20, 2008

    Call: WORKSHOP ON SOCIAL INFORMATION RETRIEVAL FOR TECHNOLOGY ENHANCED LEARNING (SIRTEL'08)

    Good news, we are ready to roll out the call for contributions for our 2nd workshop! This time we are planning more time for discussions and brainstroming type of exercises that participants can lead! This was the feedback from last year, so you see that we are taking it seriously :)

    Check out the format for contributions; Research papers and System Demos are the more conventional stuff that we welcome, whereas Hands-On proposals are there to let us all loose and to think how could we use ideas from some exiting, existing systems to enhance and support learning and teaching. Oh then, there are of course the Pecha Kucha talks. That makes me really curious: someone said that they would not really work with computer science. I hope we are able to prove that wrong ;)


    WORKSHOP ON SOCIAL INFORMATION RETRIEVAL FOR TECHNOLOGY ENHANCED LEARNING (link)

    in the 3rd European Conference on Technology Enhanced Learning (EC-TEL08), Maastricht, The Netherlands

    IMPORTANT DATES

    • Contribution Submission: June 29, 2008
    • Results Notification: August 3, 2008
    • Camera Ready Submission: August 31, 2008
    • Workshop date: September 17, 2008
    • Main conference dates: September 18-19, 2008

    CALL FOR WORKSHOP CONTRIBUTIONS

    After the successful first SIRTEL workshop last year, we are delighted to welcome
    exciting new contributions for the 2nd Social Information Retrieval for Technology Enhanced Learning (SIRTEL) workshop:

    • Research papers
    • System Demos
    • Hands-On proposals
    • "Pecha Kucha" talks*


    RATIONALE

    Learning and teaching resources are available on the Web - both in terms of digital learning
    content and people resources (e.g. other learners, experts, tutors). They can be used to
    facilitate teaching and learning tasks. Developing, deploying and
    evaluating Social information retrieval (SIR) methods, techniques and systems that provide
    learners and teachers with guidance in potentially overwhelming variety of choices remains to be tackled.

    The aim of the SIRTEL’08 workshop is to look onward beyond recent achievements to discuss
    specific topics, emerging research issues, new trends and endeavors in SIR for Technology Enhanced Learning (TEL). The
    workshop will bring together researchers and practitioners to present, and more importantly,
    to discuss the current status of research in SIR and TEL and its implications for science
    and teaching.


    TOPICS OF INTEREST (but not limited to):


    Technology Enhanced Learning (TEL) and Social Information Retrieval (SIR) techniques such as:

    • Recommender systems
    • Social collaborative searching, browsing and sharing of queries
    • Social network analysis
    • Game-theoretic approaches to select learning materials and learning partners in the long tail
    • Social bookmarking and tagging, folksonomies
    • Annotations, ratings and evaluations


    Concepts for Social Information Retrieval (SIR)

    • Defining the scope, purpose and objects of social information retrieval in TEL
    • Defining user requirements for the deployment of SIR systems in a learning setting
    • Current and new trends in SIR methods for TEL
    • Approaches to TEL metadata that reflect social ties and collaborative experiences in the field of education
    • Analytical modelling of strategic intentions in TEL communities
    • Interoperability of SIR systems for TEL


    Implementation of SIR in TEL

    • Methods and models of SIR in the area of learning and teaching
    • Social processes and metaphors in learning communities and social networks for searching, acquiring and sharing information
    • Pedagogical aspects of SIR in TEL; how to scaffold students, activity patterns, etc.
    • Integrating SIR services in existing learning platforms
    • Visualisation techniques to support SIR in TEL
    • Successful scaffolding techniques for SIR implementation

    Evaluation of SIR in TEL

    • Ideas on how can we get more empirical on evaluation
    • Best practices
    • Evaluation of the success and acceptance of SIR systems in the context of teaching,learning and/or TEL community building
    • Challenges and enablers
    • Evaluating the performance and measuring the effectiveness of SIR systems in learning applications;
    • Evaluation the user satisfaction with SIR system in supporting learning and teaching, etc.


    WORKSHOP SUBMISSIONS

    This year we base our call for contributions on last year’s comments, where the participants wanted more time for discussions, for picking each other’s brains and to forecast how SIR could be used in TEL. Apart from more conventional contributions, we also have new formats for you to consider!

    • Research papers (4-8 pages)
      to present exciting new work that is not mature enough for a long conference/journal paper. We especially value papers with focus on evaluating early results and making them available for further discussion among practitioners.

    • Work in progress and System demos (upto 4 pages)
      allow participants to share the basics of their SIR for TEL applications. Papers can be short (upto 4 pages), but also different ways using screencasting or YouTube-type recordings of the demo are welcome. Include also information also needed on how others can access your system and test it.

    • Hands-On proposals (1-pager)
      Got a good idea for a SIRTEL implementation? Toying with ideas for SIRTEL prototypes, either totally new ones or based on some existing application (e.g. Amazon, Flickr, Digg, ..)? Interested in “pimping-up” your current LMS or platform to support social networks?
      Create a little scenario and write it down so that others can follow your thinking. Put in a few screen shots to illustrate your point better. During the session, which you will lead, the participants will have their hands and brains-on your idea. The outcome will help you with requirements of implementations in a TEL setting. Early ideas welcome!

    • Abstract for Pecha Kucha (5 min talk)
      Want to share your discussion ideas on SIRTEL concepts with others? We are listening! To leverage on the face-to-face of the workshop, we invite you to submit an abstract for CP type of presentation-discussion moment which you will lead during the workshop. Your talk can be max. 4 minutes long, the participants will decide how much discussion will follow.


    Papers are to be submitted to: https://togather.eu/handle/123456789/274
    Accepted papers will be published online as EC-TEL workshop proceedings
    as part of the CEUR Workshop proceedings series.

    The two best papers of the workshop will be published in a special issue of
    the International Journal of Technology-Enhanced Learning (IJTEL)
    http://www.inderscience.com/browse/index.php?journalCODE=ijtel

    More information at the submission site. All questions and submissions should be sent to: sirtel @ cs.kuleuven.be


    PROGRAM COMMITTEE

    • Alexander Felfernig, University of Klagenfurt, Germany
    • Barry Smyth, University College Dublin, Ireland
    • Brandon Muramatsu, Utah State University, USA
    • Clemens Cap, University of Rostock, TBC
    • Frans van Assche, European Schoolnet, Belgium
    • Fridolin Wild, Vienna University of Economics and Business Administration, Austria
    • Hendrik Drachsler, Open University of the Netherlands, The Netherlands
    • Jon Dron, Athabasca University, Canada
    • Lisa Petrides, ISKME, USA
    • Marc Spaniol, Max-Planck-Institute for Informatics, Germany
    • Markus Strohmaier, Technical University of Graz, TBC
    • Martin Memmel, DFKI, Germany
    • Wolpers, Fraunhofer, Germany
    • Miguel-Angel Sicilia, University of Alcala, Spain
    • Nikos Manouselis. Greek Research & Technology Network, Greece
    • Oliver Bohl, Accenture GmbH, Germany
    • Rick D. Hangartner, MyStrands, USA
    • Selmin Nurcan, University of Paris 1, France
    • Yiwei Cao, RWTH Aachen University, Germany

    ORGANISERS

    • Riina Vuorikari, Katholieke Universiteit Leuven (K.U.Leuven) & European Schoolnet (EUN), Belgium
    • Barbara Kieslinger, Centre for Social Innovation (ZSI), Austria
    • Ralf Klamma, RWTH Aachen University, Germany
    • Prof. Erik Duval, Katholieke Universiteit Leuven (K.U.Leuven), Belgium & ARIADNE Foundation

    Wednesday, April 02, 2008

    My PhD dissertation, a new take on defining it

    How Social Information Retrieval (SIR) can be used to enhance the discovery of large-scale collections of multilingual digital learning resources

    The PhD dissertation deals with the discovery of digital learning resources and flexible access to large-scale collections of multilingual digital educational content. The thesis attempts to prove that we can use information deduced from social bookmarks and tags to better select suitable learning resources to users, who come from a variety of countries, speak different languages and whose educational context vary.

    The first step towards proving this thesis statement is to better understand whether there are digital learning resources that afford a good usage also in a context other than the one they were originally intended for. We call this type of educational content “travel well” resources because they cross borders easily; those borders can be national, linguistic, educational or socio-cultural.

    Upon better understanding of how users agree on “travel well” resources, we can explore the ways to identify them. Two different sources of information can be used for this purpose: looking at the properties of these resources (e.g. Learning Object Metadata), as well as attentional metadata collected from users interactions with the resources on the portal (Najjar, 2006). Our interest is in attentional metadata that we can gather from users' social bookmarks, from their personal collections of educational resources that they create, and from tags that they add to these resources (Vuorikari and Van Assche, 2007, Vuorikari et Poldoja, submitted).

    One major contribution of this thesis is the better understanding of how users (e.g. teachers) tag educational resources in a multilingual environment and whether a multilingual context has any implication on the tagging behaviour (e.g. in what languages do users tag) (Vuorikari, et al., submitted). Secondly, we are interested in the value that a multilingual tagging system provides; on the one hand, we want to know what kind of information multilingual tags can yield about the resources and their possible use in different contexts. On the other hand, we are interested in their value for resource discovery and as a navigational tool to allow cross-language and country exploration of new resources in multiple languages.

    Better understanding of tagging behaviour and creation of personal collections of learning resources will help us to create metrics that can be used to calculate “travel well” value of resource. Our hypothesis is that we can define a “travel well” resource when we use information deduced from social bookmarks, users’ personal collections of educational resources, and from tags that they have added to these resources. We will be watching the following variables:
    • The resource is from a different country than the user is
    • The resource is in a different language than user’s mother tongue,
    • The resource has tags in different language(s) than that of the item language
    The metrics used to calculate the “travel well” value of digital learning resources would be used to create a TravelRank algorithm that allows identifying learning resources that “travel well”, and which can be used to compliment the LearnRank algorithm (Duval, 2006). Identifying these resources from large collections of digital learning content from different countries and in different languages has a potential to allow a more flexible access to large-scale collections of resources. The final part of the thesis is to validate this claim and to evaluate its usefulness for a large audience of users from different countries.
    References:

    Najjar J., Wolpers M., and Duval E. Towards Effective Usage-Based Learning Applications: Track and Learn from User Experience(s). IEEE International Conference on Advanced Learning Technologies, (2006) (ICALT '06).

    Duval E. LearnRank: Towards a real quality measure for Learning. In U. Ehlers & J.M. Pawlowski (eds.), European Handbook for Quality and Standardization in E-Learning. Springer (2006), 379-384.

    other non-published, submitted papers at my site:
    http://www.cs.kuleuven.be/~riina/

    Monday, November 05, 2007

    Open social and education

    I wonder who is going to come up with the first OpenSocial app or widget for educational use? We certainly are talking about it, for example for our eTwinning platform. It could be cool to be able to use information about teachers collaborative networks to allow, say, better retrieval of learning resources relevant for the project, purpose or task that teachers are undertaking; link with some other sources that teachers are working on through cool widgets, etc.

    I never thought that Facebook, which has lately become really popular among my friends (not early adapters), would be the seul app that would "take it all". I was glad to read this:
    "The market has already decided that there's going to be a long tail of social networks, and that people are going to belong to more than one. As soon as you belong to more than one, this kind of interoperability is critical," Dash says. "Open standards win every time." wired

    Hurray for open standards!

    Saturday, September 29, 2007

    Notes on Smart Indicators on Learning Interactions

    Smart Indicators on Learning Interactions by Clahn et al. (2007) discusses how indicators can be used to help learners, or groups of learners, to organise, orientate and navigate through learning environments by providing contextual information that is relevant for performing learning tasks. Indicators are part of the interaction between a learner and a system (social or technical).

    Indicator system is defined as a system that informs a user on a status, on past activities or on events that have occurred in a context; and helps the user to orientate, orgaise or navigate in that context without recommending specific actions.

    So, it is not:
    • a feedback system (analyse user interactions to inform learners on thier performance on a task and to guide the learners though it) or
    • a recommender system (analyses interactions in order to recommend suitable follow-up activities),
    • instead it provides information about past actions or the current state of the learning process.
    • Moreover, smart indicator systems adapt their approach of information aggregation and indication according to a learner's situation and context.

    The paper draws heavily on the notion of social navigation, interaction history and footprints, and offers a good review of this literature (ToRead).

    The paper offers an architecture of smart indicators, where different layers are defined to support user modeling (first two) and helping the system to adapt to better decision making process (last two). Four layers:
    • sensor layer
    • semantic layer
    • control layer, where a strategy defines the conditions according to learner's context
    • indicator layer, presents aggregated information to the learner.

    This approach of smart indicators adapts the strategies on the control layer (as opposed to semantic layer) to meet the changing needs of a learner.

    SENSOR AND SEMANTIC LAYER

    The paper further presents the information aggregates of sensor and semantic layers. The idea is to classify and organise the user's engagement (interaction foot prints) with the system, e.g. contributions, tagging activities. In the sensor layer, there is a division between "learner interaction" and "contextual sensors", e.g. location tracker, tagging activities (in my case this is considered direct) and contributions of peer-learners.

    I am doing the same with my research data, and I call it the "user engagement" following the Yahoo!'s idea on STAR-metadata (kind of attentional and explicit metadata about users actions).

    I tried to apply the classes of Chlan's prototype to my research data (learning repositories) that I collect using our CAM framework. Our focus being somewhat different, it did not really work out that well. The attempt below, though:

    Direct: accessing resources through browsing, tag cloud, search result list, other user's favourites (implicit interest)
    • user views metadata
    • user views tags
    • user views resource ("entry selection sensor")
    • (timestamp on everything)

    Direct: higher level interaction with a resource (explicit interest)
    • user adds a resource to favourites and tags it ("entry contribution sensor", "tag selection sensor", "tagging sensor" or"tag tracing sensor", hard to say in my case)
    • user rates the resource ("entry contribution sensor")
    • user comments on the resource ("entry contribution sensor")
    • shares resource with network ("entry contribution sensor")
    • (timestamp on everything)

    Contextual sensors could be (here I'm blending them with user information):
    • context of a project within which the user access resources
    • the information about the country and school from where the user is from

    SEMANTIC LAYER

    The semantic layer users the information from Sensor layer and transforms it into meaningful information by using an "activity aggregator". This calculates the activity for a given period of time for an individual learner or the whole community according to different ratings that each activity has (beginners have different way of counting activity from power-users).

    CONTROL LAYER

    In this prototype the control layer defines how the indicators adapt to the learner behaviour. There are two elemental strategies:
    • motivate learners to participate to the community activity
    • raise awareness on the personal interest profile and stimulate reflection on the learning process
    Moreover, a third level control strategy uses the activity aggregator as well as the interest aggregator.

    INDICATOR LAYER

    This layer embeds the indicators into the user interface of the community system. The prototype is being tested by a group of PhD students now.

    Glahn, Christian, Specht, Marcus, Koper, Rob (2007) Smart Indicators on Learning Interactions
    http://hdl.handle.net/1820/941

    Thursday, September 27, 2007

    Some thoughts after SIRTEL07

    Last week the SIRTEL workshop took place. The papers are found here and the slides, well, most of them, at the EC-TEL07 conference wiki. I have pretty good feeling about the workshop, it was one day long, we had about 20 people participating, some of whom chose to stay with us for the whole time, and some who were hopping between workshops. For me that is totally fine, we all are responsible for our own learning! Especially in conferences where many parallel sessions are running, I would encourage people to try to get best out of them.

    For those who could not make it at all, you can soon find recording on the SIRTEL site.

    The workshop had four main sessions:
    • We started with a keynote address from people who work with music recommenders. MyStrands people talked about applying social recommender systems to technology enhanced learning. It was an interesting talk that challenged all of us to think what are recommenders for learning purposes in the first place (goal) and what kind of data do we want to use to do that.

      As any good keynote, this one gave more ideas to think than answers. It nicely set the base for the further discussions during the workshop that focused on the need to define the field of Social Information Retrieval for Technology Enhanced Learning, and to establish a baseline so that we know what are we really set to do.

    • The second session was about Tagging and Visualisation. We had my presentation about the user behaviour on tagging in multiple languages; then there was a presentation from COSL that talked about "Activities of Daily Living on the Web", Brandon also showed a few demos of the widgets that can be used to rate or recommend related content. That was followed by a talk on reward structures to encourage teachers to share open educational material. Finally, we listened about Visualisation of social bookmarks, a work that leads into visualising bookmarks in an educational repository.

    • The 3rd session was on Recommender Systems. Here we first heard about some R&D work that OU NL is carrying out using the idea of learning paths to better support learning activities of students. Then, there was a study about using affiliation networks as a mechanism for collaborative filtering (understood largely). This was followed by a study on simulating recommendations based on multi-attribute ratings on learning resources by teachers. Finally, we had a system demo of Daffodil that supports collaborative information seeking.

    • The final session was what we called "Enablers and Challenges". It was a discussion session, and as we advanced, it was clear that people had a lot to say. It might even have been better to allow more time for this, but hey, you live and you learn.
    I try to sum-up, but basically it illustrates the main topics that we talked about. If you look at the left side, there are the fundamental questions:
    • How to define and chart out the area of Social Information Retrieval (SIR) for learning?
    • Is this application domain different from other SIR, on micro and macro level?
    • What do we recommend?
    • In what context?
    • and based on what?
    On the right hand, there are the issues related to implementation and evaluation of it. These are:
    • What are the best SIR methods for TEL?
    • And what is the data that we should use? The "data issue" was something that was heavily emphasised by the MyStrands folks, who obviously speak of experience.
    • The questions rouse also: when do we start implementing these for real or are we just over-engineering and never ready to launch?
    • Evaluation and empirical data for real evidences was on the focus a lot.











    More will follow. This is quick and dirty now, hopefully I will get more input from people participating in order to get more depth on our summary.

    Sunday, September 16, 2007

    SIRTEL'07: la raison d'etre

    "We use people to find content. We use content to find people."*

    On Sept 18 our SIRTEL workshop takes place. It's gonna be "Serious Fun"! Let me just outline why:

    SIRTEL'07: Raison d'etre

    Recommender systems, as well as social navigation, have been around since the popularisation of WWW, that's some 15-20 years now. The idea is to help people choose the right stuff from a potentially overwhelming set of choices. To facilitate that users could be helped with information from other users, the choices made before (by themselves or similar users), the ratings or reviews other people had done, etc. (Rescnik et al., 1997)

    The field of learning technologies has seen recommenders of some sort being discussed and prototyped since the late nineteens. In the review of the field in Manouselis et al (2008) we identified about 10 recommenders, and even more conceptual papers of them, but very little has matierialised so far.

    Since the last few years recommenders have made a second arrival into the discussion topics of technology, or network, enhanced learning. Undoubtedly, this has been influenced by the arrival "Web 2.0" with all its ideas:

    - Collaborative tagging, for example, has changed lots of ideas of how metadata should be produced and how static a metadata record should be: it's not anymore one metadata record produced by a librarian, but lots of annotational and attentional metadata by lots of users.

    - Other annotations by users that express their subjective judgements have seen a huge growth too, we don't only talk about ratings or reviews in their traditional sense, but also tumbs-up or down, giving pokes to people or objects, etc.

    - Social bookmarking, which allows users to create easy references to their own collections of digital resources (photos, books, links, music,..), has given a new dimension to the concept of social navigations. The link between resource-user(-tag) allows users to navigate other people's collections and thus find novel resources. Also, the same resource-user-tag link gives researchers an itch to use this information to group similar users for recommendation purposes, as well as to study the emerging networks.

    - Expressing social ties between people has also brought new possibilities along. We are not only seeing networks of friends, but there are new possibilities where people can express different networks, ones for professional use, others for personal, recreational, etc purposes. Also, portability of these networks has become an issue discussed for better designs (social-network-portability group, PeopleWeb ,..).

    - Something else is also happening behind the scenes. Clicksteam and user behaviour on the Web is not anymore a property of the commercial portal on which users are, but users are starting to take seriously how their "attention" is being used, who owns it, etc. Attentional metadata is a huge source of information that educationalists are also starting to take more seriously and thinking how it could be used for better serving learners and teachers (Contextual Attention Matadata, Attention Profiling Mark-up Language, Attention Trust,..). Attentional metadata can also become crucial when it comes to better understanding the intentions of a user, why are they, for example, looking for some information and for what task at hand!

    - Finally, content for educational use, or rather its production, is also seeing a change. Users generate more and more of the content on the Web in general, a trend which is also seen in the e-learning. Of course, traditionally teachers have always produced lots of their own material, but now its re-use also has been facilitated (e.g. repositories/referatories). Also, the collaboration aspect is facilitated by the Web, it has become easier for people to work together on things (e.g. wikis, collaborative platforms,..). Additionally, learners produce plenty of material which also should be seen and used as educational content.

    To sum-up: two main topics evolve around social context and social content. Social context is how we express the who, where and with whom, and social content are the objects or digital artefacts that are in the center of the communication, exchange and networks.

    All the above has hopefully also changed how we will see the future of social information retrieval for technology enhanced learning. This workshop will all be about that! Serious Fun!

    -------

    N. Manouselis, R. Vuorikari, F. Van Assche, “Collaborative Filtering of Learning Objects for Online Communities: An Experimental Investigation”, accepted for publication in Computers in Human Behavior, Special Issue on ‘Advances of Knowledge Management and Semantic Web for Social Networks’, 2008.

    P.Morville, 2004

    Resnick P. & Varian H.R., “Recommender Systems”, Communications of the ACM, 40(3),1997

    Monday, July 23, 2007

    Multilinguality of tags and etiquetas

    I'm currently looking into the multilingual use of tags in different applications that allow users to taguér les favoris (=signet sociaux), music, photos, etc. and that display these etiquetas in a Nube de etiquetas = Tag-Wolke = nuage de tags. E.g. I took at quick tour on 10 collaborative tagging services to see how European multi-linguality is reflected on these services.

    Lots of collaborative tagging efforts on the Web take place in English. English being the lingua-franca of the Internet, it is probably not that surprising. However, we here in Europe live in a multi-cultural and lingual environment, so traces of that should be found from here and there.

    Fair enough, I was able to find French and German services (blogmarks.net, MisterWong.de, Oneview.de,..) that harbor communities of users who tag in their native language - as well as in English, too. Some people do not seem to have any difficulties in adding tags in a few different languages, e.g. talking about blogging, they add tags like "blogue", "blog" for for e-commerce "shopping" and "achat".

    I also looked at some "global" services such as Yahoo!, del.icio.us, Amazon and LibraryThing.com to see how do they deal, if at all, with the issue of tags being in different languages, and in the case of Amazon and Yahoo!, who offer localised sites, how are tags managed and provided in different languages. Below I give a few examples.

    Yahoo!'s MyWeb

    Yahoo!'s MyWeb, which is their social bookmarking and tagging service. The service does not exist in all their localised sub-sites, but with a quick look I found it at least in Yahoo.fr, Yahoo.es, Yahoo.uk, Yahoo.de. On these respective MyWeb sites one can find popular bookmarks and tag clouds in the language of the sub-site.

    In all the above mentioned sub-sites, regardless in which language I was looking at, 1116436 tags were registered. This first lead me to think that they actually have that many tags in German, Spanish and in French. Then I realised that it was a total of the tags in any languages, and they had much less tags in other languages than in English.
    • 1079 in German
    • 1052 in Spanish
    • 873 in French
    What is remarkable in Yahoo!, though, is that they seem to have a way to recognise the language of the tag somehow, as they are able to display German tags in their .de service and French tags in their .fr service. This is transparent to the user, so I have only little idea (although a few guesses) how they do that. This focus on languages is rather unique on Yahoo!'s service, it is not the case with any of the other services that I looked. I'll be interested in knowing what they will come up with this!

    LibraryThing.com

    LibraryThing.com users, as well as developers, love tags! Users actually use them, after all, most of them are book freaks who probably hang out a lot in libraries, so tagging comes easy to these folks. But also, the developers have done a few really clever things with tags and objects to tag, like the concept of "work" (a work brings together all different copies of a book, regardless of edition, title variation, or language) and combining tags (see "concepts"). The difference here to delicious "bundled" tags is that it is done once for all users.

    By combining tags, some multilinguality is also taken advantage of.

    Like in this example, 19th century also includes 19. Jahrhundert and 19eme siecle.

    del.icio.us

    Take Delicious as an example of a different approach. It is one of the most used social bookmaking services, but there is little indication to be found of hundreds of languages that exist in the world. Well, I dug a bit harder and found two different types of indications of multilinguality existing.

    Either people added tags in a few languages (achat, shopping), or they added a tag like "lang:fi" to indicate the that source site was in Finnish. Also lang:fr, lang:es, lang:pt, .. were there with various amounts of bookmarks. If you know if this was a user initiated activity or whether delicious encourages it, let me know. I could not find anything on it.

    Amazon

    I also looked at Amazon.com and Amazon.fr, funnily enough, the Frencheis do not even get to tag! (not that tags have been taken up in Amazon in the first place...) Moreover, they only get the reviews that are done by the users of the fr-site, not by the users of the .com-site. That sounds a bit silly, but maybe preserves a way to keep the cultural taste intact.

    Different strokes for different folks

    It seems to me that not many sites have paid much attention to the issue of multilinguality, however, it might also well be that it is not an important issue to them. Take for example these two German bookmarking sites, Mister Wong and Oneview.de, and compare the Top-50 tag lists.

    MrWong.de (image on the left) has mostly English terms in their Top-50 and they are very Web2.0 and developer oriented.

    Oneview.de has many German terms (image on the right) and they seem to cover larger area than only Web2.0, there are terms about holidays, music, etc.

    So, each user group seem to share the terms that are important for them. Most likely this is also the case in Delicious, by using English tags I, as a Finn, can easily share interesting bookmarks on, say, folksonomies with everyone else in the world who is interested in it.

    This to say, I also think that multilinguality has a place in tagging and we in EUN are very interested in the issue. Currently, we are running a pilot where teachers can add tags in their own language(s) to learning resources. We try our best to be able to recognise the language of the tags, as we think cool things can be done with it. But more about that another time.

    Sites that I looked at

    French:
    • social bookmarks =signet sociaux
    • tag, taguer

    Blogmarks, http://blogmarks.net/
    - most popular tags are in English
    - 577366 bookmarks
    - people sometimes add tags both in French and English (e.g. blogue, blog; tag cloud nuage de tags)
    - Tags in French also, but they don't seem to have that many users, tried some about 10 or so


    Bookmakrs, http://bookmarks.fr
    - call bookmarks "favoris" and tags "tags"
    - most popular tags in different languages, in 50 most popular tags: 14 in French and 1 in Russian
    - maybe only about 2500 bookmarks (25x107 pages)
    - some people had indicated the language of the source page in En

    In German:
    • Tags as "Schlagworte", "tags"
    • "Tag-Wolke"
    • bookmarks " Lesezeichen" (Soziale Lesezeichen)

    Netselektor, http://netselektor.de/

    Mister Wong, http://www.mister-wong.de/?tag_type=list
    - 2.142.735 Bookmarks
    - very techy, lots of En tags

    Oneview http://www.oneview.de/home/index.jsf
    - Trendwolke http://www.oneview.de/home/discover.jsf
    - more tags in German in general areas


    Yahoo! MyWeb

    Yahoo! Fr
    - top tags http://fr.myweb2.search.yahoo.com/myweb?dg=6&sort=pop
    - 1 116 436 tags in the "nuage de tags", although 873 in French.
    - in top 50 tags mostly in French, however, many terms are easily understandable like web2.0, internet, google, yahoo..

    Yahoo! De
    - top tags http://de.myweb2.search.yahoo.com/myweb?ei=UTF-8&dg=6&dmode=vtags&sortby=count
    - 1.116.436 tags in "Tag-Wolke", although only 1079 in German
    - in top 50 most entirely in German, but many terms like computer, software

    Yahoo! es
    - toptags http://es.myweb2.search.yahoo.com/myweb?dg=6&sort=pop
    - call tags "Etiquetas" and tagcloud
    - 1.116.436 tags in "Nube de etiquetas", 1052 in Spanish
    - the top 50 mostly in Spanish, only a few terms like web2.0, blog, yahoo



    Flickr
    http://www.flickr.com/photos/tags/

    - interface offered in different languages, however,
    - most popular tags are the same in all different interface languages, thus could think that there is no separation of languages.


    Delicious

    - All top 50 tags in English
    - no other interface languages
    - some tags do exist in different languages e.g. achat, shopping; ..
    - some tags indicate the language of the source; lang:fi,..

    Google

    if I understand right, google does not even display the tags from their users? Please, correct me if you know anything about using their bookmarks.

    Last.fm

    - in LastFM all the tags are the same, although they offer different interface languages
    - tags in many language exist, especially if you look for bands from different countries e.g. suomipoppia