Showing posts with label PhD. Show all posts
Showing posts with label PhD. Show all posts

Monday, November 16, 2009

My degree of Doctor at Open Universiteit Nederland

"..you hereby receive all rights associated with the degree of Doctor either by law or custom"

My Doctorate Board and me with the degree and a huge smile - what a day!

From left Prof. dr. G. Conole, Prof. dr. A. Littlejohn, Prof. dr. B. Berendt , me, Prof. dr. P.B. Sloep, Prof. dr. E.J.R. Koper (supervisor), Prof. dr. J. van Marle (chairman) and Beadle E. Vinken. And no, even if I'm a doctor now, I don't get to wear a funny hat!

I thank everyone involved in my research and all who sent me kind wishes and congratulations!



Slides and PhD download

Thursday, October 22, 2009

Most scholarly opponent - PhD Ceremony in OUNL Nov 13 2009

It's pretty exiting to prepare for the PhD ceremony. There are all kinds of little things to think about. I've never gotten married, but it sure sounds like the same thing. In the Netherlands, for example, we need to have paranymphs at the PhD ceremony.

Yes, I know, it sounds like I need to have two dwarfs by my side there, but apparently they are like a bride's maid or a best man. With a twist that in case there was a heated fight between me and the opponents during the defence, the paranymphs would defend me with swords. Or, you know, something similar. So far, all what I've seen is that they hand over a class of water, but - we've ain't seen nothing yet..

Another thing is the manner of speech, I am, for example, to address the members in my committee by saying "most scholarly opponent", if they are a professor, otherwise just "scholarly opponent" is fine. And they call me "esteemed candidate". Pretty theatrical :)

Finally, I needed to prepare 10 statements in advance, they are called stellingen. 4 of them are about my thesis and the rest are more general (some can even be a bit funny, see #10). But in general, they all should be something that I can defend.

The purpose of this is that when the opponent has not had time to read my thesis and come up with a question, they can read one from the list. This way they can have a well formulated question to ask. My advantage - I can prepare for it in advance. Kind of a funny game! It seems to me that the more bold the statement is, the better chances there are that one of the opponents reads it out loud. Let's see...

1. A learning resource portal in a multilingual context can be made more robust and flexible by interrelating conventional metadata and social tags. (this thesis)

2. Socal tags, represented as a triple (user,item,tag), open more sophisticated avenues for resource discovery across contexts, especially when it applies to cross-language and cross-country discoveries. (this thesis)

3. The triple (user,item,tag) can be used as a parameter to measure links between cross-language content that reside on heterogeneous repositories. It can be created a posteriori to content creation and link-setting, and it can be used to support and enhance a new type of link-following behaviour by end-users. (this thesis)

4. The discovery strategies based on Social Information Retrieval (SIR) methods allow users to spend less effort in finding relevant resources on a multilingual portal. (this thesis)

5. The notion of learning resources as content is too limiting.

6. Even if the current trend in information seeking behaviour is the Web, interpersonal ties still drive and support information seeking. Social search should be considered beyond individual’s Web-behaviour.

7. While studying the impact of new information and communication technologies on education and attainment, studying their out-of-school use should be considered as important as their in-school use.

8. In Technology-Enhanced Learning, like in other lo-fi high tech, the next best thing is “good ‘nuf” for most users.

9. Languages both unite and divide people.

10. Anyone involved in decision-making for educational purposes should read science fiction.

Tuesday, September 29, 2009

Invitation to the public defence of my PhD

Hurrah, finally the day when I can announce the public defence of my PhD: Nov 13 2009 at 13.30. You are welcome to join the public defence of my PhD at the OUNL premises in Heerlen, NL!

Today I also submitted my manuscript to print. Here you can see the cover of the publication, which is also downloadable here.

Run of the day:
13.30 Public defence in OUNL
15.00 Reception (all welcome, RSVP)
19.30 Dinner in Brussels (RSVP)

Tuesday, September 01, 2009

Best Paper Award in ICWL'09

The paper that I presented in ICWL'09, Are tags from Mars and descriptors from Venus? A study on the ecology of educational resource metadata, was awarded the Best Paper Award. That was pretty dam cool! The picture below shows how psyched I was to pick up the award at the gala dinner. You'll find the paper from here and a news release from here.

The picture is by Kevin Chen from RWTH, Aachen.

Ok, this is an inside joke: I've named the pic "la vengeance se mange très-bien froide", for those who know the story, you guessed that the timing of this award could not have been better! Thanks for the co-authors and the jury :)

Friday, July 10, 2009

Tags and self-organisation: a metadata ecology for learning resources in a multilingual context

I think I finally came up with a title for my PhD. You know, the type of title that says it all. It's a bit long, but "correct and descriptive", like Matt said. So here it goes: Tags and self-organisation: a metadata ecology for learning resources in a multilingual context.

Here is a wordle, it was extracted from a paper that summaries the research. Looks pretty accurate :)

Tuesday, June 30, 2009

Study on contexts in tracking usage and attention metadata in multilingual Technology Enhanced Learning

Just submitted the final version of the paper to a workshop on Exploitation of Usage and Attention Metadata (EUAM 09). Here is a one-pager about it and the link to the paper.

Study on contexts in tracking usage and attention metadata in multilingual Technology Enhanced Learning

“Context” is widely accepted to be important for correctly interpreting user input and for improving predictive and possibly also diagnostic models. But what is context, and how can it be measured? By measuring we mean to operationalise the construct and data gathering to provide values for the desired variables.

In this study, we consider the intersection of the areas of digital learning resource repositories, digital libraries and social tagging systems where users from a variety of countries use technology enhanced learning (TEL) offerings in a variety of languages. We consider usage and attention metadata as an example of the wider notion of context adapting the definition of context as “any information that can be used to characterise the situation of entities” [Dey01]. We give an overview of dimensions of context that are relevant in TEL, specifically arguing that context comprises the usage situation and environment as well as persistent and transient properties of the user. Therefore, distinguishing between the macro-context and the micro-context of TEL is useful.

TEL and the analysis of the data it generates take place in different types of educational settings which we call the macro-context of TEL. We use the term micro-context to denote the context that is relevant for interpreting a specific user input and for designing adequate system responses and other output. The micro-context is subdivided into user models, material/environment models, interaction models, and background knowledge, showing that usage and attention metadata are of different types and play different roles for learning about context.

We then concentrate on teachers using learning-resource repositories as an important use-case example of TEL and focus on language and country as context variables. We describe different ways in which these variables are operationalised, and we outline ways in which TEL use such context information to improve the use and reuse of repositories by supporting users in a multilingual and multicultural context. A key theme of our article is the central role that social tagging can play in this process: on the one hand, tags describe usage, attention, and other aspects of context, on the other, they can help to exploit context data towards making repositories more useful, and thus enhance the reuse.

Riina Vuorikari 1,2, Bettina Berendt3
1 European Schoolnet, Brussels, Belgium,
2 OUNL, Heerlen, Netherlands,
3 KU Leuven, Belgium

Thursday, June 25, 2009

My tag paper nominated for best paper award 2009

I'm pretty exited that one of my papers for ICWL 09 was among the 5 best paper nominees. For a some time now I've been wondering what does it take to write a paper that arises above the general mass of papers. Well, now I have a bit better idea :)

What does it take? Reading tons of research papers, write a few (un)successful ones to practice, a good inspiring topic, some research work with ppl who are truly interested in what they are doing, and voila!

I also like how Celstec, OUNL (where I study), picked it up for their news feed. I think that over all, they have a pretty neat way to recognise what's going on and make others aware of it too. A modest person as I am, I would never make any fuss about it.... right.. ;)

Monday, June 22, 2009

Wiley calls it “dirty secret” of OER

Just picked up a fresh PhD study by S. M. Duncan from USU, a student of D.Wiley's. The study is called Patterns of Learning Object Reuse in the Connexions Repository. The punch line is that there is very little reuse of LOs among the repository studied.

What new? Similar findings have been discovered here in Europe (end elsewhere) for a while now. Ochoa (2008), for example, found in his PhD dissertation that reuse in general remains low, about 20%, across all sizes of collections. This was interesting not only for how low the reuse is (20%, common!), but also because since forever folks have been saying that resources with smaller granularity are more reusable, as they lack context, etc (insert here the infamous graph of "modular content hierarchy", the most used LO). Well, according to Ochoa (2008), this was not the case.

I also looked at the reuse on 2 different platforms: LeMill and Calibrate from European Schoolnet. My twist was to study the cross-boundary use and reuse, i.e. teachers reusing learning resources that are in a language other than their mother tongue and originate from different countries than they do. I used the same reuse definition as Ochoa (2008), which basically is the same as in Duncan's study.

The finding was that the general reuse was around 20%, but NOT across all collections. For example, in LeMill, "Multimedia material" was used more often, but in Calibrate, the smaller granularity was seldom added to Collections. The cross-boundary reuse was notably less (37% to 55% of it). Moreover, in some of the collections only around 10% of resources were ever added to a collections, which makes you really think hard about the efficiency of this all..

Anyway, the good news in Duncan's study is this:

There was a common author in 3,722 module uses, while there were only 1,013 module uses where there was no common author. This means that modules were included in collections 3.67 times more often when there was at least one person in common with both the module and the collection.p.32


So if people know each other, they are more likely to reuse material from each other! This shows that social is important when we are talking about the use and reuse of learning resources! This is similar to what I am saying in my PhD thesis, which hopefully will come out one day soon. My twist of course is that tags can make those social connections between people, and by taking advantage of these underlying social connections, we can make the learning resource discovery much better - and hopefully also more useful for teachers.

Vuorikari, R., Koper, R. Evidence of cross-boundary use and reuse of digital educational resources. Link to a revised version of the paper, not reviewed yet!

Tuesday, May 05, 2009

Challenges and lessons learned from Social tagging in MELT

Social tagging in MELT, how do we want to take the social tagging work forward

Social tagging of educational resources potentially offers new ways for:
  • Individuals to
    1.1) better manage their digital learning resources that reside in different repositories and platforms, and

    1.2) discover and access new resources from different contexts (e.g. different language, educational system) through tags and other users.

  • LOR managers to
    2.1) get third party metadata on learning resources (either the ones that already reside on their repository, or the possible new ones to be added to collections,

    2.2) create affinities (e.g. link structure) between separate pieces of resources (either on their own repository, or the ones that reside on other repositories on the federation or on the Web) that were not cross-referenced before.

  • In the MELT project so far, we have only been able to see the peak of these potentials emerging. We list issues that we see important for future work in the field, for the clarity, we only list one of the main issues for each topic:

    • 1.1 To fully support users in their knowledge management task on digital learning resources, the bookmarks (including title, url and tags) should be exportable in standard Webfeed formats. This would allow users to access and manage their MELT resources as part of their other resources collections, whereas now users need to be logged on to the MELT portal to do this.

    • 1.2 Pivotal browsing of social bookmarks takes advantage of the affinities between the user, resource and tags. In the MELT context, more metadata could also be added to support pivotal browsing, such as the country of the user, interest topics; resource metadata such as multilingual indexing keywords. This would allow novel ways to access resources that other users have already discovered within the federation, and thus build on users’ social interactions and co-construction of knowledge.

    • 2.1 Tags by end-users on the MELT portal have been shown to be of good quality as additional metadata descriptors of resources. We have enumerated possibilities of metadata ecology that the use of multilingual Thesaurus can offer to a federation such as LRE. Apart from working on ways to automatically generate LOM from tags, we urge on using the hierarchical structure and multilingual features to leverage user-generated tags.

    • 2.2 Why not do Google for learning resources? Using PageRank-like algorithms on a learning resource repository or federation has been impossible for a number of reasons, the most important is the lack of a link-structure that cross-references resources. Tags, creating underlying connections between seemingly random pieces of content in different languages, on repositories in different countries and other platforms on the Web, rely on humans’ subjective idea of its importance for a given information seeking task. Using this new, emerging link-structure with tags as “anchor texts” offers totally new ways to “organise the world's learning resources and make them universally accessible and useful”. A new tag line could be “From teachers to teachers”.

    Monday, May 04, 2009

    Link structure and anchor text

    I read that Brin & Page (1998) paper again. A few guidelines to keep in mind:
    ..our notion of "relevant" to only include the very best documents since there may be tens of thousands of slightly relevant documents. This very high precision is important even at the expense of recall (the total number of relevant documents the system is able to return).


    Two features to produce high quality precision:
    • Link structure is used to create objective measure of its citation importance that corresponds well with people’s subjective idea of importance. Well, it's that simple..

    • Anchor text:
      ..anchors often provide more accurate descriptions of web pages than the pages themselves. Second, anchors may exist for documents which cannot be indexed by a text-based search engine, such as images, programs,..
    The point about the anchor text is so interesting, I wonder how well does it apply to tags? I bet really well..

    I also found this interesting: "it has location information for all hits and so it makes extensive use of proximity in search"

    Differences Between the Web and Well Controlled Collections
    • extreme variation internal to the documents: documents differ internally in their language (both human and programming), vocabulary (email addresses, links, zip codes, phone numbers, product numbers), type or format (text, HTML, PDF, images, sounds), and may even be machine generated (log files or out putfrom a database).
    • external meta information as information that can be inferred about a document, but is not contained within it. Examples of external meta information include things like reputation of the source, update frequency, quality, popularity or usage, and citations. Not only are the possible sources of external meta information varied, but the things that are being measuredvary many orders of magnitude as well.


    http://www.scribd.com/doc/3208417/The-Anatomy-of-a-LargeScale-Hypertextual-Web-Search-Engine

    Saturday, May 02, 2009

    Cross-language use of the Web; users behaviours and attitudes

    Berendt & Kralisch (2009) A user-centric approach to identifying best deployment strategies for language tools: the impact of content and access language on Web user behaviour and attitudes

    The results indicate that non-English languages are under-represented on the Web and that this is partly due to content-creation, link-setting and link-following behavoiur. User satisfaction is influenced both by the cognitive effort of searching and the availability of alternative information in that language.

    Cost=time+cognitive effort

    Not only capacities to access the site but also opportunities to access it, thus language is only one factor.
    • Language can be expected to not only influence the total amount of information available to Web users, but also how information sources (i.e. Websites) are linked among each other and therefore how easy/likely it is to find and access a certain Web site.
    • Bharat et al. and Halavis are first indicators of the potential impact of language: Website in different languages are less connected than sites in the same language (note: studied data aggregated on the national level and therefore only limited insight into the role of language.
    "Web sites are, in most cases more likely to link to another site hosted in the same country than to cross national borders. When they do cross national borders, they are more likely to lead to pages hosted in the United States than to pages anywhere else in the world." (Halavais, A, 2000, p. 7)

    Behavioural aspects of information seeking:
    1. Users' information seeking behaviour,
    2. information and information flow on the Web,
    Attitudinal aspects of information seeking:
    1. "usefulness", i.e. the language related value of information decreases as more information is offered in that language on the Web. "..value perceptions are also determined by topic; thus a large amount of content on a topic in a native language may also reduce the value of content on that topic in other languages.

    2. "ease of use", i.e. the cost of language processing during information seeking can be expected to affect attitudes in Web search.
    Results on behavioural aspects

    1. Non-English languages are under-represented on the Web in terms of the amount of content supplied.
    2. Search engines do not register all pages linking to the site, and many links known to the search engine were not used. This indicates that non-English language s are under-represented on the Web in terms of the links that content creators set to content in those languages.
    3. Users have a clear preference to navigate in their native language when it is available via a link, but if that is not available they accept the necessity to navigate in English.
    4. This all means: behavioural tendencies both of content providers and of content users lead to mutually reinforcing under-representation of non-English languages. Compared to the respective market size or available options, there is less content in these languages, this content is linked to less and the links are followed less often.
    Results on Attitudinal stuff:
    1. A complex interplay of English language skills, the perceived saved effort of using native-language content, the perceived overall supply in that language on the Web, and satisfaction:

    2. People who are proficient in English often prefer to navigate in English (even if offered content is their own language) and are more scrutinised of the quality of Web content. Do not care much about whether sites make efforts to provide them with content in their own languages.

    3. People who are not so proficient in English do perceive the (real) scarcity of information in their native language and are highly appreciative of content in this language.
    This means that content and search-tool designers should not draw simplistic conclusions based on behaviour alone, because this is not a reliable indicator of attitudes and preferences. In the absence of links and/or content in their native languages, users will acquiesce to English-language content. However, their preference will persist.

    Berendt, B., & Kralisch, A. (2009). A user-centric approach to identifying best deployment strategies for language tools: the impact of content and access language on Web user behaviour and attitudes. Inf. Retr., 12(3), 380-399.



    HALAVAIS, A. (2000). National Borders on the World Wide Web.New Media Society, 2 (1), 7-28. doi: 10.1177/14614440022225689.



    Bharat, K., Chang, B., Henzinger, M. R., and Ruhl, M. 2001. Who Links to Whom: Mining Linkage between Web Sites. In Proceedings of the 2001 IEEE international Conference on Data Mining (November 29 - December 02, 2001). N. Cercone, T. Y. Lin, and X. Wu, Eds. ICDM. IEEE Computer Society, Washington, DC, 51-58.

    Friday, February 27, 2009

    Are tags from Mars and descriptors from Venus?

    A study on the ecology of educational resource metadata.

    I just finished a paper on the tag evaluations that we did in the MELT project. We had lots of fun with the name of the paper :) the main question being which one, tags or descriptors, should be from Venus...?

    Anyway, we were able to show that not all the tags are as far from the Thesaurus descriptors as Mars is from Venus. We had different perspectives for evaluations: end-users, expert indexers and repository owners. For me the most interesting thing that came up was that 11% of end-user generated tags are actually terms that we can find in our multilingual Thesaurus! I assume teachers are "better taggers" than average, usually there is lots of talk about the gap between end-users' language and the one deployed by experts.

    Abstract. pdf. In this study, over a period of six months, we gathered empirical data from more than 200 users on a learning resource portal with a social bookmarking and tagging feature. Our aim was to look at the tags from different stakeholders’ points of view; end-users, librarians/expert indexers and repository owners. We first look how users tag resources, and then conduct an evaluation with indexers to understand how they perceive the value of tags as descriptors. We then present a case study from a repository owner’s point of view. Lastly, we study users’ clickstream when searching resources. We find that, even though end-users and expert evaluators apply very different strategies when adding metadata, (end-users have a rather synthetic approach whereas expert indexers an analytical one) there is an overlap in the information in tags and the official descriptors, this overlap is even up to 51%, creating an ecology of metadata.

    Keywords: Learning resource metadata, tags, folksonomy, clickstream,
    thesaurus, evaluation.





    Wednesday, January 07, 2009

    Out Now: Special Issue on Social Information Retrieval for Technology Enhanced Learning

    I am glad to announce the Special Issue on Social Information Retrieval for Technology Enhanced Learning (SIRTEL) which just came out today in Journal of Digital Information (JoDI) Vol 10, No 2 (2009)!

    I co-editored it with Erik Duval and Nikos Manouselis. The following stuff's in it, enjoy!

    Special Issue on Social Information Retrieval for Technology Enhanced Learning HTML
    Erik Duval, Riina Vuorikari, Nikos Manouselis

    Articles

    Identifying the Goal, User model and Conditions of Recommender Systems for Formal and Informal Learning Abstract PDF
    Hendrik Drachsler, Hans G. K. Hummel, Rob Koper
    The Pedagogical Value of Papers: a Collaborative-Filtering based Paper Recommender Abstract PDF
    Tiffany Y Tang, Gordon McCalla
    Lost in social space: Information retrieval issues in Web 1.5 Abstract HTML
    Jon Dron, Terry Anderson
    Exploratory Analysis of the Main Characteristics of Tags and Tagging of Educational Resources in a Multi-lingual Context Abstract HTML
    Riina Vuorikari, Xavier Ochoa
    Visualising Social Bookmarks Abstract PDF
    Joris Klerkx, Erik Duval


    A Special thank to people who participated in the PC:
    • Alexander Felfernig, University of Klagenfurt, Germany
    • Brandon Muramatsu, Utah State University, USA
    • David Massar, European Schoolnet, Be
    • Hendrik Drachsler, Open University of the Netherlands, The Netherlands
    • Jon Dron, Athabasca University, Canada
    • Marc Spaniol, Max-Planck-Institute for Informatics, Germany
    • Martin Wolpers, Fraunhofer, Germany
    • Miguel-Angel Sicilia, University of Alcala, Spain
    • Nikos Manouselis, Greek Research & Technology Network, Greece
    • Rick D. Hangartner, MyStrands, USA
    • Salvador Sanchez, University of Alcala, Spain
    • Xavier Ochoa, Escuela Superior Politécnica del Litoral, Ecuador
    • Yiwei Cao, RWTH Aachen University, Germany

    Sunday, January 04, 2009

    New users on the portal and resource discovery

    I looked at 18 new users on the portal, and studied resources that they bookmarked. I wondered how many of these resources had previous annotations by users? By annotations I mean that previous users had added ratings on them and Favourited these resources. If this is the case, it's clearly shown on the portal.

    But are these annotations persuasive? Do they help users make their decision better or faster?



    I had two sets of data:
    • last 3.5 months (Aug to Nov 17 2008)
    • Nov 18 to Dec 18. This are my 18 users who had bookmarked 114 resources.

    Out of these results, it seems that 2/3 of the resources that these new users bookmarked had no previous annotations on them! I have to verify this finding, because currently I lack data from March to July to see what was bookmarked then.

    Table 1


    Anyhow, let's see what the current mini-study holds. Out of the third of resources that were discovered by this group, about 25% had previous annotations on them. They were mostly done by users before this group got initiated on the portal, however, some were also discovered thanks to the bookmarking by this group.

    Only about 8% of resources were discovered through a special list called "Travel well" resources. These resources have been added there by "EUNRecommender" which currently is hand operated, but mainly based on picks by other users from at least two different countries. I find this figure rather surprising, as this "Travel well" list is the first thing that the user sees when they come to the portal.

    Anyhow, I find it cool that 25% of bookmarks by this group were resources that had previous annotations. What we cannot say, though, is whether these users could have found these resources without these annotations displayed publicly on the search result list. However, I think that my measures like "pick-up rate" and "overlap" among Favourites will help me sort that out.

    Here is a visualisation of this. The huge node is "sos-rec" which means that these resources were previously annotated (rate, bookmark). What is called "EUN Recommender" are resources from the "Travel well list". Other than that, the new users are the nodes which are connected by edges to resources that they have bookmarked.

    Right from the bat, we can see that 5 users had taken their totally own trails and bookmarked resources that no one had bookmarked before - quite cool!

    Additionally, there are few users on the lower right hand corner who are only connected to the whole graph through one resource (user: 192682) and (user: 217391).

    Monday, December 22, 2008

    Share early: Paper on Evidence of cross-boundary use and reuse of digital educational resources

    I finally sent off my paper to a journal. Exiting. The first comment was to cut it shorter by about 2500 words, even before they started reviewing it. Outch, I think I managed to do it, I have a copy of it here:

    Vuorikari, R., Koper, R. (submitted). Evidence of cross-boundary use and reuse of digital educational resources. pdf
    ABSTRACT: In this study we conducted an investigation on the server-end log-files of teachers’ Collections of educational resources in a number of content platforms. Our goal was to find empirical evidence from the field that teachers use and reuse learning resources that are in a language other than their mother tongue and originate from different countries than they do. We call these cross-boundary learning resources. We compared the cross-boundary reuse of educational resources to the general reuse figure of 20%, and find that it was either equal to or less than the general reuse. We further studied the coverage, the overlap and the pick-up rate of these resources, and propose steps that could improve the probability of discovery, use and reuse of cross-boundary resources.

    I actually have a new academic homepage too, check it out http://elgg.ou.nl/rvu

    Friday, December 19, 2008

    How different is user behaviour on a portal from the ones who log-in to ones how do not?

    I've recently done quite a few studies on users of learning resources portals, I've looked for example how do they tag resources in a multilingual context or how much use and reuse is there across the borders. In all cases the studies have concentrated on the small amount of the (minority) users who actually log in and had created Collections of resources: in Calibrate that was about 30% and in LeMill about 10%.

    Now in MELT we've revised the logging scheme to collect the click-stream from users who don't log-in. We also have Google analytics, but I don't have those at hand right now. I looked at the data from last 3 months, from Aug 18 to Dec 18, and then only from the last month (Table 1).

    What do users do on the portal?

    The most popular activity on the portal is search, 64.29% of all actions on the portal are different types of searches. They result in "playing" the resources in 18.31% of all actions on the portal. 13.09% of all actions are contributing actions on the portal, this means adding a tag, bookmarking or rating it. The figure of contributing actions is actually a bit distorted, we count each tag, rating and bookmarking there. As each bookmark has average of 4.3 tags attached to it, it brings up the figure. Actually, the number of actions that contribute to "acting with an individual resource" is around 4.3% of all actions (i.e. add rate and bookmark). Other includes activities like view evaluation, view other users who have bookmarked the resources, etc.

    Table 1


    The downside here is not having the stats from Google analytics, so I cannot exclude our internal usage, which I know has been quite a lot, since we've been testing the portal internally. So the figures might be somewhat distorted...


    What about users who log-in and the ones who don't?

    About a month ago I invited some 260 teachers on the portal, so I was intrigued to see what had happened. 2 weeks ago I checked that 11% of these teachers had started their own account. But, it seems like much more have come about and cruised around the portal.

    Table 2 presents the data from the last 4.5 months (Aug-December) where I have divided it in two slots: first months include pilot teachers and lots of testing, in the table it's erroneously called "First 2 months". The second slot covers the time from Nov 18 to Dec 18 when we invited the new teachers (Nov 18/19 in 4 different patches of invites). It is called "the 3rd month" in the tables (again, my mistake). Moreover, the top half of the table has data regarding users who log-in and the bottom with users who did not log in.

    I have mostly the same attributes for both, how did they search; advanced, browsing by category and by tag cloud and how many resources they clicked on (play). The table also contains the number of sessions and number of actions. A session is one consecutive event when the user does something, it's logged. If left idel, the user is logged off in some time. An action is anything, a search, a click on a resource, on a tag, etc. Additionally, we have the contribution by logged-in users, these are tags, bookmarks and ratings.

    Table 2


    As you can see, most of the sessions (above 86%, the second last row) in both slots take place when users are not logged in. Actually, the percentage of sessions stays pretty regular in both slots. Moreover, regarding the actions, we can see that during first months they mostly (70%) came from non-logged in users. However, when we invited the new teachers, we see 10% increase in actions by logged in users (from 30% to 40%). That's positive, as it shows that some of the invited teachers were motivated to contribute.

    There is actually quite big differences in what do these two groups of users do when they are on the portal. Where logged-in users spend about 1/3 of their actions in searching, non logged-in users spent about 2/3 of their actions in searching. Chart 1 shows this clearly, however, I must say that most likely the disparity between the number of searches and plays by non-logged in users in the first months are due to our internal testing. If you compare that to non logged-in users in the 3rd month, you see that there is already less searches and more plays.

    There has been a difference since the new comers (3rd month): within the logged in users, the number of searches executed has gone down (10%), whereas the number of plays has gone up (from 17% to 23%). Among non-logged in users there is the same 10% drop in searches, but plays have gone up by 10% (from 16% to 27%)! That shows that the new comers were interested in seeing what kind of resources were out there in the portal.

    Chart 1 can maybe be used to illustrate

    Chart 1


    One difference can be observed in how differently these two groups seem to search: with logged-in users the advanced search seems to be the more popular way to search (more than 50% of searches are advanced), whereas with the users who are not logged-in browsing (both by category and tag cloud) is more popular. During first months 54% of searches were browsing, which went slightly up (to 56%) during the 3rd month. The tag cloud was the biggest winner in both groups (logged-in and not) at the cost of advanced search. I assume the difference is due to the fact that people who are not logged in are interested in seeing what is out there and browse around to discover learning material.

    In any case, if we look at the figures of non logged-in users within the 3rd month, it's intriguing how equally the searches are distributed across these different ways of searching. We'll keep an eye on this in the future (e.g. when we know that most non-logged clicks come from us testing the portal).

    Consumers and contributors

    In Table 3, where I again have data for users logged-in and not, and by periods of first months and the 3rd month, we see that when users are logged in, they do things differently. First of all, the logged-in users spare much smaller percentage of their actions in searching (average 33% to 75%), however, bizarrely, they still seem to "play" about the same amount of resources (around 20%).

    Table 3


    Within the 3rd month we see the percentage of plays growing. We can assume that the logs from the first months period are most likely influenced due to our internal testing of the portal, which often times includes making searches. We see that the percentage of plays go up for the non logged-in users within the last month (from 16% to 27%), which, I assume shows a more normal user-behaviour than what could be observed before.

    This still indicates that there is lots of inefficiency when non-logged in users search: on the average during the 3rd month, for those logged-in, one "play" was a result of 1.2 searches, whereas with those not logged-in, one "play" was a result of 2.6 searches - lot of time lost in searching. From Table 2 we can observe that there was more browsing (non logged-in users 1 month), I wonder if that was the reason? Have to keep on eye at that one!

    Most interestingly, 40% of actions by logged in users contribute are the ones that contribute something to the portal, they rate, tag and bookmark. Folks who do not log-in are consumers: they only search, click and leave (- which is fine too).

    So all in all, if we look at all the actions on the portal, the contributing actions by logged in users amount to about 17%. Too bad that this figure did not go up in the 3rd month like some others did. Anyhow, it seems to follow the power-law of distribution (20-80), where small amount of people contribute a lot so that other people can take advantage of this work, also know as participation inequality by J.Nilsen (2006).



    J.Nilsen (2006) Participation inequality: Encouraging More Users to contribute

    Wednesday, November 05, 2008

    Power of pursuation/example/recommendation

    We plan to invite about 130 teachers to the portal and I'd like to make a little experiment on this. The idea is to show two different examples of how to access and discover learning resources on the portal, and see whether that has an influence on
    • Uptake: users sign up
    • Retention: users come back
    • Different ways to discover resources (social cues vs. conventional search strategies)
    • Focus on the system (discovery e.g. clicking on resources vs. contributing, e.g. bookmark and tag, rate, etc)
    The point would be to test the statement from the results reported in Harper et al (2007), where an email newsletter with manipulated social comparison made no difference to a member's interest in using the system, but it changed their focus within the system.
    While subjects who received an email message with the comparison manipulation were no more likely to click on one of the links or log in to the system, they were more likely to rate movies.
    There are differences, but in grosso modo the idea has similarities: In Harper et al (2007) the manipulated social information in the email concerned the subject him/herself (treatment group), whereas in my case the information would be about some other teachers that the subjects would be able to relate with (e.g. see favourites from a science or language teacher). The biggest difference would be that there is no comparison aspect of the subject's performance to the other users in the system, which was one of the central features of Harper et al study.

    Study design

    The randomly selected treatment group would receive an invitation with set goals of expectations on the use of the portal. Examples will be given on how to access and discover learning resources based on social navigation cues, show examples of how to browse other user's Favourites and how to access resources through a tag cloud.

    The randomly selected control group will
    receive an invitation with set goals of expectations on the use of the portal. Examples will be given on how to access and discover learning resources would receive an invitation where examples would be given on how to access and discover learning resources based on conventional free text search or browsing keywords
    (to be reworked, just initial ideas based on tracking methods that I could use).
    • Immediate reaction: number of people who access the portal through the direct links on the invitation as opposed to the number of people who access the portal though the main page. the time these people spend on the portal on the first time and how they search, how many searches they execute and how many resources they click on
    • Uptake: is there difference between the groups on signing up on the portal?
    • Retention: on the longer run, say, within a month, do they still come back
    • Different ways to discover resources (social cues vs. conventional search strategies)
    • Focus on the system (discovery vs. contributing in terms of ratings, tagging, etc) : do people who see examples of social navigation use these methods more than the control group? are there any differences how many resources the subjects in both groups tag and rate?

    Hypothesis to test
    (to be reworked, just basic ideas)

    hypotheses would be that subjects in the treatment group will discover more resources than the control due to social navigation cues made readily available to them. By discovering I mean that they click on these resources on the portal to view them. I also would hypothesise that the retention rate is better among the treatment group, as they get a direct access to selected resources whereas the control group has to look for the interesting resources and might get diverted there.

    Relevant studies in this direction, to be completed

    A study towards this direction was reported by Harper et al (2007). They study the effect of email newsletters that told the community members whether their contribution was above or below average. They report that a) previous studies has shown that information about social norms can affect contributions, e.g. people recycled more material when they were provided with information about how much other people had recycled (Schultz, 1999).


    Social Comparisons to Motivate Contributions to an Online Community.
    , Harper, F.; Li, X.; Chen, Y.; Konstan, J. , Persuasive Technology, 26/04/2007, Palo Alto, CA, (2007)

    Using Social Psychology to Motivate Contributions to Online Communities, Ling, K.; Beenen, G.; Ludford, P.; Wang, X.; Chang, K.; Li, X.; Cosley, D.; Frankowski, D.; Terveen, L.; Rashid, A.M.; Resnick, P.; Kraut, R. , Journal of Computer-Mediated Communication, Volume 10, Issue 4, (2005)

    Changing Behavior With Normative Feedback Interventions: A Field Experiment on Curbside Recycling (1999)
    by P Schultz

    Friday, September 12, 2008

    Cross-boundary ranking of learning resources

    Based on the idea of Interest Indicators, like social bookmarks and ratings, I've looked at the data so that we can make the cross-boundary resources better available on the MELT portal.
    The aim is that we can, based on previous users' behaviour :

    a) make separate "travel well" lists of resources that have a potential to cross-borders better,

    b) use this information to rank resources better in the normal search result list,

    c) allow users search for resources that have a good "travel well" value (e.g. give me resources in math that can cross-borders)

    This is the data that I'm using (table below) and this is how I've defined cross-boundary (e.g. cross-country and language) learning resources before. Using that definition I have manually verified the number of cross-country resources. In the dataset, about 82% of resources were cross-country.



    Now, we have a problem, though. On our MELT portal we do not have information about the country where the resource originates from. Dah!

    This is a big blunder (in my opinion) in our Application Profile, we have not defined the country where the resource originates. We do define the provider, and the country information could be inferred from the provider, but it does not always work.

    For example, one of our providers has frequently metadata about resources that do not originate from the same country!

    I've experimented with the data using the information that we have on the portal, which is LOM about the resource including the language of the resource. As we also know the mother tongue of the registered users, this gives us a kick.

    In the table below we can see the coverage of cross-boundary actions that we can get on resources without using any manual labour or verification of the country or language. As a base-line, with manual verification I found that 82% of the actions concerned cross-border rating or bookmarking of a resource.



    The first row represents the cross-language resources (i.e. user's mother tongue is different from the resource language). Just using this information, we get about 65% of resources right, as opposed using manual checking (82%). I think it's pretty good, I'd settle for that! (although I have to look what kind of material was left out!)

    The two other comparisons in the table are based only using information about users' previous behaviour. These would be:
    • rating > 2
    • bookmark
    Only using information regarding bookmarked and rated resources results in a lousy coverage of around 20%. The problem is that 25% of resources bookmarked or rated are on more than one resource, the data still is very sparse.

    Anyway, I want to use that information to "cross-boundary rank" the resources. As we do not know the country where the resource comes from, my work-around is based on countries where these users come from.

    Here is a visualisation about resources that have been bookmarked or rated by users (see also ManyEyes link below). We can see the orange node in the middle, a learning resource called "Five Days in New York..". We see 3 edges leading out to Finland, Belgium and Hungary. This means at least one user from each of these countries has bookmarked the resource!

    So, even if we do not know the origin of the resource, we know that it has users from 3 different countries. I can infer that it is a cross-boundary resource.

    As most likely one of these 3 users come from the same country than the resource comes from, I will minus one country out of the total of countries: (number of countries -1)

    My cross-boundary rank will be the following:
    • Count the number of ratings grater than 2 and/or bookmarks for a resource (actions). Give each action one point
    • Count the number of these users and give each user one point
    • Count the number of user countries of origin. Give each country one point and then minus 1
    • Compare the mother tongue of each of these users to the language of resource. If they differ, give one point/mismatch.
    Then, count the following:
    number of users + number of actions + number of cross-language x (number of countries -1)
    Let's take the above resource "Five Days in New York.." as an example
    • Count the number of ratings grater than 2 (3) and/or bookmarks for a resource (5). Give each action one point. (8)
    • Count the number of these users and give each user one point (5)
    • Count the number of user countries of origin (Hungary, Finland, Belgium). Give each country one point and then minus 1 ( 3-1=2)
    • Count the number of user mother tongue (hu, nl, fi). Compare the mother tongue of each of these users to the language of resource (en). If they differ, give one point/mismatch (3).

    • number of users (5) + number of actions (8) + number of cross-language (3) x (number of countries -1) (2) = 32 Travel well value
    This way you can count a value of "travel well" for each resource that users have previously interacted with on the portal. The value will always be an integer, which is important from the technical implementation point of view (in Lucine index it apparently needs to be an integer).

    The down side is that we'll have a huge cold-start problem. As I said, our data is very sparse. To seed the system, I actually still manually check the new resources that users have interacted with and make a fake bookmark on them so that it looks like it has at least two users from 2 different countries. This way the resource gets a "travel well" value counted and appears on the "travel well" list and is better ranked, etc.

    Of course, at the end I will evaluate how this treatment affects on users, do we, for example, see a big amount of bookmarks on these resources that I have been able to count a travel well value?

    You can see a visualisation here. This is based on on user's country of origin.

    Tuesday, September 02, 2008

    Emerging search patterns on learning resources

    I am hugely inspired by the stuff from J.Feinberg and D.Millen, especially by the studies that they've done on doger, the IBM internal social bookmarking service. I must admit, though, that I had missed on it a bit, I cannot believe! Anyway..

    This paper is really interesting, Social bookmarking and exploratory search (2007), not least for the reason that it offers a very interesting, almost similar study design that I am planning on my log files and search pattern analysis on the MELT portal. (Great minds think a like ;p yeah, right..)

    The study design is a field study of a social bookmakring service in a large corporate (IBM) with quantitative data (click level analyses of log files and boomarking data) and qualitative data like interviews.

    And, this is the data that I collect for my study (Table 1). Quite simlar!

    They use a categorisation of search that I will adapt to my usage:
    • Community browsing (Examining bookmarks created by the community. In my case these could be lists of most boomarked items, travel well; tags; other people Favourites)
    • Personal search (Looking for bookmarks from one's own personal collecction of bookmarks)
    • Explicit search (Explicit search using traditional search box)
    A few days ago I looked at the first logs from Melt. Here is the run-down. I have not used the same types of searches, but I will explore them for the later usage. But basically what you can see:

    There were 512 search events:
    • 41% Explicit searches (adv. search)
    • 34% Community browsing (tagcloud)
    • 25% browsing categories
    What was called "click-through" in this paper is when the particular navigation path resulted in a page view. This is what I call view resources. We can see that there were 538 page views, which implies that most likely users have clicked on more than one resource as a result of a search.
    • 74% of resources were viewed in search result list (srl), this means that they were results from advanced search and browsing categories

    • 20% of resources were viewed as a result of community browsing (e.g. tagcloud, lists)

    • 5% of of resources were viewed as a result of personal search (e.g. in Favourites)

    How many searches resulted in viewing resources? I have to verify this
    • 85% of explicit searches and browsing categories
    • 65% of Community browsing (tag cloud and lists)
    Millen et al. 2007 speculate that in their study, the a higher click-through rate indicates a more purposeful searching, whereas the community browsing was used more as an exploratrory search activity. I will need to keep my attention on this and whether I can make similar conclusions.

    Most likely, anyway, I do not use click-through as an indication of intenet of using a learning resources, what I find most interesting in my study will be how many of the search activities result in a bookmark and/or rating. That is much cooler in my mind than viewing the page. Actually, I've noticed that in our system users view resources a lot, but they do not necessarily show any further interest on them.

    That is why I use both implicit and explicit Interest Indicators. By Interest Indicators I mean Explicit Interest Indicators like ratings as a subjective relevance judgment on a given learning content, Marking Interest Indicators like bookmarks and tags on educational content, and Navigational Interest Indicators such as time spent on evaluating the metadata of educational resources, as well as Repetition Interest Indicators as categorised by Claypool, et al., 2001.

    In this small study we can see that 19% of viewed resources ended up in users' Favourites. Again, ratings were much less, only about 6% of viewed resources ended up being rated. In the study from IBM system, they had 34-39% as high click-through. Will be keeping my eye on that.

    My advisor asked me whether I could find search patterns in my logs. I think I could. In this paper (Millen et al. 2007) they do that :) and here is how: "Looking for Patterns using Cluster Analysis"
    To better see the patterns of use, we performed a cluster analysis (K-means) for the different types of search activities. We first normalized the use data for each end-user by computing the percentage of each search type (i.e., community, person, topic, personal, and explicit search). The K- means cluster analysis is then performed, which finds related groups of users with
    similar distributions of the five search activities. The resulting cluster solution, shown in Table 4, shows the average percentage of item type for each of four easily interpretable clusters.
    Should not be too hard :)

    Millen, D., Yang, M., Whittaker, S., Feinberg, J., Social bookmarking and exploratory search (2007). In L. Bannon, I. Wagner, C. Gutwin, R. Harper, and K. Schmidt (eds.).
    ECSCW’07: Proceedings of the Tenth European Conference on Computer Supported Cooperative. Work, 24-28 September 2007, Limerick, Ireland

    Activity Theory helping us explain folksonomies

    I've been lately reading Engeström's (1) stuff and about Activity Theory. In this paper (2) I found a cool reference that explains how using Activity Theory as a theoretical framework we can study learning resource repositories (LOR) and their communities as one single system rather than as a loose set of instruments, subject, objects and outcomes.

    I think that is a very important point. I've been arguing for quite a while that tags and social bookmarking can be revolutionary for LOR because now we can make a connection between the user, the resource and its metadata.

    Before, it was only the resource and metadata, and a scary looking form for searching the resources (this is time before Google's simple search box, right?). Now, if implemented correctly, social bookmarking and tagging not only helps individuals with their resources management (e.g. Favourites), but also helps other folks to find resources through other users and their digital traces such as tags, number of bookmarks, etc.

    To understand Activity Theory it is important to get the bases: everything, well, everything within human activity, is based on the three dominant aspects which are production, distribution and exchange (or communication).
    The model suggest the possibility of analyzing a multitude of relations within the triangular structure of activity. However, the essential task is always to grasp the systemic whole, not just separate connections. (no page number in my print, just below Figure 2.6).

    This is also the base for the analysis in Margaryan & Littlejohn (2008) for the learning resource repositories as instrument, with rules, division of labour, outcomes, etc.

    Moreover, Engström emphasises that there is no activity without the component of production.
    The specificity of human activity is that it yields more than what goes into the immediate reproduction of the subjects of productions. One part of this "more" is the surplus product that leads to sharing and sociality, discussed by Leakey & Lewin and Ruben above...(found on the next page)

    I was thinking of folksonomies and how the production of tags is first of all good for me. People often times tag and bookmark to "keep found things found", it's part of personal knowledge management activity. Similarly like above, when referred to, for example, production of food that leads to sharing and sociality, in tags, the fact that they are made available to all, leads to sharing and sociality.

    I thought that was pretty neat.