Wednesday, July 15, 2009

Google parsing microformat, e.g. ratings

In May Google announced that they will start parsing microformats (on a small scale first), similar stuff came out from Yahoo! last year, but even on a smaller scale.

This is pretty huge for the end-user generated ratings! I must say that I did not see it coming in this way, which makes it even more exiting :)

..Google is releasing support for parsing and display of microformat data in their search results. .. anyone who marks their pages up with the appropriate microformat data will be able to make their information understandable by Google. This technology would allow you to explicitly search, for example, for only printers that had an average customer review of 3 stars or higher.

Holy smokes! This is cool, can't wait to see when it will first pop up in my search :)

So, since a long time it's been problematic to get enough ratings on items, this is a known problem especially in the field of Recommender systems. They talk about "sparse data". An example, you want to make a recommendation on music, but the item x has 3 ratings, item y 2 ratings, etc. This is way too little to be used to create recommendations using the algorithms that are out there. Take another example, a camera shop, they let users rate their cameras, but they get very little reviews from users.

Now, however, there are other camera shops who are struggling with the same problem. Essentially, they all are selling the same camera brands, and they all have only a few ratings on it, and at the end, non of them can do much fun with this small amount of rather anecdotal information.

There has been talk about a unique identifier set by industry, for example, so that all camera sellers could use them and thus aggregate all the reviews and ratings together. Yep, you guessed it, there's maybe that one shop down the blog who does not want to use it. I think a couple of years back Yahoo! came up with a very compelling paper reiterating the idea and trying to muster up enough consensus among industry and other players. Not much happened - and then, here is Google and microformats... beautiful :)

Why I'm interested in this is that with the idea of federating learning resource metadata across repositories, we face the same problem. As a result of sharing metadata, the same resource might end up used in many different repositories, where users might be allowed to rate them. But that metadata on ratings or evaluations is VERY seldom shipped back to the mother board.

The same with tags and bookmarking (other other tools that allow users to create collections or playlists). That could be valuable information for the repository who first federated the resource metadata out. By collecting back the varied annotations from different repositories, they could gain interesting information, and eventually overpass the sparse data problem. Moreover, they would gain data about what works and in which context, which makes me think of "travel well" resources.

In action

Here is an example of a search for Palm's new phone that I'm contemplating on. I search for reviews only and the result list shows the ratings, but I cannot yet make a query saying "palm pre" ratings grater than 3. Nice in any case.




I've had a few ideas on this with some colleagues and I really look forward to seeing what Google comes up with that. And how are they going to solve the issue of different rating scales used, and multi-attribute ratings.

Vuorikari, R., Manouselis, N., & Duval, E. (2007). Metadata for social recommendations: storing, sharing and reusing evaluations of learning resources. In D. H. Goh & S. Foo (Eds.), Social Information Retrieval Systems: Emerging Technologies and Applications for Searching the Web Effectively (pp. 87-107). Hershey, PA: Idea Group Inc. Retrieved from http://elgg.ou.nl/rvu/files/20/144/SIR_vuorikari_manouselis_duval_web.pdf.


Manouselis, N., & Vuorikari, R. (2009). What if annotations were reusable: a preliminary discussion. In M. Spaniol (Ed.), Advances in Web-Based Learning - ICWL 2009, Lecture Notes in Computer Science (Vol. 5686, pp. 255–264). Berlin Heidelberg: Springer-Verlag.

Friday, July 10, 2009

Tags and self-organisation: a metadata ecology for learning resources in a multilingual context

I think I finally came up with a title for my PhD. You know, the type of title that says it all. It's a bit long, but "correct and descriptive", like Matt said. So here it goes: Tags and self-organisation: a metadata ecology for learning resources in a multilingual context.

Here is a wordle, it was extracted from a paper that summaries the research. Looks pretty accurate :)

Tuesday, June 30, 2009

Study on contexts in tracking usage and attention metadata in multilingual Technology Enhanced Learning

Just submitted the final version of the paper to a workshop on Exploitation of Usage and Attention Metadata (EUAM 09). Here is a one-pager about it and the link to the paper.

Study on contexts in tracking usage and attention metadata in multilingual Technology Enhanced Learning

“Context” is widely accepted to be important for correctly interpreting user input and for improving predictive and possibly also diagnostic models. But what is context, and how can it be measured? By measuring we mean to operationalise the construct and data gathering to provide values for the desired variables.

In this study, we consider the intersection of the areas of digital learning resource repositories, digital libraries and social tagging systems where users from a variety of countries use technology enhanced learning (TEL) offerings in a variety of languages. We consider usage and attention metadata as an example of the wider notion of context adapting the definition of context as “any information that can be used to characterise the situation of entities” [Dey01]. We give an overview of dimensions of context that are relevant in TEL, specifically arguing that context comprises the usage situation and environment as well as persistent and transient properties of the user. Therefore, distinguishing between the macro-context and the micro-context of TEL is useful.

TEL and the analysis of the data it generates take place in different types of educational settings which we call the macro-context of TEL. We use the term micro-context to denote the context that is relevant for interpreting a specific user input and for designing adequate system responses and other output. The micro-context is subdivided into user models, material/environment models, interaction models, and background knowledge, showing that usage and attention metadata are of different types and play different roles for learning about context.

We then concentrate on teachers using learning-resource repositories as an important use-case example of TEL and focus on language and country as context variables. We describe different ways in which these variables are operationalised, and we outline ways in which TEL use such context information to improve the use and reuse of repositories by supporting users in a multilingual and multicultural context. A key theme of our article is the central role that social tagging can play in this process: on the one hand, tags describe usage, attention, and other aspects of context, on the other, they can help to exploit context data towards making repositories more useful, and thus enhance the reuse.

Riina Vuorikari 1,2, Bettina Berendt3
1 European Schoolnet, Brussels, Belgium,
2 OUNL, Heerlen, Netherlands,
3 KU Leuven, Belgium

Thursday, June 25, 2009

My tag paper nominated for best paper award 2009

I'm pretty exited that one of my papers for ICWL 09 was among the 5 best paper nominees. For a some time now I've been wondering what does it take to write a paper that arises above the general mass of papers. Well, now I have a bit better idea :)

What does it take? Reading tons of research papers, write a few (un)successful ones to practice, a good inspiring topic, some research work with ppl who are truly interested in what they are doing, and voila!

I also like how Celstec, OUNL (where I study), picked it up for their news feed. I think that over all, they have a pretty neat way to recognise what's going on and make others aware of it too. A modest person as I am, I would never make any fuss about it.... right.. ;)

Monday, June 22, 2009

Wiley calls it “dirty secret” of OER

Just picked up a fresh PhD study by S. M. Duncan from USU, a student of D.Wiley's. The study is called Patterns of Learning Object Reuse in the Connexions Repository. The punch line is that there is very little reuse of LOs among the repository studied.

What new? Similar findings have been discovered here in Europe (end elsewhere) for a while now. Ochoa (2008), for example, found in his PhD dissertation that reuse in general remains low, about 20%, across all sizes of collections. This was interesting not only for how low the reuse is (20%, common!), but also because since forever folks have been saying that resources with smaller granularity are more reusable, as they lack context, etc (insert here the infamous graph of "modular content hierarchy", the most used LO). Well, according to Ochoa (2008), this was not the case.

I also looked at the reuse on 2 different platforms: LeMill and Calibrate from European Schoolnet. My twist was to study the cross-boundary use and reuse, i.e. teachers reusing learning resources that are in a language other than their mother tongue and originate from different countries than they do. I used the same reuse definition as Ochoa (2008), which basically is the same as in Duncan's study.

The finding was that the general reuse was around 20%, but NOT across all collections. For example, in LeMill, "Multimedia material" was used more often, but in Calibrate, the smaller granularity was seldom added to Collections. The cross-boundary reuse was notably less (37% to 55% of it). Moreover, in some of the collections only around 10% of resources were ever added to a collections, which makes you really think hard about the efficiency of this all..

Anyway, the good news in Duncan's study is this:

There was a common author in 3,722 module uses, while there were only 1,013 module uses where there was no common author. This means that modules were included in collections 3.67 times more often when there was at least one person in common with both the module and the collection.p.32


So if people know each other, they are more likely to reuse material from each other! This shows that social is important when we are talking about the use and reuse of learning resources! This is similar to what I am saying in my PhD thesis, which hopefully will come out one day soon. My twist of course is that tags can make those social connections between people, and by taking advantage of these underlying social connections, we can make the learning resource discovery much better - and hopefully also more useful for teachers.

Vuorikari, R., Koper, R. Evidence of cross-boundary use and reuse of digital educational resources. Link to a revised version of the paper, not reviewed yet!

Monday, June 15, 2009

Testing LeMill for embedding content

LeMill is one of my favourite tools to create online material. It's so simple and easy to use. What I like a lot is that they always follow their time, like here I'm testing how the embedding of my content happens in another platform.

Here is one "Collection" that I've created called "hansin kamaa" = stuff from Hans. It includes two different pieces of content. Apart from exporting the content as a zip-file, I can now also just embed it somewhere, for example in my blog. Pretty neat - and useful!



Testing something else here:

Monday, May 11, 2009

ICT Call 5 info days: European Schoolnet

I'm attending Call 5 infodays tomorrow for European Schoolnet. Here are a few things that we've been working with lately that could be relevant. The speakers look interesting, check them here.

For Large data sets:
  • eTwinning schools: more than 60 000 teachers have signed up. Stats available here. Now with eTwinning 2.0, new data will be available for new types of "connections" and "links" that teachers have.
  • Social Bookmarking data by teachers on learning resources residing on a number of different learning resource repositories in Europe. Some ManyEyes visualisation available to see different types of "connections" or "links" created.
Personal sphere
  • Would be interesting to study how is the personal sphere of a teachers in these days!
Some related slideshows:

Wednesday, May 06, 2009

SIRTEL'09: 3rd Workshop on Social Information Retrieval for Technology-Enhanced Learning

Paper Submission by June 14, 2009
in the International Conference on Web-based Learning (ICWL) 2009
Aachen, Germany, August 21, 2009
http://celstec.org/sirtel

IMPORTANT DATES

Contribution Submission: June 14, 2009
Results Notification: July 13, 2009
Camera Ready Submission: July 31, 2009
Workshop date: August 21, 2009

CALL FOR WORKSHOP CONTRIBUTIONS

We are delighted to welcome exciting new contributions for the 3rd SIRTEL workshop
- Research papers
- System Demos
- Hands-On proposals
- Abstracts for "Pecha Kucha"


RATIONALE

Learning and teaching resource are available on the Web - both in terms of digital learning content and people resources (e.g. other learners, experts, tutors). They can be used to facilitate teaching and learning tasks. The remaining challenge is to develop, deploy and evaluate Social information retrieval (SIR) methods, techniques and systems that provide learners and teachers with guidance in potentially overwhelming variety of choices.

The aim of the SIRTEL’09 workshop is to look onward beyond recent achievements to discuss specific topics, emerging research issues, new trends and endeavors in SIR for TEL. The workshop will bring together researchers and practitioners to present, and more importantly, to discuss the current status of research in SIR and TEL and its implications for science and teaching.

The proceedings from the last years:
· SIRTEL'07 http://ceur-ws.org/Vol-307
· SIRTEL'08 http://ceur-ws.org/Vol-382


TOPICS OF INTEREST (but not limited to):

  • Recommender systems and collaborative filtering in educational settings
  • Defining the scope, purpose and objects of social information retrieval in TEL
  • Novel ways of generating input for recommenders (explicit and implicit methods)
  • Ranking of search results to support individualised learning needs
  • Integrating SIR services in existing educational platforms
  • Folksonomies, tagging and other collaboration-based information retrieval systems
  • Social navigation processes and metaphors for searching information related to teaching and learning
  • Social networks and interactions in learning communities to facilitate information sharing and retrieval
  • Approaches to TEL metadata reflecting social ties and collaborative experiences in the field of education
  • Pedagogic decisions, recommender systems and how to contextualise recommender system to support learning processes.
  • Interoperability of SIR systems for TEL
  • Visualisation techniques in learning and teaching
  • Semantic annotation and tagging for social information retrieval purposes
  • Evaluating the performance of SIR systems in educational applications
  • Measuring the effectiveness of SIR systems in supporting learning and teaching
  • Evaluation the user satisfaction with SIR systems in supporting learning and teaching


WORKSHOP SUBMISSIONS

The workshop invites several types of contributions which allow a wide level of participation:
· Research papers (upto 8 pages)
· System Demos (upto 2 pages)
· Hands-On proposals (1-pager)
· Abstract for Pecha Kucha (1-pager)

The workshop proceedings will be published as CEUR Workshop Proceedings online at http://ftp.informatik.rwth-aachen.de/Publications/CEUR-WS/. Copy rights will be reserved.

Please use the same template as the one for the main conference with details at http://www.hkws.org/events/icwl2009/submission.html. Workshop paper length is not limited.

All questions and submissions should be sent to: sirtelworkshop@gmail.com

PROGRAM COMMITTEE
  • Alexander Felfernig, Graz University of Technology, Austria
  • Brandon Muramatsu, MIT, USA
  • Frans van Assche, European Schoolnet, Belgium
  • John Dron, Athabasca University, Canada
  • Lloyd Rutledge, OUNL, The Netherlands
  • Markus Strohmaier, Graz University of Technology, Austria
  • Markus Weimer, Technical University of Darmstadt, Germany
  • Martin Wolpers, Fraonhofer-Institut, Germany
  • Miguel-Angel Sicilia, University of Alcala, Spain
  • Olga Santos, UNED, Spain
  • Rick D. Hangartner, MyStrands,USA
  • Rosta Farzan, University of Pittsburgh, USA
  • Styliani Kleanthous, University of Leeds, UK
  • Tiffany Tang, Hong Kong Polytechnic University, China
  • Wolfgang Reinhardt, University of Paderborn, Germany
  • Xavier Ochoa, Escuela Superior Politecnica del Litoral, Ecuador
  • Yiwei Cao, RWTH Aachen University, Germany
  • Zinayida Petrushyna, RWTH Aachen University, Germany


ORGANISERS

* Riina Vuorikari, European Schoolnet (EUN), Belgium and CELSTEC, OUNL, Netherland
* Hendrik Drachsler, CELSTEC, OUNL,The Netherlands
* Nikos Manouselis, Greek Research & Technology Network
* Rob Koper, CELSTEC, OUNL, The Netherlands


ABOUT ICWL 2009

ICWL is an annual international conference on web-based learning. Since the first ICWL was held in Hong Kong in 2002, it has been held in Australia (2003), China (2004), Hong Kong (2005), Malaysia (2006), United Kingdom (2007), and China (2008). The 8th ICWL 2009 will be held in Aachen, Germany, a city with rich culture, high-tech research, and a truly European spirit. ICWL 2009 will be jointly organized by Hong Kong Web Society, RWTH Aachen University, and Max-Planck-Institute for Computer Science.
TAG THIS

Feel free to blog about this and social bookmark the call! Use the tag "sirtel09".
*************************************************************

Tuesday, May 05, 2009

Challenges and lessons learned from Social tagging in MELT

Social tagging in MELT, how do we want to take the social tagging work forward

Social tagging of educational resources potentially offers new ways for:
  • Individuals to
    1.1) better manage their digital learning resources that reside in different repositories and platforms, and

    1.2) discover and access new resources from different contexts (e.g. different language, educational system) through tags and other users.

  • LOR managers to
    2.1) get third party metadata on learning resources (either the ones that already reside on their repository, or the possible new ones to be added to collections,

    2.2) create affinities (e.g. link structure) between separate pieces of resources (either on their own repository, or the ones that reside on other repositories on the federation or on the Web) that were not cross-referenced before.

  • In the MELT project so far, we have only been able to see the peak of these potentials emerging. We list issues that we see important for future work in the field, for the clarity, we only list one of the main issues for each topic:

    • 1.1 To fully support users in their knowledge management task on digital learning resources, the bookmarks (including title, url and tags) should be exportable in standard Webfeed formats. This would allow users to access and manage their MELT resources as part of their other resources collections, whereas now users need to be logged on to the MELT portal to do this.

    • 1.2 Pivotal browsing of social bookmarks takes advantage of the affinities between the user, resource and tags. In the MELT context, more metadata could also be added to support pivotal browsing, such as the country of the user, interest topics; resource metadata such as multilingual indexing keywords. This would allow novel ways to access resources that other users have already discovered within the federation, and thus build on users’ social interactions and co-construction of knowledge.

    • 2.1 Tags by end-users on the MELT portal have been shown to be of good quality as additional metadata descriptors of resources. We have enumerated possibilities of metadata ecology that the use of multilingual Thesaurus can offer to a federation such as LRE. Apart from working on ways to automatically generate LOM from tags, we urge on using the hierarchical structure and multilingual features to leverage user-generated tags.

    • 2.2 Why not do Google for learning resources? Using PageRank-like algorithms on a learning resource repository or federation has been impossible for a number of reasons, the most important is the lack of a link-structure that cross-references resources. Tags, creating underlying connections between seemingly random pieces of content in different languages, on repositories in different countries and other platforms on the Web, rely on humans’ subjective idea of its importance for a given information seeking task. Using this new, emerging link-structure with tags as “anchor texts” offers totally new ways to “organise the world's learning resources and make them universally accessible and useful”. A new tag line could be “From teachers to teachers”.

    Monday, May 04, 2009

    Link structure and anchor text

    I read that Brin & Page (1998) paper again. A few guidelines to keep in mind:
    ..our notion of "relevant" to only include the very best documents since there may be tens of thousands of slightly relevant documents. This very high precision is important even at the expense of recall (the total number of relevant documents the system is able to return).


    Two features to produce high quality precision:
    • Link structure is used to create objective measure of its citation importance that corresponds well with people’s subjective idea of importance. Well, it's that simple..

    • Anchor text:
      ..anchors often provide more accurate descriptions of web pages than the pages themselves. Second, anchors may exist for documents which cannot be indexed by a text-based search engine, such as images, programs,..
    The point about the anchor text is so interesting, I wonder how well does it apply to tags? I bet really well..

    I also found this interesting: "it has location information for all hits and so it makes extensive use of proximity in search"

    Differences Between the Web and Well Controlled Collections
    • extreme variation internal to the documents: documents differ internally in their language (both human and programming), vocabulary (email addresses, links, zip codes, phone numbers, product numbers), type or format (text, HTML, PDF, images, sounds), and may even be machine generated (log files or out putfrom a database).
    • external meta information as information that can be inferred about a document, but is not contained within it. Examples of external meta information include things like reputation of the source, update frequency, quality, popularity or usage, and citations. Not only are the possible sources of external meta information varied, but the things that are being measuredvary many orders of magnitude as well.


    http://www.scribd.com/doc/3208417/The-Anatomy-of-a-LargeScale-Hypertextual-Web-Search-Engine

    Saturday, May 02, 2009

    Cross-language use of the Web; users behaviours and attitudes

    Berendt & Kralisch (2009) A user-centric approach to identifying best deployment strategies for language tools: the impact of content and access language on Web user behaviour and attitudes

    The results indicate that non-English languages are under-represented on the Web and that this is partly due to content-creation, link-setting and link-following behavoiur. User satisfaction is influenced both by the cognitive effort of searching and the availability of alternative information in that language.

    Cost=time+cognitive effort

    Not only capacities to access the site but also opportunities to access it, thus language is only one factor.
    • Language can be expected to not only influence the total amount of information available to Web users, but also how information sources (i.e. Websites) are linked among each other and therefore how easy/likely it is to find and access a certain Web site.
    • Bharat et al. and Halavis are first indicators of the potential impact of language: Website in different languages are less connected than sites in the same language (note: studied data aggregated on the national level and therefore only limited insight into the role of language.
    "Web sites are, in most cases more likely to link to another site hosted in the same country than to cross national borders. When they do cross national borders, they are more likely to lead to pages hosted in the United States than to pages anywhere else in the world." (Halavais, A, 2000, p. 7)

    Behavioural aspects of information seeking:
    1. Users' information seeking behaviour,
    2. information and information flow on the Web,
    Attitudinal aspects of information seeking:
    1. "usefulness", i.e. the language related value of information decreases as more information is offered in that language on the Web. "..value perceptions are also determined by topic; thus a large amount of content on a topic in a native language may also reduce the value of content on that topic in other languages.

    2. "ease of use", i.e. the cost of language processing during information seeking can be expected to affect attitudes in Web search.
    Results on behavioural aspects

    1. Non-English languages are under-represented on the Web in terms of the amount of content supplied.
    2. Search engines do not register all pages linking to the site, and many links known to the search engine were not used. This indicates that non-English language s are under-represented on the Web in terms of the links that content creators set to content in those languages.
    3. Users have a clear preference to navigate in their native language when it is available via a link, but if that is not available they accept the necessity to navigate in English.
    4. This all means: behavioural tendencies both of content providers and of content users lead to mutually reinforcing under-representation of non-English languages. Compared to the respective market size or available options, there is less content in these languages, this content is linked to less and the links are followed less often.
    Results on Attitudinal stuff:
    1. A complex interplay of English language skills, the perceived saved effort of using native-language content, the perceived overall supply in that language on the Web, and satisfaction:

    2. People who are proficient in English often prefer to navigate in English (even if offered content is their own language) and are more scrutinised of the quality of Web content. Do not care much about whether sites make efforts to provide them with content in their own languages.

    3. People who are not so proficient in English do perceive the (real) scarcity of information in their native language and are highly appreciative of content in this language.
    This means that content and search-tool designers should not draw simplistic conclusions based on behaviour alone, because this is not a reliable indicator of attitudes and preferences. In the absence of links and/or content in their native languages, users will acquiesce to English-language content. However, their preference will persist.

    Berendt, B., & Kralisch, A. (2009). A user-centric approach to identifying best deployment strategies for language tools: the impact of content and access language on Web user behaviour and attitudes. Inf. Retr., 12(3), 380-399.



    HALAVAIS, A. (2000). National Borders on the World Wide Web.New Media Society, 2 (1), 7-28. doi: 10.1177/14614440022225689.



    Bharat, K., Chang, B., Henzinger, M. R., and Ruhl, M. 2001. Who Links to Whom: Mining Linkage between Web Sites. In Proceedings of the 2001 IEEE international Conference on Data Mining (November 29 - December 02, 2001). N. Cercone, T. Y. Lin, and X. Wu, Eds. ICDM. IEEE Computer Society, Washington, DC, 51-58.

    Thursday, April 02, 2009

    A touch screen for schools for less than 50€

    I love the do-it-yourself attitude of some e-learning tech support guys! Marko Puusaar just showed to us here in the e-university conference what he hacked together based on http://johnnylee.net/projects/wii.

    In the picture you can see the touch screen for less than 50€. It took him about 15min, and that is with all the explanations of parts included!!



    The pieces needed are: a Wii remote control, which will work through blue-tooth with your computer, a bit of software (there are pay versions, or free ones for educational use), and a Infra-red pen, which he did himself to look like a normal pen.


    Tuesday, March 24, 2009

    The Ada Lovelace pledge "Ms. Mayer"

    Some time ago I pledged to this one: "I will publish a blog post on Tuesday 24th March about a woman in technology whom I admire but only if 1,000 other people will do the same." I do to honor Ada Lovelace.

    So here I go: since the first moment I set my eyes on it, I thought there was something that set it apart from the crowd. It must have been sometimes in 2000 or so. The name was catchy too, but what I most admired was the plain, simplistic look, almost too little, and yet, everything was there. Ever since I've admired the almost iconic look of it.

    A couple of weeks back when visiting D&D in San Fransisco, we were drinking coffee and reading the Sunday edition of New York Times, I got across an article about Google's design. I learned that "Ms. Mayer controls the look, feel and functionality of the Internet’s most heavily trafficked search engine."

    Of course, I thought, that is why the interface looks so DAM GOOD, it's a she!

    The article got a few "Hyvä Suomi!" when I learned that since a kid she had admired the design of Marimekko, something that every Finnish kid from the seventies has imprinted in their brain. So, she's got Finnish ancestors too! "Hyvä Suomi!"

    Apart from being the employee no: 20 and Google's first female engineer, she seems to like a good party and is comfortable on skis. Petty darn impressive! Thanks for being there!

    http://en.wikipedia.org/wiki/Ada_Lovelace
    an interesting interview on Marissa Mayer (where she wears an awful shirt, oups...)
    http://www.charlierose.com/view/clip/10136

    Monday, March 23, 2009

    Sneak preview: SIRTEL'09

    Workshop on Social Information Retrieval for Technology-Enhanced Learning (SIRTEL'08) in the International Conference on Web-based Learning (ICWL) 2009 in Aachen, Germany, August 21, 2009
    http://www.hkws.org/events/icwl2009/workshops.html


    IMPORTANT DATES

    Contribution Submission: June 14, 2009
    Results Notification: July 13, 2009
    Camera Ready Submission: July 31, 2009
    Workshop date: August 21, 2009

    CALL FOR WORKSHOP CONTRIBUTIONS

    We are delighted to welcome exciting new contributions for the 3rd SIRTEL workshop
    • Research papers
    • System Demos
    • Hands-On proposals
    • Abstracts for "Pecha Kucha"
    RATIONALE

    Learning and teaching resource are available on the Web - both in terms of digital learning content and people resources (e.g. other learners, experts, tutors). They can be used to facilitate teaching and learning tasks. The remaining challenge is to develop, deploy and evaluate Social information retrieval (SIR) methods, techniques and systems that provide learners and teachers with guidance in potentially overwhelming variety of choices.

    The aim of the SIRTEL’09 workshop is to look onward beyond recent achievements to discuss specific topics, emerging research issues, new trends and endeavors in SIR for TEL. The workshop will bring together researchers and practitioners to present, and more importantly, to discuss the current status of research in SIR and TEL and its implications for science and teaching.

    The proceedings from the last years:
    - SIRTEL'07 http://ceur-ws.org/Vol-307
    - SIRTEL'08 http://ceur-ws.org/Vol-382


    TOPICS OF INTEREST (but not limited to):

    • Recommender systems and collaborative filtering in educational settings
    • Defining the scope, purpose and objects of social information retrieval in TEL
    • Novel ways of generating input for recommenders (explicit and implicit methods)
    • Ranking of search results to support individualised learning needs
    • Integrating SIR services in existing educational platforms
    • Folksonomies, tagging and other collaboration-based information retrieval systems
    • Social navigation processes and metaphors for searching information related to teaching and learning
    • Social networks and interactions in learning communities to facilitate information sharing and retrieval
    • Approaches to TEL metadata reflecting social ties and collaborative experiences in the field of education
    • Pedagogic decisions, recommender systems and how to contextualise recommender system to support learning processes.
    • Interoperability of SIR systems for TEL
    • Visualisation techniques in learning and teaching
    • Semantic annotation and tagging for social information retrieval purposes
    • Evaluating the performance of SIR systems in educational applications
    • Measuring the effectiveness of SIR systems in supporting learning and teaching
    • Evaluation the user satisfaction with SIR systems in supporting learning and teaching

    WORKSHOP SUBMISSIONS

    The workshop invites several types of contributions which allow a wide level of participation:
    • Research papers (upto 8 pages)
    • System Demos (upto 2 pages)
    • Hands-On proposals (1-pager)
    • Abstract for Pecha Kucha (1-pager)

    The workshop proceedings will be published as CEUR Workshop Proceedings online at http://ftp.informatik.rwth-aachen.de/Publications/CEUR-WS/. Copy rights will be reserved.

    Please use the same template as the one for the main conference with details at http://www.hkws.org/events/icwl2009/submission.html. Workshop paper length is not limited.

    All questions and submissions should be sent to: sirtelworkshop@gmail.com
    The website with the call will be up shortly too.

    PROGRAM COMMITTEE:
    • Alexander Felfernig, University of Klagenfurt, Germany
    • Brandon Muramatsu, MIT, USA
    • Frans van Assche, European Schoolnet (EUN), Belgium
    • John Dron, Athabasca University, Canada
    • Lloyd Rutledge, Open University of the Netherlands, NL
    • Markus Strohmaier, Technical University of Graz, Austria
    • Markus Weimer, Technische Universität Darmstadt, Germany
    • Martin Wolpers, Fraonhofer-Institut, Germany
    • Miguel-Angel Sicilia, University of Alcala, Spain
    • Olga Santos, UNED, Spain
    • Rick D. Hangartner, MyStrands, USA
    • Rosta Farzan, University of Pittsburgh, USA
    • Wolfgang Reinhardt, Universität Paderborn, Germany
    • Xavier Ochoa, Escuela Superior Politécnica del Litoral, Ecuador
    • Yiwei Cao, RWTH Aachen University, Germany
    • Zinayida Petrushyna, RWTH Aachen University, Germany

    Wednesday, March 04, 2009

    OER Creating connections: content, users, tags

    I'm in the Hewlett grantees meeting now and marveling all the work that has been done in the area of Open Educational Resources. I had a chance to present some of the work here that we are doing with OER. I focused on creating connections, which I think is one of the most important things for the content. Knowing how it all is connected together is important in order to make sense of all the small pieces of separate content. Here are the slides, I give a few explanations.



    The LRE is an access point to 16 content providers who make resources available to teachers in Europe and elsewhere. It's pretty much like any conventional content portal, but we have added a social tagging and bookmarking tool on top of it.

    A good part of the slides show how we can make connections between pieces of content from different repositories, and how those content pieces can bring users togeher across country and language borders. In the visualisations, the little dots (nodes) are resources that are connected to other resources or users through tags and bookmarks.

    Moreover, I give a few pointers to research that we have done on creating better navigation to the content, evaluating whether it is efficient or not, and also looking into the quality of tags.

    Friday, February 27, 2009

    Are tags from Mars and descriptors from Venus?

    A study on the ecology of educational resource metadata.

    I just finished a paper on the tag evaluations that we did in the MELT project. We had lots of fun with the name of the paper :) the main question being which one, tags or descriptors, should be from Venus...?

    Anyway, we were able to show that not all the tags are as far from the Thesaurus descriptors as Mars is from Venus. We had different perspectives for evaluations: end-users, expert indexers and repository owners. For me the most interesting thing that came up was that 11% of end-user generated tags are actually terms that we can find in our multilingual Thesaurus! I assume teachers are "better taggers" than average, usually there is lots of talk about the gap between end-users' language and the one deployed by experts.

    Abstract. pdf. In this study, over a period of six months, we gathered empirical data from more than 200 users on a learning resource portal with a social bookmarking and tagging feature. Our aim was to look at the tags from different stakeholders’ points of view; end-users, librarians/expert indexers and repository owners. We first look how users tag resources, and then conduct an evaluation with indexers to understand how they perceive the value of tags as descriptors. We then present a case study from a repository owner’s point of view. Lastly, we study users’ clickstream when searching resources. We find that, even though end-users and expert evaluators apply very different strategies when adding metadata, (end-users have a rather synthetic approach whereas expert indexers an analytical one) there is an overlap in the information in tags and the official descriptors, this overlap is even up to 51%, creating an ecology of metadata.

    Keywords: Learning resource metadata, tags, folksonomy, clickstream,
    thesaurus, evaluation.





    Monday, February 09, 2009

    "Thesaurus-tags"

    One of the particularities of the MELT portal is that apart from being a "traditional" resource portal, we also have social tagging-features implemented. This creates a situation where resources have both indexing terms that come from our multilingual Thesaurus, as well as teacher generated tags, that, btw, are also multilingual.

    I was looking at the tags today with a specific question in mind: "How many of the user-generated tags are actually terms that exist in the Thesaurus?" If there is a tag that is added by a teacher, and if it exists in the Thesaurus, I will call it a "Thesaurus-tag".

    Here are the figures:
    • Distinct tags: 4428
    • Tags applied: 5009
    • Distinct "Thesaurus-tags": 505
    • "Thesaurus-tags" applied: 714
    I was actually really impressed: 11.4% of distinct tags are "Thesaurus-tags"! And if we look at tag application, "Thesaurus-tags" amount to 14.25% of all tags.

    Moreover, 22.37% of distinct "Thesaurus-tags" are applied more than once. This amounts to 45.1% of all "Thesaurus-tag" applications! The top "Thesaurus-tag" were:
    Europe (10), music (8), test (8), Vocabulary (8), Internet (7), art (6), biology (6), history (6), Australia (5), chemistry (5).

    This is really quite interesting. On the one hand, we always ask ourselves how to make the Learning resource indexing better, and here we can totally "crowd source" part of the indexing, that is usually done by the experts", to end-users. If we think an end-user thought that a Thesaurus-tag was good enough to add for a resource, I can be pretty sure that it is also good enough for being an indexing term. The story can be VERY different for all tags, we do not think that ALL tags could become indexing terms, although we know they are good for other stuff.

    On the other hand, being in the multi-lingual context, we have always a bit hard time with tags in different languages. At least with "Thesaurus-tags" we could easily show a translation of the tag, as we are certain it to be a good one.

    There are lots of interesting things to see, for example, whether these resources previously had the same indexing term as the "Thesaurus-tag" was, i.e. was the tag redundant or does it really add some value to our system. It will also be interesting to see whether there was a trend; the resources that had poor/little indexing terms received more "Thesaurus-tags" from the end-users. Well, lots more, I guess, but that will be for another time.

    Tuesday, January 13, 2009

    Modelling the portal ecology: What goes around comes around

    The three main actions on the portal: discover resources, play them and annotate. The two main group of users: ones logged-in and the others not.

    I've divided the resource discovery process in three slots (Millen et al.):
    1. Explicit search
    2. Community search
    3. Personal search
    Play is when the user clicks on the link. We also call this implicit interest indicator, however, we are not sure whether it was relevant to the user or not. Worth noting anyway. This is also called "hits" or "click-through" in some lingo.

    Annotation is when the user makes an explicit interest marking (indicator) on the resource, this can currently be either a rating (usefulness, scale 1 to 5) and bookmark with tags. Both of these actions are public.

    About users and logs

    In general terms we record all kinds of clicks and actions on the portal (see here). I studied the logs from the last 2,5 months. We know that we have 340 users who have a user name, excluding staff, etc., we have 168 "real" users. Out of them 82 had clicked on a resource on the portal at least once, so these users are included in the logs. Additionally, there are users who do not log, but I do not have any idea currently how many they are (check Analytics). There were 13 604 actions recorded, 40% from the ones who logged in and 60% by others who did not log in.

    In general, we can think that the relationship between these 3 actions is important on the portal and can indicate something about its efficiency for users to get what they want, as well as for the system to get what is needed to keep it going. In our case, we are in the process of looking at how Social information can help the discovery. So, a perquisite is to have SI available, thus the system needs ratings and bookmarks.

    With the contributing users (=logged-in) on the MELT portal:
    • 2 searches result to one play;
    • 2.6 searches result to one annotation, this can be either rating or bookmarking;
    • 1.3 plays result to one annotation.

    For the comparison, in Calibrate the figures were the following:
    • 0.5 searches result to one play;
    • 5.7 searches result to one annotation, this can be either rating or bookmarking;
    • 11.3 plays result to one annotation.
    I will study this further too. A quick look would say that a system, which emphasises Social Information for users own benefits (Favourites) and for everyone's benefits (allows Community browsing) like is the case with MELT, the loop for getting annotations is more efficient than in the system which does not make use of such information (e.g. Calibrate). The ration of search-to-annotation is 2.6 to 1 in MELT, whereas the same in Calibrate is 5.7 to 1.

    However, if we look at the ration of "hits", the Calibrate system has been four times more efficient: it took one search to play 2 resources, whereas in MELT it took 2 searches for one play. The MELT system search function has been under constant development for speed, which has been somewhat problematic due to huge amount of content. I will report later on the same ration after our last optimisation effort.

    What goes around comes around

    Graph 1 depicts what is going on on the portal. I will explain this later in details. For each action I have indicated the percentage of total, e.g. Explicit search 78% is from all Explicit searches executed by non-logged in users.


    Graph 1


    Monday, January 12, 2009

    Discovering cross-boundary resources on the portal

    This is continuation for the previous post, the same data-

    The question now is how and where do users discover cross-border resources? By this I mean that the user and the content come from different countries, and/or that the content is in a language other than the user’s mother tongue.

    One challenge is discovering resources in general (previous post), another one is to discover resources that are in a different language and come from different countries. This is the case on our portal, so I'm interested in how to facilitate the discovery process that involves crossing those boundaries (language and national, which might have implications to the educational content of the resource), and hopefully make it more efficient.

    My take is social information: making it readily available to all to leave cues from other users. I bet that this should facilitate the discovery process and thus also make it more efficient (more resources and faster).

    So I looked at the previous data and the cross-boundary bookmarks from users. This is relative to the user, of course, so with every bookmark I compare whether the resource and the user come from the same country and/or is the users mother tongue different from the resource language.

    41 out of 48 active users had bookmarked cross-boundary resources. They had added
    • 299 distinct resources into their collections 350 times. Out of these resources
    • 163 were cross-boundary resources, which had been added to their collections 198 times.
    • This means that 55% of distinct resources obtained during the period of 1.5 months were cross-boundary resources.
    So it was interesting to see how many of these resources had Social information added to them?

    Table 1


    Table 1 shows that about a third of the bookmarked cross-boundary resources had Social Information on them, they were either bookmarked by previous users or existed in the "travel well" list. This is cool! Although we cannot say that the users only discovered these resources because of Social Information, it is important to know that it has had added value for users to discover them. I also found out that about 10% of bookmarks on resources with SI had been previously bookmarked by someone from the same country as the resource was. The fact that they were bookmarked, although still almost dismissible small (10%), is still good news for SI and social navigation based on it.

    I then looked at where the resources are discovered: 62% of cross-boundary discoveries were done in Search Result List (SRL) whereas 37% took place in Community searches, most of them in the tag cloud (30%) and 7.5% on the Travel well and Most bookmarked lists (only one case in the latter :(

    As comparison for the resource discoveries that did not cross any borders (e.g. German teacher found German resources), 90% of them took place in SRL. So it seems that for discovering cross-boundary resources the Social Information is important, as it allows users to do Community searches to discover these resources. We do see, though, that cross-boundary discovery is efficient also within the resources that cannot leverage the previous user experiences, as 63% of cross-boundary resources do not have any SI.

    Interestingly, when we look at Table 1, we see that for the resource discovery that does not imply any cross-border action, users do not seem to care much for SI. Actually, more than 90% of these bookmarked resources had no SI. This is cool, as it seems that we need users to discover and annotate resources among their comfort zone (e.g. national and regional educational material in their own language) in order to make it more readily available to others.

    Another thing that I've looked at are these measures for
    • usage coverage within the repository,
    • how many resources are shared among collections (e.g. Favourites) and
    • what I call the pick-up rate, this is how many of distinct used resources are reused. This could happen when someone discovers a resource that has SI related to it (e.g. in the travel well list or from other user's favourites, or just picks it up from SRL based on someone else's annotations). These were discussed in more details in the last paper.
    Table 2



    Table 2 shows these measure in LeMill, Calibrate and delicious, additionally, the gray column indicates this 1.5 month trial in MELT and the last column has all the data from MELT, which includes the pilot teachers, staff, partners, etc.

    We can see that as the initial amount of resources is so big in MELT, we still cover only a minimal amount of resources, from 1 to 2 %. This does not even include the assets, which more than doubles the amount. Used here means that the resource is added to Favourites once (reuse more than once).

    We can see, though, that even if the resources coverage is not that high, there is still quite a lot of sharing among used resources. This figure still remains low with 1.5month trial (about 15%, same as in Calibrate which did not make the SI available!), but if we look at all the usage so far, the sharing is at 43%. This is somewhat artificial, though, as there is a lot of staff use, but still I hope it indicates that making SI available on the long run helps sharing resources (or, I need to look better solutions on the portal for sharing, which is also planned but super delayed because of all the other dev programme).

    We can already see that the pick-up rate is higher in MELT trail than in the 3 other platforms that I've looked previously. This is an indication that SI works, I hope.

    Btw, I could not find any correlation between the act of putting the resources in Favourites and whether they have Social information related to them (I got some lousy 0.18 even if I removed 2 outliers who outperformed anyone else). Also, in some previous test I had hard time finding significant changes, so I might need to seddle with increased reuse rate, which has previously been shown to be abotu 20% across collections. I found that it was about half of this (or the same as general reuse) in my previous paper.

    Saturday, January 10, 2009

    Where does a CC-licence take you?

    ..or better, your photos? Some time ago one of my pics was "selected" for a travel guide from Prague. That was pretty cool and I really appreciated it. Tonight, I was checking my Flickr stats and noticed that one photo had been viewed many times recently. So I looked what's up with that.















    I followed the referrals. Here are the top 3 sources, some other hits were through search engines for generic keywords like ski, etc.

    1. It was the main picture on the Facebook group for Sankt-Anton fans.


    2. It was on a Yahoo! travel website for best skiing in America. They had a Flickr badge with some random ski shots, et voila moi! (see the small image on the right corner)



    3. The funniest of all, though, is that it is one some German photo website where the dude discussed how the picture should have been framed differently! Hmm..I think my boyfriend, who took the picture, did not appreciate this lesson..










    It's just kinda weird to find yourself in odd places on the web, but hey, isn't that what the creative commons license is supposed to allow. So this is just one consequence of it, I guess.

    Talking about odd places to be on the Web, the other night my Google alerts had picked up that my blog post appeared on a porno site. Sure enough, there after lots of photos of youknowwhat, were feeds from my latest blog post. Pretty hilarious... I don't recommend clicking on the link, but you can read the paper though!