Showing posts with label network. Show all posts
Showing posts with label network. Show all posts

Tuesday, November 15, 2016

Tools for Calculating Academic Collaboration Distance

I think most of you have heard about the Erdös number. The Erdös number is the number of edges between you an Erdös in an author collaboration graph.
This is an undirected network where every published paper defines egdes between their authors. Having a low Erdös number somehow became a status symbol for researchers. Since Erdös already passed away, there is no way to get an Erdös number of one today, unless you hope for a Zombie apocalypse with the death rising:

Excerpt from "Apocalypse" by Randall Munroe at xkcd.com under CC-BY-NC 2.5


Due to Paul Erdös' outstanding publication productivity, there are quite a number of people with an Erdös number of 1, so if you find the right collaborator, you can reach an Erdös number of 2, if you like. But even beyond the fad on Erdös numbers, author collaboration graphs and distances between authors are an interesting way to define closeness between the work that two academics are doing.

What are good tools to calculate author collaboration distance?

There is MathSciNet, but their database only includes mathematical journals. Since my research is mostly published in computer science/embedded systems journals, this site doesn't work for me.
The zbMATH page offers a similar tool, again it seems to include only mathamatical journals. I should publish more there.

Previously, Microsoft Academic Research had a nice author collaboration search that graphically displayed the connections between any two authors. However, this feature is currently not available, since the page was restructured to work without the Silverlight plugin. I hope the feature comes back someday.

Distance calculator at csauthors.net

Currently, the best tool for computer scientists is the distance calculator at csauthors.net. It works with a database that seems to be more complete than the ones used by the sites cited above. The database is however far from being complete, so that distances are sometimes reported to be longer than they actually are.

My Erdös number

Thanks for asking! It is 3, for example via the following papers:


All papers are on the topic of networks or networked systems. How fitting.

Sunday, October 19, 2014

On the Road to your PhD

In the blog Between a rock and a hard place, James Hickey posted a valuable list of tipps for succeeding in your PhD:

1. Learn Latex
2. Use Bibtex
3. Keep your papers organised
4. Keep a formatted list of your own publications and conference abstracts as you go along
5. Always give conference abstracts different titles
6. Keep on top of your emails
7. Manage time
8. Hypothesis testing
9. Keep detailed notes
10. Avoid perfectionism
11. Always give deadlines when you want feedback
12. Source additional funding
13. Write as you go
14. Don’t be scared of your supervisors
15. Log out of Facebook
16. Keep an eye on your budget
17. Diversify yourself
18. Music
19. Get your workstation set up
20. Take notes in meetings
21. Read around your subject
22. Write a literature review
23. Socialise!
24. Sport and/or hobbies
25. Go to conferences and workshops
26. Network
27. Establish a routine
28. Take the lead
29. Practice presenting your work
30. Be prepared for the worse
31. Back up, and back up again
32. Small steps to success
33. Keep on top of admin
34. ENJOY IT!

Being somebody who already has his PhD, I find the list very useful, the title "Things I wish I knew when I started my PhD..." definitely holds some truth, although I accidently did many of the good things mentioned there. See the original post for a detailed explanation of every item!

Monday, September 16, 2013

The Complex Systems Community Explorer

If you are working in the field of complex systems, you face network analysis and graphical data representation. So why not use these features to organize your research network and to identify possible collaborators?
The complex systems community explorer developed at ISC-PIF by Julian Bilke and David Chavalarias is doing exactly this. In particular, it visualizes data are taken from the complex systems registry directory. This directory is an open directory maintained by several complex systems organisations and coordinated by the Complex Systems Society. After registering your data and interests, you can explore other scholars graphically. Links symbolize how semantically close two researchers are. The more shared keywords match, the stronger the link.


Wednesday, February 13, 2013

You don’t cite me anymore - Scientific publications and the ravages of time

One of the most specific things about scientific literature is that scientific papers and books contain references to other papers. The number of citations has become an indicator for the impact of a result, the more other papers cite an article, the higher is its considered impact.

The number of citations a scientific paper gets is a result of interesting effects and interactions:

First, there is the Matthew effect, also known by the proverb "the rich get richer and the poor get poorer" can be observed, where a preferential attachment to larger nodes causing a power-law distribution of node degrees rather than a normal distribution which would be expected for any repeated random experiment with statistically independent trials. Due to this effect the average paper does not get the mean value of all citations, no, it gets close to zero. Most papers do not get more than 5 citations. But a few papers get cited a thousand times or even more often. Models assume that highly cited papers have a better chance of being cited in new papers can explain this behavior and predict a smooth power law distribution for paper citations.

However, to make the model accurate, there is another factor: time.

While the total number of citations for a given paper naturally can only increase over the years, the actual ability of papers to attract further citations dimishes over time - the paper "ages" (see Citation averages, 2000-2010). This applies even to classic papers, for example from Einstein or Hawking, which are no longer cited as they once were.

Matúš Medo and his colleagues from the University of Fribourg in Switzerland developed a model taking this aging factor into account. They found that a paper’s relevance decreases dramatically a few years after its publication.

Especially in our time of instant communication of results, it thus becomes very unlikely that a scientific paper gains in popularity after some time has passed. Sorry to crush your hopes, but if you have a meagerly cited paper now, it most likely won't become more popular in the future ;-)

 Links
  1. Matthew Effect. Wikipedia
  2. Matúš Medo, Giulio Cimini, and Stanislao Gualdi.Temporal Effects in the Growth of Networks. Phys. Rev. Lett. 107, 2011
  3. W. Elmenreich. Why is it important to get cited?. Self-Organizing Networked Systems Blog. October 2012
  4. Citation averages, 2000-2010, by fields and years. Times Higher Eduction 2011.

Sunday, October 14, 2012

Why is it important to get cited?

♫ "Hey! I just met you, and this is crazy,
      but here's my paper, so cite me, maybe?" ♫


(to be sung to the tune of "Call me maybe" by Carly Rae Jepsen, idea for text adaptation by Nikolaj Marchenko)

If this would work, you would find a lot of people singing that tune at conferences. The reason why it is important to get cited is because the number of citations of your publications have become an assessment of your scientific performance. The most famous indicator is the h-index [1], stating the largest number h for which there are at least h papers with at least h citations each.

20 years ago, a scientist was assessed by the number of papers she or he managed to write and publish. Getting a paper published was difficult, because there was limited space in the journals and each issue was a costly and time-consuming endeavor including typesetting, printing, distribution. In economic terms this means we had a shortage of a resource which made it valuable.

Today, things are better with regard to cost and effort for a publication - typsetting software is fast and easy to use and costs are lower than in the past. And when the Internet became the main medium instead of paper, printing costs vanished. If you like, you can found a new journal by just investing some time into the setup of a webpage template. Except from your own work time, personal costs would be no issue, since traditionally, being a journal's editor or reviewer is considered an honorary but unpaid job. This issues a quality assurance problem: If everybody can publish by themselves or provide an easy publication opportunity for others, the number of publications lose their status as a criterion for scientific quality and success. Therefore, attention has shifted to measure the actual impact of a publication in order to infer about its quality. The easy formula is: the more other works are influenced by a publication, the better this publication/work must have been.

This concept has its pros and cons. On the positive side, for at least all publications on the internet, the number of citations can be calculated automatically - Google scholar does it for you. Second, there is a good correlation between successful scientists and their number of citations. Negative aspects are that citations from publications which are not online are usually not included. There is also a bias depending on the scientific field, although there is work suggesting correction factors for this bias. The method further counts citations without distinction of the quality of the citation (be it positive, negative, long, brief, etc.) And finally, counting citations primarily measures the popularity of a paper, which explains why successful (popular) scientists have lots of citations. Still it appears that counting citations is currently the best way to assess publications with low effort. And it is a nice application of network theory.

Citations can be also used to assess journals. The more the publications in a journal are cited by others, the better is the journal. If everybody tries to get their papers published in journals with high impact, i.e. many citations, the competition leads to situation with a shortage on excellent publication places. Interesting is that a top journal does not require more effort than one with lower impact. The self-organizing effect of authors competing for the 'best' journals puts these journals in the convenient situation that they can pick the best papers - which in turn helps them in keeping their position. Regardless of the flaws of the citation-based impact analysis, as long as it is used by so many people, you have to play along.

Finally, some tips you might have been waiting for:
How can you push your h-index by maximizing the chance to get cited?
  • Make your publications available online (mine are here btw)
  • Discuss your work with others
  • Write good papers. Interesting comprehensive work is more likely to be cited.
  • Avoid low-impact journals and conferences
  • Publish in the language which is most common for your field of research. In most cases this is english. 
  • Add your paper as reference to appropriate pages in social networks (e.g., Wikipedia)
Note that these tips in general are part of serious research work. They make sense either you believe in the h-index or not. Don't try to fake you h-index, e.g. by massively citing yourself. Self-citations are likely to be excluded in future h-index calculations. Technically, excluding self-citations will be easy to implement by Google Scholar and co.

Links
  1. h-index. Wikipedia
  2. Google Scholar citation count (took myself as example)
  3. J. E. Iglesias and C. Pecharromán. Scaling the h-index for different scientific ISI fields. Scientometrics, Vol. 73, No. 3 (2007)
  4. W. Elmenreich. Google Scholar, Citation Indices, and the University of Klagenfurt. TEWI-Blog. November 2011

Thursday, November 11, 2010

A self-organizing algorithm for video distribution networks

The way how users consume videos has changed with the availability of large repositories with a high number of more or less related videos. Many users are only interested in tiny fractions of a video and not even necessarily in the original temporal order. Moreover, they might wish to dynamically compose portions of di erent videos into one presentation.
For example, take a video recording of a ski-jumping competition. Some users might be interested in watching it sequentially. A trainer might be interested in studying the jumping-off technique of athletes in parallel. Another user might be interested in the performance of several jumpers from one country.
In order to keep up with this emergent access patterns, we invented a self-organizing video delivery network that is based on artificial hormones which are spread throughout the network when a particular video is requested. The hormone spreading is affected by the bandwidth and delay parameters of the network edges, thus indirectly help in searching for the (currently) best path to transmit a video.
The interactions between nodes like spreading/evaporating hormone or moving a video according to the neighbor with highest hormone gradient are all local within a node's neighborhood. Still, the system is able
to guide the overall transportation and placement of units in the system up to near optimum.