idw – Informationsdienst Wissenschaft

Nachrichten, Termine, Experten

Grafik: idw-Logo
Grafik: idw-Logo

idw - Informationsdienst
Wissenschaft

Science Video Project
idw-Abo

idw-News App:

AppStore

Google Play Store



Instanz:
Teilen: 
13.09.2010 10:51

Computer Scientists from Saarbrücken Search Through Large Datasets using "Good Trojans"

Friederike Meyer zu Tittingdorf Pressestelle der Universität des Saarlandes
Universität des Saarlandes

    Social networks, search engines, digital archives, and several global-scale internet companies host very large data collections. In order to search through this data companies like Facebook, Ebay, Yahoo, and Twitter use the freely available Hadoop software - a variant of MapReduce originally proposed by Google. Several database experts, however, criticize Hadoop for being inefficient. Computer scientists from Saarland University (Germany) are now proposing a new system coined Hadoop++. It allows users to search through very large datasets much faster than before. Hadoop++ improves over Hadoop by up to a factor of 20.

    Internet companies need to process data volumes on the order of millions of Gigabytes (Petabytes) on a daily basis. In order to search through this data effectively, Google proposed the MapReduce programming model. MapReduce divides the input data into smaller chunks that are then distributed over a large network of machines and processed in parallel. The open-source counterpart of MapReduce is called Hadoop. Even though Google holds a patent on MapReduce, it granted a license to Hadoop. Therefore, Hadoop may still be used by companies free of charge. However, "Database experts who are fluent in SQL consider MapReduce a major step backwards towards the database stone age", explains Jens Dittrich, Professor of Information Systems at Saarland University. He adds, "MapReduce disregards considerable wisdom from database research. As a consequence MapReduce is often slow and inefficient."

    Despite this criticism, over the past years Hadoop has gained considerable attention from both industry and academia. Hadoop is very popular among programmers throughout the world. Professor Dittrich explains why, "The reason is its ease of use: the user neither has to learn a complex database language nor a data model. Furthermore, Hadoop is relatively easy to administrate. In summary, Hadoop allows even database-illiterate people to search through billions of records on very large computer clusters."

    However, this comes at a price: "When compared to modern relational database management systems, Hadoop is just too slow." Therefore the researcher and his team at Saarland University have developed a new system, coined Hadoop++. It aims to eliminate the performance deficiencies encountered in Hadoop. The creativity of the new approach lies in how the problem is tackled: Hadoop++ works similarly to a "trojan", i.e. a computer virus which infects a computer system clandestinely and may then cause considerable harm. In contrast to these bad trojans, Hadoop++ injects hidden code into a system in order to "heal" it, ie. dynamically accelerating the underlying Hadoop. Professor Dittrich emphasizes, "One could say that Hadoop++ is a good trojan".

    The Saarbrücken approach has the considerable advantage that Hadoop's tested code base does not have to be modified and retested. Thus, Hadoop++ avoids complex changes to a working system and unforeseen consequences. Hadoop++ will be presented at this year's International Conference on Very Large Databases (VLDB) - one of the world's most prestigious database conferences to be held in Singapore from September 13-17.

    Background

    There have been heated discussions about the pros and cons of MapReduce/Hadoop when compared to traditional database management systems. This discussion was led by database professors in the US. Please refer to the links below for details. Recently there have been attempts to improve the runtime efficiency of MapReduce. However, as Professor Dittrich explains, "those attempts could not really marry MapReduce with database technology."
    http://www.heise.de/newsticker/meldung/HadoopDB-versoehnt-SQL-mit-Map-Reduce-668...
    http://databasecolumn.vertica.com/database-innovation/mapreduce-a-major-step-bac...

    Press Pictures: www.uni-saarland.de/pressefotos

    For questions, contact:
    Jens Dittrich
    Professor of Information Systems
    Saarland University
    Tel. (+49) 681 302 70141


    Weitere Informationen:

    http://infosys.cs.uni-saarland.de/hadoop++.php
    http://hadoop.apache.org/
    http://www.mapreduce.org/
    http://www.vldb2010.org


    Bilder

    Social networks and search engines host very large data collections. Computer scientists of Saarland University are able to search through very large datasets much faster than before.
    Social networks and search engines host very large data collections. Computer scientists of Saarland ...
    bellhäuser - das bilderwerk
    None

    Jens Dittrich, Professor for Information Systems
    Jens Dittrich, Professor for Information Systems
    Universität des Saarlandes
    None


    Merkmale dieser Pressemitteilung:
    Informationstechnik
    überregional
    Forschungsergebnisse, Forschungsprojekte
    Englisch


     

    Hilfe

    Die Suche / Erweiterte Suche im idw-Archiv
    Verknüpfungen

    Sie können Suchbegriffe mit und, oder und / oder nicht verknüpfen, z. B. Philo nicht logie.

    Klammern

    Verknüpfungen können Sie mit Klammern voneinander trennen, z. B. (Philo nicht logie) oder (Psycho und logie).

    Wortgruppen

    Zusammenhängende Worte werden als Wortgruppe gesucht, wenn Sie sie in Anführungsstriche setzen, z. B. „Bundesrepublik Deutschland“.

    Auswahlkriterien

    Die Erweiterte Suche können Sie auch nutzen, ohne Suchbegriffe einzugeben. Sie orientiert sich dann an den Kriterien, die Sie ausgewählt haben (z. B. nach dem Land oder dem Sachgebiet).

    Haben Sie in einer Kategorie kein Kriterium ausgewählt, wird die gesamte Kategorie durchsucht (z.B. alle Sachgebiete oder alle Länder).