Back
Data mining of high content screening data
Together with my classmate Oscar we conducted our master thesis at the Karolinska Institute where we made a web application to enable load and analyze the test result data from high content screening. High content screening is a method within cell biology where cells are exposed for different substances, photos with image processing extracts data about the cells behaviors. The screening took pictures of up to hundreds of test tubes which followed by image processing resulted in a csv file of data could conclude of millions of lines of data with multiple columns (Filesize > 3GB). This amount of data was impossible for a human to analyze manually, so data mining came to the rescue in an attempt to solve this issue. Our mission to provide a software which could handle such an amount of data and evaluate different algorithms that could provide the scientists with useful findings within their screening.








