Genomics Data Mining and Statistics

SPH BS 831

The goal of this course is for the students to develop a good understanding and hands-on skills in the design and analysis of data from microarray and high-throughput sequencing experiments, including data collection and management, statistical techniques for the identification of genes that have differential expression in different biological conditions, development of prognostic and diagnostic models for molecular classification, and the identification of new disease taxonomies based on their molecular profile. These topics will be taught using real examples, extensively documented hands-on's, class discussion and critical reading. Students will be asked to analyze real gene expression data sets in their homeworks and final project. Principles of reproducible research will be emphasized, and students will become proficient in the use of the statistical language R (an advanced beginners knowledge of the language is expected of the students entering the class) and associated packages (including Bioconductor), and in the use of R markdown (and/or electronic notebooks) for the redaction of analysis reports.