Zhengdeng Lei, PhD

Zhengdeng Lei, PhD

2009 - Present Research Fellow at Duke-NUS, Singapore
2007 - 2009 High Throughput Computational Analyst, Memorial Sloan-Kettering Cancer Center, New York
2003 - 2007 PhD, Bioinformatics, University of Illinois at Chicago

Thursday, June 14, 2012

Overlap of gene signatures of GC and PDA

1. Inv overlap signature_GC and signature_PDA, then DAVID/Gather


KEGG Pathway# Genesp ValueBayes Factor
1.path:hsa04010: MAPK signaling pathway19[show]      0.0051

2.path:hsa04020: Calcium signaling pathway16[show]      0.0071

3.path:hsa04210: Apoptosis10[show]     0.0071

4.path:hsa04610: Complement and coagulation cascades8[show]     0.010




Gene Ontology# Genesp ValueBayes Factor
1.GO:0007154 [3]: cell communication156[show]          < 0.000115
2.GO:0007275 [2]: development88[show]          < 0.00019
3.GO:0007165 [4]: signal transduction123[show]          < 0.00018
4.GO:0009653 [3]: morphogenesis62[show]          < 0.00018
5.GO:0009887 [4]: organogenesis52[show]          < 0.00017
6.GO:0048513 [3]: organ development52[show]          < 0.00017
7.GO:0006956 [6]: complement activation8[show]         0.00025
8.GO:0007242 [5]: intracellular signaling cascade50[show]        0.00064


2. Pro


KEGG Pathway# Genesp ValueBayes Factor
1.path:hsa04110: Cell cycle16[show]          < 0.000111
2.path:hsa04080: Neuroactive ligand-receptor interaction2[show]       0.0012
3.path:hsa00100: Biosynthesis of steroids4[show]     0.0071



Gene Ontology# Genesp ValueBayes Factor
1.GO:0000278 [6]: mitotic cell cycle35[show]          < 0.000136
2.GO:0007067 [8]: mitosis29[show]          < 0.000135
3.GO:0000087 [7]: M phase of mitotic cell cycle29[show]          < 0.000135
4.GO:0000279 [6]: M phase32[show]          < 0.000133
5.GO:0000280 [7]: nuclear division31[show]          < 0.000133
6.GO:0007049 [5]: cell cycle62[show]          < 0.000130
7.GO:0008283 [4]: cell proliferation73[show]          < 0.000124
8.GO:0000910 [5]: cytokinesis18[show]          < 0.000114


3.Meta


KEGG Pathway# Genesp ValueBayes Factor
1.path:hsa00480: Glutathione metabolism3[show]         0.00025
2.path:hsa00260: Glycine, serine and threonine metabolism2[show]      0.0032
3.path:hsa00791: Atrazine degradation1[show]      0.0041
4.path:hsa00460: Cyanoamino acid metabolism1[show]     0.0071
5.path:hsa00430: Taurine and hypotaurine metabolism1[show]     0.0091
6.path:hsa00625: Tetrachloroethene degradation1[show]     0.0091
7.path:hsa00363: Bisphenol A degradation1[show]     0.010




Gene Ontology# Genesp ValueBayes Factor
1.GO:0007586 [4]: digestion3[show]       0.0013
2.GO:0008203 [7]: cholesterol metabolism3[show]       0.0013
3.GO:0016125 [6]: sterol metabolism3[show]       0.0013
4.GO:0006848 [8]: pyruvate transport1[show]      0.0052
5.GO:0007039 [7]: vacuolar protein catabolism1[show]      0.0052
6.GO:0007165 [4]: signal transduction3[show]     0.011
7.GO:0045494 [8]: photoreceptor maintenance1[show]     0.011
8.GO:0006621 [6]: protein-ER retention1[show]     0.011




Monday, June 11, 2012

survival rate of gastric cancer


Statistics
This year, an estimated 21,320 adults (13,020 men and 8,300 women) in the United States will be diagnosed with stomach cancer. It is estimated that 10,540 deaths (6,190 men and 4,350 women) from this disease will occur this year.
The incidence of stomach cancer varies in different parts of the world. Although it is decreasing in the Western world, it is still one of the most common cancer types worldwide.
The five-year survival rate (percentage of people who survive at least five years after the cancer is detected, excluding those who die from other diseases) of people with stomach cancer is about 26%. This statistic reflects the fact that most people with stomach cancer are diagnosed after the cancer has already spread to other parts of the body. If stomach cancer is found before it has spread, the five-year survival rate is generally higher but depends on the stage of the cancer found during surgery.
Cancer survival statistics should be interpreted with caution. These estimates are based on data from thousands of people with this type of cancer in the United States each year, but the actual risk for a particular individual may differ. It is not possible to tell a person how long he or she will live with stomach cancer. Because the survival statistics are measured in five-year intervals, they may not represent advances made in the treatment or diagnosis of this cancer. Learn more about understanding statistics.
Statistics adapted from the American Cancer Society's publication, Cancer Facts & Figures 2012.

http://www.cancer.net/patient/Cancer+Types/Stomach+Cancer/ci.Stomach+Cancer.printer

Friday, June 8, 2012

read cel file



library(affxparser)
celfile = "E:\\CEL\\BreastCancer\\GSE28821_LCM\\OneBatch\\GSM713753.CEL"
dat.header <- readCelHeader(celfile)$datheader
split.header <- strsplit(dat.header, " ")[[1]]
scan.date <- grep("\\d+\\/\\d+\\/\\d+", split.header, perl=T, value = T)
scan.time <- grep("\\d+:\\d+:\\d+", split.header, perl=T, value = T)

Thursday, June 7, 2012

Wednesday, June 6, 2012

Principal Investigator@Ontario Institute for Cancer Research

http://bioinformatics.ca/resources/jobs/principal-investigator


Institution/Company: 
Ontario Institute for Cancer Research
Location: 
Downtown Toronto
Job Description: 
Position: Principal Investigator
Site: MaRS Centre, Toronto
Department: Informatics & Bio-computing
Reports To: Director, Informatics & Bio-computing Platform
Salary: Commensurate with level of experience
Hours: 35 Hrs/week
Status: Full-time, Permanent
The Ontario Institute for Cancer Research (OICR) is seeking Junior, Intermediate and Senior Principal Investigators (PIs) in Bioinformatics, Computational Biology and Biostatistics to undertake world-class computational research in a wide range of research areas, including any of the following: (1) discovery of key genetic alterations in the initiation or progression of cancer; (2) identification of biomarkers indicative of tumour subtypes or predictive of response to targeted therapy; (3) modeling of regulatory networks relevant to disease pathways; (4) analysis of genetic and environmental risk factors for cancer in patient populations; (5) use of machine learning and/or biostatistical approaches to develop novel algorithms and computational techniques for genomic and/or epigenomic data; (6) development of interoperability standards for exchanging and collaboratively annotating genome-scale data sets; or (7) development of software engineering techniques for managing and manipulating genome-scale datasets and complex analytic workflows. We expect to appoint up to five PIs over the period 2012-2014.
PIs will be expected to mentor trainees, and to build collaborations both within and outside the OICR community. In addition to base funds provided by the Institute to support the PI's salary and personnel, PIs are expected to raise additional research funds from external competitive granting agencies. The OICR will assist PIs in obtaining faculty appointments at the University of Toronto or another affiliated academic institution.
QUALIFICATIONS
• An MD or PhD with a proven track record in computational biology, bioinformatics, or biostatistics;
• For new PIs, a record of independent research and either first-author peer reviewed publications or the publication of software, databases or other significant community resources.
• For senior and intermediate-level PIs, international recognition and a strong publication record of relevance, proven leadership and management experience including the building of strong research teams, as well as a strong record of mentorship and/or teaching;
• Eligible to hold the rank of assistant, associate or full professor at an Ontario university;
• Excellent communication and presentation skills.
OICR is an innovative cancer research institute located in the MaRS Centre in the Discovery District in downtown Toronto. OICR is addressing significant challenges in cancer research with multi-disciplinary, multi-institutional teams. New discoveries to prevent, detect and treat cancer will be moved from the bench to practical applications in patients. The OICR team is growing quickly. We are innovative, dedicated professionals who bring expertise to each of our roles. We are looking for individuals interested in being part of a culture of excellence that will result in Ontario being recognized internationally as a leading jurisdiction for cancer research.
Launched in December 2005, OICR is an independent institute funded by the Government of Ontario through the Ministry of Economic Development and Innovation.
For more information about OICR, please visit the website at www.oicr.on.ca.
POSTED DATE: June 1, 2012
CLOSING DATE: Posted until filled
Interested candidates may apply here
https://www.recruitingsite.com/csbsites/oicr/JobDescription.asp?JobNumber=675388
OICR has a diverse workforce and is an equal opportunity employer.
The Ontario Institute for Cancer Research thanks all applicants. However, only those under consideration will be contacted. Candidates will be expected to provide their current employer as a reference.
Resume Format: If you elect to apply, you will need a text or HTML version of your resume so that you can cut and paste it into the application box provided. Before you submit the completed application, you will be asked to attach one or two files to your application. Please attach your resume as a .doc file.

try

1. use microdissected BC-class to predict NCI60/GEMINI
ANSWER: looks like not good.

2. use macrodissected BC-class to predict BC57 with some dropped samples (to be population balance)
Answer: tried, but not good result, so it is not because of population bias.

Tuesday, June 5, 2012

Normalization before NTP

Before NTP, is it better to do standardization on gene(row) then on array(column) to avoid population bias.
Or may median polish (an iterative method)?

ANSWER: It is NOT GOOD to do standardization on gene(row) then on array(column) to avoid population bias.


################################################
# standardization by row and column, respetively
#By row (gene)
std.data.by.row <- t(scale(t(data.filtered), scale=T))
#By column (array)
std.data <- scale(std.data.by.row, scale=T)