Healthcare analytics, AI solutions for biological big data, providing an AI platform for the biotech, life sciences, medical and pharmaceutical industries, as well as for related technological approaches, i.e., curation and text analysis with machine learning and other activities related to AI applications to these industries.
Crowdsourcing Genetic Data Yields Discovery of DNA loci associated with Major Depressive Disorder (MDD) in European Descendants, Volume 2 (Volume Two: Latest in Genomics Methodologies for Therapeutics: Gene Editing, NGS and BioInformatics, Simulations and the Genome Ontology), Part 1: Next Generation Sequencing (NGS)
Crowdsourcing Genetic Data Yields Discovery of DNA loci associated with Major Depressive Disorder (MDD) in European Descendants
Reporter: Kelly Perlman, Life Sciences Student and Research Assistant, McGill University
UPDATED on 11/24/2019
Can AI help diagnose depression? It’s a long shot
At the moment, machine intelligence is just as subjective as human intelligence
Researchers from Pfizer Global Research and Development, 23andMe, and the Massachusetts General Hospital have published a study in Nature Genetics, pinpointing 15 genetic loci associated with the risk of developing major depressive disorder (MDD) in individuals of European ancestry. Evidence from previous research suggests that MDD is heritable, but the details of the specific gene correlates are unclear. The identification of loci where single nucleotide polymorphisms (SNPs) related to MDD exist could provide better insight into the neurobiology of depression, and therefore better treatment options.
23andMe, a private biotechnology company situated in California, offers a DNA sequencing service in which consumers send in a saliva swab for testing, and later receive a report listing the findings of the analysis related to ancestry, physical and behavioral traits, along with risk of inheriting certain diseases. The participants of this study had agreed to provide the results of their genetic testing for scientific research.
The results of 75,607 participants with self-reported diagnoses of depression were compared to the results of 231,747 participants reporting having never experienced depression. This data was combined with the results of previously published MDD genome-wide association studies (GWAS). To test the whether these results could be replicated, another set of results from 23andMe was analyzed, in which there were 45,773 MDD subjects, and 106,354 controls.
After the joint analysis, 17 SNPs were identified at 15 different loci. Tissue and gene enrichment assays showed that the genes that were over-expressed in the CNS were related to functions including neurodevelopment, histone methylation, neurogenesis and synaptic modification.
The team then created a weighted genetic risk score (GRS) in which they compared the 17 SNPs with factors including medication use, comorbid diseases and behavioral phenotypes, all of which were correlated with the GRS. Of note, the GRS was very highly correlated with age of onset of MDD.
The crowdsourcing of genetic data proves to be an efficient and powerful tool for large-scale MDD studies. Pooling large subject databases together is essential in order to account for the heterogeneous nature of the disease. Despite not being able to precisely assess each subject’s disease phenotype, scientists can make more rapid headway by collaborating with biotechnology companies in the quest to better understand the biological mechanisms of depression. Ron Perlis, M.D., M.Sc., of the Massachusetts General Hospital and co-author of this paper explained that “finding genes associated with depression should help make clear that this is a brain disease, which we hope will decrease the stigma still associated with these kinds of illnesses”.
Hyde, C. L., Nagle, M. W., Tian, C., Chen, X., Paciga, S. A., Wendland, J. R., . . . Winslow, A. R. (2016). Identification of 15 genetic loci associated with risk of major depression in individuals of European descent. Nature Genetics Nat Genet. doi:10.1038/ng.3623
mRNA Data Survival Analysis, Volume 2 (Volume Two: Latest in Genomics Methodologies for Therapeutics: Gene Editing, NGS and BioInformatics, Simulations and the Genome Ontology), Part 1: Next Generation Sequencing (NGS)
mRNA Data Survival Analysis
Curators: Larry H. Bernstein, MD, FCAP and Aviva Lev-Ari, PhD, RN
SURVIV for survival analysis of mRNA isoform variation
The rapid accumulation of clinical RNA-seq data sets has provided the opportunity to associate mRNA isoform variations to clinical outcomes. Here we report a statistical method SURVIV (Survival analysis of mRNA Isoform Variation), designed for identifying mRNA isoform variation associated with patient survival time. A unique feature and major strength of SURVIV is that it models the measurement uncertainty of mRNA isoform ratio in RNA-seq data. Simulation studies suggest that SURVIV outperforms the conventional Cox regression survival analysis, especially for data sets with modest sequencing depth. We applied SURVIV to TCGA RNA-seq data of invasive ductal carcinoma as well as five additional cancer types. Alternative splicing-based survival predictors consistently outperform gene expression-based survival predictors, and the integration of clinical, gene expression and alternative splicing profiles leads to the best survival prediction. We anticipate that SURVIV will have broad utilities for analysing diverse types of mRNA isoform variation in large-scale clinical RNA-seq projects.
Eukaryotic cells generate remarkable regulatory and functional complexity from a finite set of genes. Production of mRNA isoforms through alternative processing and modification of RNA is essential for generating this complexity. A prevalent mechanism for producing mRNA isoforms is the alternative splicing of precursor mRNA1. Over 95% of the multi-exon human genes undergo alternative splicing2, 3, resulting in an enormous level of plasticity in the regulation of gene function and protein diversity. In the last decade, extensive genomic and functional studies have firmly established the critical role of alternative splicing in cancer4, 5, 6. Alternative splicing is involved in a full spectrum of oncogenic processes including cell proliferation, apoptosis, hypoxia, angiogenesis, immune escape and metastasis7, 8. These cancer-associated alternative splicing patterns are not merely the consequences of disrupted gene regulation in cancer but in numerous instances actively contribute to cancer development and progression. For example, alternative splicing of genes encoding the Bcl-2 family of apoptosis regulators generates both anti-apoptotic and pro-apoptotic protein isoforms9. Alternative splicing of the pyruvate kinase M (PKM) gene has a significant impact on cancer cell metabolism and tumour growth10. A transcriptome-wide switch of the alternative splicing programme during the epithelial–mesenchymal transition plays an important role in cancer cell invasion and metastasis11, 12.
RNA sequencing (RNA-seq) has become a popular and cost-effective technology to study transcriptome regulation and mRNA isoform variation13, 14. As the cost of RNA-seq continues to decline, it has been widely adopted in large-scale clinical transcriptome projects, especially for profiling transcriptome changes in cancer. For example, as of April 2015 The Cancer Genome Atlas (TCGA) consortium had generated RNA-seq data on over 11,000 cancer patient specimens from 34 different cancer types. Within the TCGA data, breast invasive carcinoma (BRCA) has the largest sample size of RNA-seq data covering over 1,000 patients, and clinical information such as survival times, tumour stages and histological subtypes is available for the majority of the BRCA patients15. Moreover, the median follow-up time of BRCA patients is ~400 days, and 25% of the patients have more than 1,200 days of follow-up. Collectively, the large sample size and long follow-up time of the TCGA BRCA data set allow us to correlate genomic and transcriptomic profiles to clinical outcomes and patient survival times.
To date, systematic analyses have been performed to reveal the association between copy number variation, DNA methylation, gene expression and microRNA expression profiles with cancer patient survival16, 17. By contrast, despite the importance of mRNA isoform variation and alternative splicing, there have been limited efforts in transcriptome-wide survival analysis of alternative splicing in cancer patients. Most RNA-seq studies of alternative splicing in cancer transcriptomes focus on identifying ‘cancer-specific’ alternative splicing events by comparing cancer tissues with normal controls (see refs 18, 19, 20, 21, 22, 23 for examples). A recent analysis of TCGA RNA-seq data identified 163 recurrent differential alternative splicing events between cancer and normal tissues of three cancer types, among which five were found to have suggestive survival signals for breast cancer at a nominal P-value cutoff of 0.05 (ref. 21). Some other studies reported a significant survival difference between cancer patient subgroups after stratifying patients with overall mRNA isoform expression profiles24, 25. However, systematic cancer survival analyses of alternative splicing at the individual exon resolution have been lacking. Two main challenges exist for survival analyses of mRNA isoform variation and alternative splicing using RNA-seq data. The first challenge is to account for the estimation uncertainty of mRNA isoform ratios inferred from RNA-seq read counts. The statistical confidence of mRNA isoform ratio estimation depends on the RNA-seq read coverage for the events of interest, with larger read coverage leading to a more reliable estimation14. Modelling the estimation uncertainty of mRNA isoform ratio is an essential component of RNA-seq analyses of alternative splicing, as shown by various statistical algorithms developed for detecting differential alternative splicing from multi-group RNA-seq data14, 26, 27, 28,29. The second challenge, which is a general issue in survival analysis, is to properly model the association of mRNA isoform ratio with survival time, while accounting for missing data in survival time because of censoring, that is, patients still alive at the end of the survival study, whose precise survival time would be uncertain. To date, no algorithm has been developed for survival analyses of mRNA isoform variation that accounts for these sources of uncertainty simultaneously.
Here we introduce SURVIV (Survival analysis of mRNA Isoform Variation), a statistical model for identifying mRNA isoform ratios associated with patient survival times in large-scale cancer RNA-seq data sets. SURVIV models the estimation uncertainty of mRNA isoform ratios in RNA-seq data and tests the survival effects of isoform variation in both censored and uncensored survival data. In simulation studies, SURVIV consistently outperforms the conventional Cox regression survival analysis that ignores the measurement uncertainty of mRNA isoform ratio. We used SURVIV to identify alternatively spliced exons whose exon-inclusion levels significantly correlated with the survival times of invasive ductal carcinoma (IDC) patients from the TCGA breast cancer cohort. Survival-associated alternative splicing events are identified in gene pathways associated with apoptosis, oxidative stress and DNA damage repair. Importantly, we show that alternative splicing-based survival predictors outperform gene expression-based survival predictors in the TCGA IDC RNA-seq data set, as well as in TCGA data of five additional cancer types. Moreover, the integration of clinical information, gene expression and alternative splicing profiles leads to the best prediction of survival time.
SURVIV statistical model
The statistical model of SURVIV assesses the association between mRNA isoform ratio and patient survival time. While the model is generic for many types of alternative isoform variation, here we use the exon-skipping type of alternative splicing to illustrate the model (Fig. 1a). For each alternative exon involved in exon-skipping, we can use the RNA-seq reads mapping to its exon-inclusion or -skipping isoform to estimate its exon-inclusion level (denoted as ψ, or PSI that is Per cent Spliced In14). A key feature of SURVIV is that it models the RNA-seq estimation uncertainty of exon-inclusion level as influenced by the sequencing coverage for the alternative splicing event of interest. This is a critical issue in accurate quantitative analyses of mRNA isoform ratio in large-scale RNA-seq data sets14, 26, 27, 28, 29. Therefore, SURVIV contains two major components: the first to model the association of mRNA isoform ratio with patient survival time across all patients, and the second to model the estimation uncertainty of mRNA isoform ratio in each individual patient (Fig. 1a).
Figure 1: The statistical framework of the SURVIV model.
(a) For each patient k, the patient’s hazard rate λk(t) is associated with the baseline hazard rate λ0(t) and this patient’s exon-inclusion level ψk. The association of exon-inclusion level with patient survival is estimated by the survival coefficient β. The exon-inclusion level ψk is estimated from the read counts for the exon-inclusion isoform ICk and the exon-skipping isoform SCk. The proportion of the inclusion and skipping reads is adjusted by a normalization function f that considers the lengths of the exon-inclusion and -skipping isoforms (see details in Results and Supplementary Methods). (b) A hypothetical example to illustrate the association of exon-inclusion level with patient survival probability over time Sk(t), with the survival coefficient β=−1 and a constant baseline hazard rate λ0(t)=1. In this example, patients with higher exon-inclusion levels have lower hazard rates and higher survival probabilities. (c) The schematic diagram of an exon-skipping event. The exon-inclusion reads ICk are the reads from the upstream splice junction, the alternative exon itself and the downstream splice junction. The exon-skipping reads SCk are the reads from the skipping splice junction that directly connects the upstream exon to the downstream exon.
Briefly, for any individual exon-skipping event, the first component of SURVIV uses a proportional hazards model to establish the relationship between patient k’s exon-inclusion level ψk and hazard rate λk(t).
For each exon, the association between the exon-inclusion level and patient survival time is reflected by the survival coefficient β. A positive β means increased exon inclusion is associated with higher hazard rate and poorer survival, while a negative β means increased exon inclusion is associated with lower hazard rate and better survival. λ0(t) is the baseline hazard rate estimated from the survival data of all patients (see Supplementary Methods for the detailed estimation procedure). A particular patient’s survival probability over time Sk(t) can be calculated from the patient-specific hazard rate λk(t) as . Figure 1b illustrates a simple example with a negative β=−1 and a constant baseline hazard rate λ0(t)=1, where higher exon-inclusion levels are associated with lower hazard rates and higher survival probabilities.
The second component of SURVIV models the exon-inclusion level and its estimation uncertainty in individual patient samples. As illustrated in Fig. 1c, the exon-inclusion level ψk of a given exon in a particular sample can be estimated by the RNA-seq read count specific to the exon inclusion isoform (ICk) and the exon-skipping isoform (SCk). Other types of alternative splicing and mRNA isoform variation can be similarly modelled by this framework29. Given the effective lengths (that is, the number of unique isoform-specific read positions) of the exon-inclusion isoform (lI) and the exon-skipping isoform (lS), the exon-inclusion level ψk can be estimated as . Assuming that the exon-inclusion read count ICk follows a binomial distribution with the total read count nk=ICk+SCk, we have:
The binomial distribution models the estimation uncertainty of ψk as influenced by the total read count nk, in which the parameter pk represents the proportion of reads from the exon-inclusion isoform, given the exon-inclusion level ψk adjusted by a length normalization function f(ψk) based on the effective lengths of the isoforms. The definitions of effective lengths for all basic types of alternative splicing patterns are described in ref. 29.
Distinct from conventional survival analyses in which predictors do not have estimation uncertainty, the predictors in SURVIV are exon-inclusion levels ψk estimated from RNA-seq count data, and the confidence of ψk estimate for a given exon in a particular sample depends on the RNA-seq read coverage. We use the statistical framework of survival measurement error model30 to incorporate the estimation uncertainty of isoform ratio in the proportional hazards model. Using a likelihood ratio test, we test whether the exon-inclusion levels have a significant association with patient survival over the null hypothesis H0:β=0. The false discovery rate (FDR) is estimated using the Benjamini and Hochberg approach31. Details of the parameter estimation and likelihood ratio test in SURVIV are described in Supplementary Methods.
Figure 2: Simulation studies to assess the performance of SURVIV and the importance of modelling the estimation uncertainty of mRNA isoform ratio.
We compared our SURVIV model with Cox regression using point estimates of exon-inclusion levels, which does not consider the estimation uncertainty of the mRNA isoform ratio. (a) To study the effect of RNA-seq depth, we simulated the mean total splice junction read counts equal to 5, 10, 20, 50, 80 and 100 reads. We generated two sets of simulations with and without data-censoring. For each simulation, the true-positive rate (TPR) at 5% false-positive rate is plotted. The inset figure shows the empirical distribution of the mean total splice junction read counts in the TCGA IDC RNA-seq data (x axis in the log10 scale). (b) To faithfully represent the read count distribution in a real data set, we performed another simulation with read counts directly sampled from the TCGA IDC data. Sampled read counts were then multiplied by different factors ranging from 10 to 300% to simulate data sets with different RNA-seq read depth. Continuous and dashed lines represent the performance of SURVIV and Cox regression, respectively. Red lines represent the area under curve (AUC) of the ROC curve (TPR versus false-positive rate plot). Black lines represent the TPR at 5% false-positive rate.
Using these simulated data, we compared SURVIV with Cox regression in two settings, without or with censoring of the survival time. In the setting without censoring, the death and survival time of each individual is known. In the setting with censoring, certain individuals are still alive at the end of the survival study. Consequently, these patients have unknown death and survival time. Here, in the simulation with censoring, we assumed that 85% of the patients were still alive at the end of the study, similar to the censoring rate of the TCGA IDC data set. In both settings and with different depths of RNA-seq coverage, SURVIV consistently outperformed Cox regression in the true-positive rate at the same false-positive rate of 5% (Fig. 2a). As expected, we observed a more significant improvement in SURVIV over Cox regression when the RNA-seq read coverage was low (Fig. 2a).
To more faithfully recapitulate the read count distribution in a real cancer RNA-seq data set, we performed another simulation study with read counts directly sampled from the TCGA IDC data. To assess the influence of RNA-seq read depth on the performance of SURVIV and Cox regression, sampled read counts were then multiplied by different factors ranging from 10 to 300% to simulate data sets with different RNA-seq read depths (Fig. 2b). The TCGA IDC data set has an average RNA-seq depth of ~60 million paired-end reads per patient. Thus, the read depth of these simulated RNA-seq data sets ranged from ~6 million reads to 180 million reads per patient, representing low-coverage RNA-seq studies designed primarily for gene expression analysis32 up to high-coverage RNA-seq studies designed primarily for alternative isoform analysis29. At all levels of RNA-seq depth, SURVIV consistently outperformed Cox regression, as reflected by the area under curve of the receiver operating characteristic (ROC) curve as well as the true-positive rate at 5% false-positive rate (Fig. 2b). The improvement of SURVIV over Cox regression was particularly prominent when the read depth was low. For example, at 10% read depth, SURVIV had 7% improvement in area under curve (68% versus 61%) and 8% improvement in the true-positive rate at 5% false-positive rate (46% versus 38%). Collectively, these simulation results suggest that SURVIV achieves a higher accuracy by accounting for the estimation uncertainty of mRNA isoform ratio in RNA-seq data.
SURVIV analysis of TCGA IDC breast cancer data
To illustrate the practical utility of SURVIV, we used it to analyse the overall survival time of 682 IDC patients from the TCGA breast cancer (BRCA) RNA-seq data set (see Methods for details of the data source and processing pipeline). We chose to analyse IDC because it is the most frequent type of breast cancer33, comprising ~70% of patients in the TCGA breast cancer data set. To control for the effects of significant clinical parameters such as tumour stage and subtype and identify alternative splicing events associated with patient outcomes across multiple molecular and clinical subtypes, we followed the procedure of Croce and colleagues in analysing mRNA and microRNA prognostic signature of IDC33 and stratified the patients according to their clinical parameters. We then conducted SURVIV analysis in 26 clinical subgroups with at least 50 patients in each subgroup. We identified 229 exon-skipping events associated with patient survival in multiple clinical subgroups that met the criteria of SURVIV P-value≤0.01 in at least two subgroups of the same clinical parameter (cancer subtype, stage, lymph node, metastasis, tumour size, oestrogen receptor status, progesterone receptor status, HER2 status and age as shown in Fig. 3). DAVID (Database for Annotation, Visualization and Integrated Discovery) Gene Ontology analyses34 of the 229 alternative splicing events suggest an enrichment of genes in cancer-related functional categories such as intracellular signalling, apoptosis, oxidative stress and response to DNA damage (Supplementary Fig. 1). Table 1 shows a few selected examples of survival-associated alternative splicing events in cancer-related genes. Using two-means clustering of each individual exon’s inclusion levels, the 682 IDC patients can be segregated into two subgroups with significantly different survival times as illustrated by the Kaplan–Meier survival plot (Fig. 4). We also carried out hierarchical clustering of IDC patients using 176 survival-associated alternative exons (P≤0.01; SURVIV analysis of all IDC patients). Using the exon-inclusion levels of these 176 exons, we clustered IDC patients into three major subgroups, with 95, 194 and 389 patients, respectively. As illustrated by the Kaplan–Meier survival plots, the three subgroups had significantly different survival times (Supplementary Fig. 2).
Figure 3: SURVIV analysis of exon-skipping events in the TCGA IDC RNA-seq data set.
IDC patients are stratified into multiple clinical subgroups based on clinical parameters including cancer subtype, stage, lymph node status, metastasis, tumour size, oestrogen receptor status, progesterone receptor status, HER2 status and age. Only clinical subgroups with at least 50 patients are included in further analyses. Numbers of patients in the subgroups are indicated next to the names of the subgroups. Shown in the heatmap are the log10 SURVIV P-values of the 229 exons associated with patient survival (P≤0.01) in at least two subgroups of the same class of clinical parameters. Turquoise colour indicates positive correlation that higher exon-inclusion levels are associated with higher survival probabilities. Magenta colour indicates negative correlation that lower exon-inclusion levels are associated with higher survival probabilities.
Figure 4: Kaplan–Meier survival plots of IDC patients stratified by two-means clustering of the exon-inclusion levels of four survival-associated alternative splicing events.
Clustering was generated for each of the four exons separately. Black lines represent patients with high exon-inclusion levels. Red lines represent patients with low exon-inclusion levels. The P-values are from SURVIV analysis of the TCGA IDC RNA-seq data. (a) ATRIP. (b) BCL2L11. (c) CD74. (d) PCBP4.
Figure 5: Alternative splicing of STAT5A exon 5 is significantly associated with IDC patient survival.
(a) The gene structure of the STAT5A full-length isoform compared to the ΔEx5 isoform skipping the 5th exon. (b) Kaplan–Meier survival plot of IDC patients stratified by two-means clustering using exon-inclusion levels of STAT5A exon 5. The 420 patients in Group 1 (average exon 5 inclusion level=95%) have significantly higher survival probabilities than the 262 patients in Group 2 (average exon 5 inclusion level=85%) (SURVIV P=6.8e−4). (c) Exon 5 inclusion levels of IDC patients stratified by two-means clustering using exon 5 inclusion levels. Group 1 has 420 patients with average exon-inclusion level at 95%. Group 2 has 262 patients with average exon-inclusion level at 85%. (d) STAT5A exon 5 inclusion levels in normal breast tissues versus breast cancer tumour samples. Exon-inclusion levels are extracted from 86 TCGA breast cancer patients with matched normal and tumour samples. Normal breast tissues have average exon 5 inclusion level at 95%, compared to 91% average exon-inclusion level in tumour samples. Error bars represent 95% confidence interval of the mean.
Figure 6: Splicing factor regulatory network of survival-associated alternative splicing events in IDC.
(a–c) Kaplan–Meier survival plots of IDC patients stratified by the gene expression levels of three splicing factors: TRA2B (a, Cox regression P=1.8e−4), HNRNPH1 (b, P=3.4e−4) and SFRS3 (c, P=2.8e−3). Black lines represent patients with high gene expression levels. Red lines represent patients with low gene expression levels. (d) The exon-inclusion levels of a DHX30 alternative exon are negatively correlated with TRA2B gene expression levels (robust correlation coefficient r=−0.26, correlation P=1.2e−17). (e) The exon-inclusion levels of a MAP3K4 alternative exon are positively correlated withHNRNPH1 gene expression levels (robust correlation coefficient r=0.16, correlation P=2.6e−06). (f) A splicing co-expression network of the three splicing factors and their correlated survival-associated alternative exons. In total, 84 survival-associated alternative exons are significantly correlated with the three splicing factors. The positive/negative correlation between splicing factors and alternative exons is represented by blue/red lines, respectively. Exons whose inclusion levels are positively/negatively correlated with survival times are represented by blue/red dots, respectively. The size of the splicing factor circles is proportional to the number of correlated exons within the network.
Figure 7: Cross-validation of different classes of IDC survival predictors measured by the C-index
A C-index of 1 indicates perfect prediction accuracy and a C-index of 0.5 indicates random guess. The plots indicate the distribution of C-indexes from 100 rounds of cross-validation. The centre value of the box plot is the median C-index from 100 rounds of cross-validation. The notch represents the 95%confidence interval of the median. The box represents the 25 and 75% quantiles. The whiskers extended out from the box represent the 5 and 95% quantiles. Two-sided Wilcoxon test was used to compare different survival predictors. The different classes of predictors are: (a) clinical information (median C-index 0.67). (b) Gene expression (median C-index 0.68). (c) Alternative splicing (median C-index 0.71). (d) Clinical information+gene expression (median C-index 0.69). (e) Clinical information+alternative splicing (median C-index 0.73). (f) Clinical information+gene expression+alternative splicing (median C-index 0.74). Note that ‘Gene’ refers to ‘Gene-level expression’ in these plots.
Next, we carried out the SURVIV analysis in five additional cancer types in TCGA, including GBM (glioblastoma multiforme), KIRC (kidney renal clear cell carcinoma), LGG (lower grade glioma), LUSC (lung squamous cell carcinoma) and OV (ovarian serous cystadenocarcinoma). As expected, the number of significant events at different FDR or P-value significance cutoffs varied across cancer types, with LGG having the strongest survival-associated alternative splicing signals with 660 significant exon-skipping events at FDR≤5% (Supplementary Data 3 and 4). Strikingly, regardless of the number of significant events, alternative splicing-based survival predictors outperformed gene expression-based survival predictors across all cancer types (Supplementary Fig. 3), consistent with our initial observation on the IDC data set.
Alternative processing and modification of mRNA, such as alternative splicing, allow cells to generate a large number of mRNA and protein isoforms with diverse regulatory and functional properties. The plasticity of alternative splicing is often exploited by cancer cells to produce isoform switches that promote cancer cell survival, proliferation and metastasis7, 8. The widespread use of RNA-seq in cancer transcriptome studies15, 47, 48 has provided the opportunity to comprehensively elucidate the landscape of alternative splicing in cancer tissues. While existing studies of alternative splicing in large-scale cancer transcriptome data largely focused on the comparison of splicing patterns between cancer and normal tissues or between different subtypes of cancer18, 21, 49, additional computational tools are needed to characterize the clinical relevance of alternative splicing using massive RNA-seq data sets, including the association of alternative splicing with phenotypes and patient outcomes.
We have developed SURVIV, a novel statistical model for survival analysis of alternative isoform variation using cancer RNA-seq data. SURVIV uses a survival measurement error model to simultaneously model the estimation uncertainty of mRNA isoform ratio in individual patients and the association of mRNA isoform ratio with survival time across patients. Compared with the conventional Cox regression model that uses each patient’s mRNA isoform ratio as a point estimate, SURVIV achieves a higher accuracy as indicated by simulation studies under a variety of settings. Of note, we observed a particularly marked improvement of SURVIV over Cox regression for low- and moderate-depth RNA-seq data (Fig. 2b). This has important practical value because many clinical RNA-seq data sets have large sample size but relatively modest sequencing depth.
Using the TCGA IDC breast cancer RNA-seq data of 682 patients, SURVIV identified 229 alternative splicing events associated with patient survival time, which met the criteria of SURVIVP-values≤0.01 in multiple clinical subgroups. While the statistical threshold seemed loose, several lines of evidence suggest the functional and clinical relevance of these survival-associated alternative splicing events. These alternative splicing events were frequently identified and enriched in the gene functional groups important for cancer development and progression, including apoptosis, DNA damage response and oxidative stress. While some of these events may simply reflect correlation but not causal effect on cancer patient survival, other events may play an active role in regulating cancer cell phenotypes. For example, a survival-associated alternative splicing event involving exon 5 of STAT5A is known to regulate the activity of this transcription factor with important roles in epithelial cell growth and apoptosis37. Using a co-expression network analysis of splicing factor to exon correlation across all patients, we identified three splicing factors (TRA2B, HNRNPH1 and SFRS3) as potential hubs of the survival-associated alternative splicing network of IDC. The expression levels of all three splicing factors were negatively associated with patient survival times (Fig. 6a–c), and both TRA2B and HNRNPH1 were previously reported to have an impact on cancer-related molecular pathways40, 41, 42, 43, 44, 45. Finally, despite the limited power in detecting individual events, we show that the survival-associated alternative splicing events can be used to construct a predictor for patient survival, with an accuracy higher than predictors based on clinical parameters or gene expression profiles (Fig. 7). This further demonstrates the potential biological relevance and clinical utility of the identified alternative splicing events.
We performed cross-validation analyses to evaluate and compare the prognostic value of alternative splicing, gene expression and clinical information for predicting patient survival, either independently or in combination. As expected, the combined use of all three types of information led to the best prediction accuracy. Because we used penalized regression to build the prediction model, combining information from multiple layers of data did not necessarily increase the number of predictors in the model. The perhaps more surprising and intriguing result is that alternative splicing-based predictors appear to outperform gene expression-based predictors when used alone and when either type of data was combined with clinical information (Fig. 7). We observed the same trend in five additional cancer types (Supplementary Fig. 3). We note that this finding was consistent with a previous report that cancer subtype classification based on splicing isoform expression performed better than gene expression-based classification25. While this trend seems counterintuitive because accurate estimation of gene expression requires much lower RNA-seq depth than accurate estimation of alternative splicing29, one possible explanation may be the inherent characteristic of isoform ratio data. By definition, mRNA isoform ratio is estimated as the ratio of multiple mRNA isoforms from a single gene. Therefore, mRNA isoform ratio data have a ‘built-in’ internal control that could be more robust against certain artefacts and confounding issues that influence gene expression estimates across large clinical RNA-seq data sets, such as poor sample quality and RNA degradation12. Regardless of the reasons, our data call for further studies to fully explore the utility of mRNA isoform ratio data for various clinical research applications.
The SURVIV source code is available for download at https://github.com/Xinglab/SURVIV. SURVIV is a general statistical model for survival analysis of mRNA isoform ratio using RNA-seq data. The current statistical framework of SURVIV is applicable to RNA-seq based count data for all basic types of alternative splicing patterns involving two isoform choices from an alternatively spliced region, such as exon-skipping, alternative 5′ splice sites, alternative 3′ splice sites, mutually exclusive exons and retained introns, as well as other forms of alternative isoform variation such as RNA editing. With the rapid accumulation of clinical RNA-seq data sets, SURVIV will be a useful tool for elucidating the clinical relevance and potential functional significance of alternative isoform variation in cancer and other diseases.
A Genetic Switch to Control Female Sexual Behavior, Volume 2 (Volume Two: Latest in Genomics Methodologies for Therapeutics: Gene Editing, NGS and BioInformatics, Simulations and the Genome Ontology), Part 2: CRISPR for Gene Editing and DNA Repair
Reporter and Curator: Dr. Sudipta Saha, Ph.D.
In an African cichlid fish, Astatotilapia burtoni, fertile females select a mate and perform a stereotyped spawning / mating routine, offering quantifiable behavioral outputs of neural circuits. A male fish attracts a fertile female by rapidly quivering his brightly colored body. If she chooses him, he guides her back to his territory, where he quivers some more as she pecks at fish egg–colored spots on his anal fin. Next, she lays eggs and quickly scoops them up in her mouth. With a mouthful of eggs, she continues pecking at the male’s spots, “believing” them to be eggs to be collected. As she does, he releases sperm from near his anal fin, which she also gathers. This fertilizes the eggs, and she carries the embryos in her mouth for two weeks as they develop.
But, the question was how these females can time their reproduction to coincide with when they are fertile. The female fish will not approach or choose males until they are ready to reproduce, so there must be something in their brains that signals when sexual behavior will be required. The scientists began by considering signaling molecules previously associated with sexual behavior and reproduction, and showed that PGF2α injection activates a naturalistic pattern of sexual behavior in female Astatotilapia burtoni. They would engage in mating behavior even if they were non-fertile, doing the quiver dance with males, but wouldn’t actually lay eggs since they had none.
The scientists also identified cells in the brain that transduce the prostaglandin signal to mate and showed that the gonadal steroid 17α, 20β-dihydroxyprogesterone modulates mRNA levels of the putative receptor for PGF2α. The scientists keyed in on a receptor for PGF2α in the preoptic area (POA) within the hypothalamus of the brain, a region involved in sexual behavior across animals. They suspected that when PGF2α levels elevated in the fish, the molecule attaches to this receptor and triggers sexual behavior. Then they used CRISPR/Cas9 to generate PGF2α receptor knockout fish. This gene deletion or knockout uncoupled the sexual behavior from fertility status to prove that the receptor of PGF2α is necessary for the initiation of sexual behavior.
The finding has parallels across all vertebrates, and might influence the understanding of social behavior in humans. The next steps for this work will involve understanding other behaviors that are regulated by this receptor, and the finding provides insight into both the evolution of reproduction and sexual behaviors. In mammals and other vertebrates, PGF2α promotes the onset of labor and motherly behaviors, and this present research, coupled with other studies, suggests that PGF2α signaling has a common ancestral function associated with birth and its related behaviors.
Disease related changes in proteomics, protein folding, protein-protein interaction, Volume 2 (Volume Two: Latest in Genomics Methodologies for Therapeutics: Gene Editing, NGS and BioInformatics, Simulations and the Genome Ontology), Part 1: Next Generation Sequencing (NGS)
Disease related changes in proteomics, protein folding, protein-protein interaction
Curator: Larry H. Bernstein, MD, FCAP
LPBI
Frankenstein Proteins Stitched Together by Scientists
The Frankenstein monster, stitched together from disparate body parts, proved to be an abomination, but stitched together proteins may fare better. They may, for example, serve specific purposes in medicine, research, and industry. At least, that’s the ambition of scientists based at the University of North Carolina. They have developed a computational protocol called SEWING that builds new proteins from connected or disconnected pieces of existing structures. [Wikipedia]
Unlike Victor Frankenstein, who betrayed Promethean ambition when he sewed together his infamous creature, today’s biochemists are relatively modest. Rather than defy nature, they emulate it. For example, at the University of North Carolina (UNC), researchers have taken inspiration from natural evolutionary mechanisms to develop a technique called SEWING—Structure Extension With Native-substructure Graphs. SEWING is a computational protocol that describes how to stitch together new proteins from connected or disconnected pieces of existing structures.
“We can now begin to think about engineering proteins to do things that nothing else is capable of doing,” said UNC’s Brian Kuhlman, Ph.D. “The structure of a protein determines its function, so if we are going to learn how to design new functions, we have to learn how to design new structures. Our study is a critical step in that direction and provides tools for creating proteins that haven’t been seen before in nature.”
Traditionally, researchers have used computational protein design to recreate in the laboratory what already exists in the natural world. In recent years, their focus has shifted toward inventing novel proteins with new functionality. These design projects all start with a specific structural “blueprint” in mind, and as a result are limited. Dr. Kuhlman and his colleagues, however, believe that by removing the limitations of a predetermined blueprint and taking cues from evolution they can more easily create functional proteins.
Dr. Kuhlman’s UNC team developed a protein design approach that emulates natural mechanisms for shuffling tertiary structures such as pleats, coils, and furrows. Putting the approach into action, the UNC team mapped 50,000 stitched together proteins on the computer, and then it produced 21 promising structures in the laboratory. Details of this work appeared May 6 in the journal Science, in an article entitled, “Design of Structurally Distinct Proteins Using Strategies Inspired by Evolution.”
“Helical proteins designed with SEWING contain structural features absent from other de novo designed proteins and, in some cases, remain folded at more than 100°C,” wrote the authors. “High-resolution structures of the designed proteins CA01 and DA05R1 were solved by x-ray crystallography (2.2 angstrom resolution) and nuclear magnetic resonance, respectively, and there was excellent agreement with the design models.”
Essentially, the UNC scientists confirmed that the proteins they had synthesized contained the unique structural varieties that had been designed on the computer. The UNC scientists also determined that the structures they had created had new surface and pocket features. Such features, they noted, provide potential binding sites for ligands or macromolecules.
“We were excited that some had clefts or grooves on the surface, regions that naturally occurring proteins use for binding other proteins,” said the Science article’s first author, Tim M. Jacobs, Ph.D., a former graduate student in Dr. Kuhlman’s laboratory. “That’s important because if we wanted to create a protein that can act as a biosensor to detect a certain metabolite in the body, either for diagnostic or research purposes, it would need to have these grooves. Likewise, if we wanted to develop novel therapeutics, they would also need to attach to specific proteins.”
Currently, the UNC researchers are using SEWING to create proteins that can bind to several other proteins at a time. Many of the most important proteins are such multitaskers, including the blood protein hemoglobin.
Histone Mutation Deranges DNA Methylation to Cause Cancer
In some cancers, including chondroblastoma and a rare form of childhood sarcoma, a mutation in histone H3 reduces global levels of methylation (dark areas) in tumor cells but not in normal cells (arrowhead). The mutation locks the cells in a proliferative state to promote tumor development. [Laboratory of Chromatin Biology and Epigenetics at The Rockefeller University]
They have been called oncohistones, the mutated histones that are known to accompany certain pediatric cancers. Despite their suggestive moniker, oncohistones have kept their oncogenic secrets. For example, it has been unclear whether oncohistones are able to cause cancer on their own, or whether they need to act in concert with additional DNA mutations, that is, mutations other than those affecting histone structures.
While oncohistone mechanisms remain poorly understood, this particular question—the oncogenicity of lone oncohistones—has been resolved, at least in part. According to researchers based at The Rockefeller University, a change to the structure of a histone can trigger a tumor on its own.
This finding appeared May 13 in the journal Science, in an article entitled, “Histone H3K36 Mutations Promote Sarcomagenesis Through Altered Histone Methylation Landscape.” The article describes the Rockefeller team’s study of a histone protein called H3, which has been found in about 95% of samples of chondoblastoma, a benign tumor that arises in cartilage, typically during adolescence.
The Rockefeller scientists found that the H3 lysine 36–to–methionine (H3K36M) mutation impairs the differentiation of mesenchymal progenitor cells and generates undifferentiated sarcoma in vivo.
After the scientists inserted the H3 histone mutation into mouse mesenchymal progenitor cells (MPCs)—which generate cartilage, bone, and fat—they watched these cells lose the ability to differentiate in the lab. Next, the scientists injected the mutant cells into living mice, and the animals developed the tumors rich in MPCs, known as an undifferentiated sarcoma. Finally, the researchers tried to understand how the mutation causes the tumors to develop.
The scientists determined that H3K36M mutant nucleosomes inhibit the enzymatic activities of several H3K36 methyltransferases.
“Depleting H3K36 methyltransferases, or expressing an H3K36I mutant that similarly inhibits H3K36 methylation, is sufficient to phenocopy the H3K36M mutation,” the authors of the Science study wrote. “After the loss of H3K36 methylation, a genome-wide gain in H3K27 methylation leads to a redistribution of polycomb repressive complex 1 and de-repression of its target genes known to block mesenchymal differentiation.”
Essentially, when the H3K36M mutation occurs, the cell becomes locked in a proliferative state—meaning it divides constantly, leading to tumors. Specifically, the mutation inhibits enzymes that normally tag the histone with chemical groups known as methyls, allowing genes to be expressed normally.
In response to this lack of modification, another part of the histone becomes overmodified, or tagged with too many methyl groups. “This leads to an overall resetting of the landscape of chromatin, the complex of DNA and its associated factors, including histones,” explained co-author Peter Lewis, Ph.D., a professor at the University of Wisconsin-Madison and a former postdoctoral fellow in laboratory of C. David Allis, Ph.D., a professor at Rockefeller.
The finding—that a “resetting” of the chromatin landscape can lock the cell into a proliferative state—suggests that researchers should be on the hunt for more mutations in histones that might be driving tumors. For their part, the Rockefeller researchers are trying to learn more about how this specific mutation in histone H3 causes tumors to develop.
“We want to know which pathways cause the mesenchymal progenitor cells that carry the mutation to continue to divide, and not differentiate into the bone, fat, and cartilage cells they are destined to become,” said co-author Chao Lu, Ph.D., a postdoctoral fellow in the Allis lab.
Once researchers understand more about these pathways, added Dr. Lewis, they can consider ways of blocking them with drugs, particularly in tumors such as MPC-rich sarcomas—which, unlike chondroblastoma, can be deadly. In fact, drugs that block these pathways may already exist and may even be in use for other types of cancers.
“One long-term goal of our collaborative team is to better understand fundamental mechanisms that drive these processes, with the hope of providing new therapeutic approaches,” concluded Dr. Allis.
Histone H3K36 mutations promote sarcomagenesis through altered histone methylation landscape
Missense mutations (that change one amino acid for another) in histone H3 can produce a so-called oncohistone and are found in a number of pediatric cancers. For example, the lysine-36–to-methionine (K36M) mutation is seen in almost all chondroblastomas. Lu et al. show that K36M mutant histones are oncogenic, and they inhibit the normal methylation of this same residue in wild-type H3 histones. The mutant histones also interfere with the normal development of bone-related cells and the deposition of inhibitory chromatin marks.
Several types of pediatric cancers reportedly contain high-frequency missense mutations in histone H3, yet the underlying oncogenic mechanism remains poorly characterized. Here we report that the H3 lysine 36–to–methionine (H3K36M) mutation impairs the differentiation of mesenchymal progenitor cells and generates undifferentiated sarcoma in vivo. H3K36M mutant nucleosomes inhibit the enzymatic activities of several H3K36 methyltransferases. Depleting H3K36 methyltransferases, or expressing an H3K36I mutant that similarly inhibits H3K36 methylation, is sufficient to phenocopy the H3K36M mutation. After the loss of H3K36 methylation, a genome-wide gain in H3K27 methylation leads to a redistribution of polycomb repressive complex 1 and de-repression of its target genes known to block mesenchymal differentiation. Our findings are mirrored in human undifferentiated sarcomas in which novel K36M/I mutations in H3.1 are identified.
Mitochondria? We Don’t Need No Stinking Mitochondria!
Diagram comparing typical eukaryotic cell to the newly discovered mitochondria-free organism. [Karnkowska et al., 2016, Current Biology 26, 1–11]
The organelle that produces a significant portion of energy for eukaryotic cells would seemingly be indispensable, yet over the years, a number of organisms have been discovered that challenge that biological pretense. However, these so-called amitochondrial species may lack a defined organelle, but they still retain some residual functions of their mitochondria-containing brethren. Even the intestinal eukaryotic parasite Giardia intestinalis, which was for many years considered to be mitochondria-free, was proven recently to contain a considerably shriveled version of the organelle.
Now, an international group of scientists has released results from a new study that challenges the notion that mitochondria are essential for eukaryotes—discovering an organism that resides in the gut of chinchillas that contains absolutely no trace of mitochondria at all.
“In low-oxygen environments, eukaryotes often possess a reduced form of the mitochondrion, but it was believed that some of the mitochondrial functions are so essential that these organelles are indispensable for their life,” explained lead study author Anna Karnkowska, Ph.D., visiting scientist at the University of British Columbia in Vancouver. “We have characterized a eukaryotic microbe which indeed possesses no mitochondrion at all.”
Mysterious Eukaryote Missing Mitochondria
Researchers uncover the first example of a eukaryotic organism that lacks the organelles.
Monocercomonoides sp. PA203VLADIMIR HAMPL, CHARLES UNIVERSITY, PRAGUE, CZECH REPUBLIC
Scientists have long thought that mitochondria—organelles responsible for energy generation—are an essential and defining feature of a eukaryotic cell. Now, researchers from Charles University in Prague and their colleagues are challenging this notion with their discovery of a eukaryotic organism,Monocercomonoides species PA203, which lacks mitochondria. The team’s phylogenetic analysis, published today (May 12) in Current Biology,suggests that Monocercomonoides—which belong to the Oxymonadida group of protozoa and live in low-oxygen environments—did have mitochondria at one point, but eventually lost the organelles.
“This is quite a groundbreaking discovery,” said Thijs Ettema, who studies microbial genome evolution at Uppsala University in Sweden and was not involved in the work.
“This study shows that mitochondria are not so central for all lineages of living eukaryotes,” Toni Gabaldonof the Center for Genomic Regulation in Barcelona, Spain, who also was not involved in the work, wrote in an email to The Scientist. “Yet, this mitochondrial-devoid, single-cell eukaryote is as complex as other eukaryotic cells in almost any other aspect of cellular complexity.”
Charles University’s Vladimir Hampl studies the evolution of protists. Along with Anna Karnkowska and colleagues, Hampl decided to sequence the genome of Monocercomonoides, a little-studied protist that lives in the digestive tracts of vertebrates. The 75-megabase genome—the first of an oxymonad—did not contain any conserved genes found on mitochondrial genomes of other eukaryotes, the researchers found. It also did not contain any nuclear genes associated with mitochondrial functions.
“It was surprising and for a long time, we didn’t believe that the [mitochondria-associated genes were really not there]. We thought we were missing something,” Hampl told The Scientist. “But when the data kept accumulating, we switched to the hypothesis that this organism really didn’t have mitochondria.”
Because researchers have previously not found examples of eukaryotes without some form of mitochondria, the current theory of the origin of eukaryotes poses that the appearance of mitochondria was crucial to the identity of these organisms.
“We now view these mitochondria-like organelles as a continuum from full mitochondria to very small . Some anaerobic protists, for example, have only pared down versions of mitochondria, such as hydrogenosomes and mitosomes, which lack a mitochondrial genome. But these mitochondrion-like organelles perform essential functions of the iron-sulfur cluster assembly pathway, which is known to be conserved in virtually all eukaryotic organisms studied to date.
Yet, in their analysis, the researchers found no evidence of the presence of any components of this mitochondrial pathway.
Like the scaling down of mitochondria into mitosomes in some organisms, the ancestors of modernMonocercomonoides once had mitochondria. “Because this organism is phylogenetically nested among relatives that had conventional mitochondria, this is most likely a secondary adaptation,” said Michael Gray, a biochemist who studies mitochondria at Dalhousie University in Nova Scotia and was not involved in the study. According to Gray, the finding of a mitochondria-deficient eukaryote does not mean that the organelles did not play a major role in the evolution of eukaryotic cells.
To be sure they were not missing mitochondrial proteins, Hampl’s team also searched for potential mitochondrial protein homologs of other anaerobic species, and for signature sequences of a range of known mitochondrial proteins. While similar searches with other species uncovered a few mitochondrial proteins, the team’s analysis of Monocercomonoides came up empty.
“The data is very complete,” said Ettema. “It is difficult to prove the absence of something but [these authors] do a convincing job.”
To form the essential iron-sulfur clusters, the team discovered that Monocercomonoides use a sulfur mobilization system found in the cytosol, and that an ancestor of the organism acquired this system by lateral gene transfer from bacteria. This cytosolic, compensating system allowed Monocercomonoides to lose the otherwise essential iron-sulfur cluster-forming pathway in the mitochondrion, the team proposed.
“This work shows the great evolutionary plasticity of the eukaryotic cell,” said Karnkowska, who participated in the study while she was a postdoc at Charles University. Karnkowska, who is now a visiting researcher at the University of British Columbia in Canada, added: “This is a striking example of how far the evolution of a eukaryotic cell can go that was beyond our expectations.”
“The results highlight how many surprises may await us in the poorly studied eukaryotic phyla that live in under-explored environments,” Gabaldon said.
Ettema agreed. “Now that we’ve found one, we need to look at the bigger picture and see if there are other examples of eukaryotes that have lost their mitochondria, to understand how adaptable eukaryotes are.”
Karnkowska et al., “A eukaryote without a mitochondrial organelle,” Current Biology,doi:10.1016/j.cub.2016.03.053, 2016.
•Monocercomonoides sp. is a eukaryotic microorganism with no mitochondria
•The complete absence of mitochondria is a secondary loss, not an ancestral feature
•The essential mitochondrial ISC pathway was replaced by a bacterial SUF system
The presence of mitochondria and related organelles in every studied eukaryote supports the view that mitochondria are essential cellular components. Here, we report the genome sequence of a microbial eukaryote, the oxymonad Monocercomonoides sp., which revealed that this organism lacks all hallmark mitochondrial proteins. Crucially, the mitochondrial iron-sulfur cluster assembly pathway, thought to be conserved in virtually all eukaryotic cells, has been replaced by a cytosolic sulfur mobilization system (SUF) acquired by lateral gene transfer from bacteria. In the context of eukaryotic phylogeny, our data suggest that Monocercomonoides is not primitively amitochondrial but has lost the mitochondrion secondarily. This is the first example of a eukaryote lacking any form of a mitochondrion, demonstrating that this organelle is not absolutely essential for the viability of a eukaryotic cell.
This method catches a bait protein together with its associated protein partners in virus-like particles that are budded from human cells. Like this, cell lysis is not needed and protein complexes are preserved during purification.
With his feet in both a proteomics lab and an interactomics lab, VIB/UGent professor Sven Eyckerman is well aware of the shortcomings of conventional approaches to analyze protein complexes. The lysis conditions required in mass spectrometry–based strategies to break open cell membranes often affect protein-protein interactions. “The first step in a classical study on protein complexes essentially turns the highly organized cellular structure into a big messy soup”, Eyckerman explains.
Inspired by virus biology, Eyckerman came up with a creative solution. “We used the natural process of HIV particle formation to our benefit by hacking a completely safe form of the virus to abduct intact protein machines from the cell.” It is well known that the HIV virus captures a number of host proteins during its particle formation. By fusing a bait protein to the HIV-1 GAG protein, interaction partners become trapped within virus-like particles that bud from mammalian cells. Standard proteomic approaches are used next to reveal the content of these particles. Fittingly, the team named the method ‘Virotrap’.
The Virotrap approach is exceptional as protein networks can be characterized under natural conditions. By trapping protein complexes in the protective environment of a virus-like shell, the intact complexes are preserved during the purification process. The researchers showed the method was suitable for detection of known binary interactions as well as mass spectrometry-based identification of novel protein partners.
Virotrap is a textbook example of bringing research teams with complementary expertise together. Cross-pollination with the labs of Jan Tavernier (VIB/UGent) and Kris Gevaert (VIB/UGent) enabled the development of this platform.
Jan Tavernier: “Virotrap represents a new concept in co-complex analysis wherein complex stability is physically guaranteed by a protective, physical structure. It is complementary to the arsenal of existing interactomics methods, but also holds potential for other fields, like drug target characterization. We also developed a small molecule-variant of Virotrap that could successfully trap protein partners for small molecule baits.”
Kris Gevaert: “Virotrap can also impact our understanding of disease pathways. We were actually surprised to see that this virus-based system could be used to study antiviral pathways, like Toll-like receptor signaling. Understanding these protein machines in their natural environment is essential if we want to modulate their activity in pathology.“
Trapping mammalian protein complexes in viral particles
Cell lysis is an inevitable step in classical mass spectrometry–based strategies to analyse protein complexes. Complementary lysis conditions, in situ cross-linking strategies and proximal labelling techniques are currently used to reduce lysis effects on the protein complex. We have developed Virotrap, a viral particle sorting approach that obviates the need for cell homogenization and preserves the protein complexes during purification. By fusing a bait protein to the HIV-1 GAG protein, we show that interaction partners become trapped within virus-like particles (VLPs) that bud from mammalian cells. Using an efficient VLP enrichment protocol, Virotrap allows the detection of known binary interactions and MS-based identification of novel protein partners as well. In addition, we show the identification of stimulus-dependent interactions and demonstrate trapping of protein partners for small molecules. Virotrap constitutes an elegant complementary approach to the arsenal of methods to study protein complexes.
Proteins mostly exert their function within supramolecular complexes. Strategies for detecting protein–protein interactions (PPIs) can be roughly divided into genetic systems1 and co-purification strategies combined with mass spectrometry (MS) analysis (for example, AP–MS)2. The latter approaches typically require cell or tissue homogenization using detergents, followed by capture of the protein complex using affinity tags3 or specific antibodies4. The protein complexes extracted from this ‘soup’ of constituents are then subjected to several washing steps before actual analysis by trypsin digestion and liquid chromatography–MS/MS analysis. Such lysis and purification protocols are typically empirical and have mostly been optimized using model interactions in single labs. In fact, lysis conditions can profoundly affect the number of both specific and nonspecific proteins that are identified in a typical AP–MS set-up. Indeed, recent studies using the nuclear pore complex as a model protein complex describe optimization of purifications for the different proteins in the complex by examining 96 different conditions5. Nevertheless, for new purifications, it remains hard to correctly estimate the loss of factors in a standard AP–MS experiment due to washing and dilution effects during treatments (that is, false negatives). These considerations have pushed the concept of stabilizing PPIs before the actual homogenization step. A classical approach involves cross-linking with simple reagents (for example, formaldehyde) or with more advanced isotope-labelled cross-linkers (reviewed in ref. 2). However, experimental challenges such as cell permeability and reactivity still preclude the widespread use of cross-linking agents. Moreover, MS-generated spectra of cross-linked peptides are notoriously difficult to identify correctly. A recent lysis-independent solution involves the expression of a bait protein fused to a promiscuous biotin ligase, which results in labelling of proteins proximal to the activity of the enzyme-tagged bait protein6. When compared with AP–MS, this BioID approach delivers a complementary set of candidate proteins, including novel interaction partners7, 8. Such particular studies clearly underscore the need for complementary approaches in the co-complex strategies.
The evolutionary stress on viruses promoted highly condensed coding of information and maximal functionality for small genomes. Accordingly, for HIV-1 it is sufficient to express a single protein, the p55 GAG protein, for efficient production of virus-like particles (VLPs) from cells9, 10. This protein is highly mobile before its accumulation in cholesterol-rich regions of the membrane, where multimerization initiates the budding process11. A total of 4,000–5,000 GAG molecules is required to form a single particle of about 145 nm (ref. 12). Both VLPs and mature viruses contain a number of host proteins that are recruited by binding to viral proteins. These proteins can either contribute to the infectivity (for example, Cyclophilin/FKBPA13) or act as antiviral proteins preventing the spreading of the virus (for example, APOBEC proteins14).
We here describe the development and application of Virotrap, an elegant co-purification strategy based on the trapping of a bait protein together with its associated protein partners in VLPs that are budded from the cell. After enrichment, these particles can be analysed by targeted (for example, western blotting) or unbiased approaches (MS-based proteomics). Virotrap allows detection of known binary PPIs, analysis of protein complexes and their dynamics, and readily detects protein binders for small molecules.
Concept of the Virotrap system
Classical AP–MS approaches rely on cell homogenization to access protein complexes, a step that can vary significantly with the lysis conditions (detergents, salt concentrations, pH conditions and so on)5. To eliminate the homogenization step in AP–MS, we reasoned that incorporation of a protein complex inside a secreted VLP traps the interaction partners under native conditions and protects them during further purification. We thus explored the possibility of protein complex packaging by the expression of GAG-bait protein chimeras (Fig. 1) as expression of GAG results in the release of VLPs from the cells9, 10. As a first PPI pair to evaluate this concept, we selected the HRAS protein as a bait combined with the RAF1 prey protein. We were able to specifically detect the HRAS–RAF1 interaction following enrichment of VLPs via ultracentrifugation (Supplementary Fig. 1a). To prevent tedious ultracentrifugation steps, we designed a novel single-step protocol wherein we co-express the vesicular stomatitis virus glycoprotein (VSV-G) together with a tagged version of this glycoprotein in addition to the GAG bait and prey. Both tagged and untagged VSV-G proteins are probably presented as trimers on the surface of the VLPs, allowing efficient antibody-based recovery from large volumes. The HRAS–RAF1 interaction was confirmed using this single-step protocol (Supplementary Fig. 1b). No associations with unrelated bait or prey proteins were observed for both protocols.
Figure 1: Schematic representation of the Virotrap strategy.
Expression of a GAG-bait fusion protein (1) results in submembrane multimerization (2) and subsequent budding of VLPs from cells (3). Interaction partners of the bait protein are also trapped within these VLPs and can be identified after purification by western blotting or MS analysis (4).
Virotrap for the detection of binary interactions
We next explored the reciprocal detection of a set of PPI pairs, which were selected based on published evidence and cytosolic localization15. After single-step purification and western blot analysis, we could readily detect reciprocal interactions between CDK2 and CKS1B, LCP2 and GRAP2, and S100A1 and S100B (Fig. 2a). Only for the LCP2 prey we observed nonspecific association with an irrelevant bait construct. However, the particle levels of the GRAP2 bait were substantially lower as compared with those of the GAG control construct (GAG protein levels in VLPs; Fig. 2a, second panel of the LCP2 prey). After quantification of the intensities of bait and prey proteins and normalization of prey levels using bait levels, we observed a strong enrichment for the GAG-GRAP2 bait (Supplementary Fig. 2).
…..
Virotrap for unbiased discovery of novel interactions
For the detection of novel interaction partners, we scaled up VLP production and purification protocols (Supplementary Fig. 5 and Supplementary Note 1 for an overview of the protocol) and investigated protein partners trapped using the following bait proteins: Fas-associated via death domain (FADD), A20 (TNFAIP3), nuclear factor-κB (NF-κB) essential modifier (IKBKG), TRAF family member-associated NF-κB activator (TANK), MYD88 and ring finger protein 41 (RNF41). To obtain specific interactors from the lists of identified proteins, we challenged the data with a combined protein list of 19 unrelated Virotrap experiments (Supplementary Table 1 for an overview). Figure 3 shows the design and the list of candidate interactors obtained after removal of all proteins that were found in the 19 control samples (including removal of proteins from the control list identified with a single peptide). The remaining list of confident protein identifications (identified with at least two peptides in at least two biological repeats) reveals both known and novel candidate interaction partners. All candidate interactors including single peptide protein identifications are given in Supplementary Data 2 and also include recurrent protein identifications of known interactors based on a single peptide; for example, CASP8 for FADD and TANK for NEMO. Using alternative methods, we confirmed the interaction between A20 and FADD, and the associations with transmembrane proteins (insulin receptor and insulin-like growth factor receptor 1) that were captured using RNF41 as a bait (Supplementary Fig. 6). To address the use of Virotrap for the detection of dynamic interactions, we activated the NF-κB pathway via the tumour necrosis factor (TNF) receptor (TNFRSF1A) using TNFα (TNF) and performed Virotrap analysis using A20 as bait (Fig. 3). This resulted in the additional enrichment of receptor-interacting kinase (RIPK1), TNFR1-associated via death domain (TRADD), TNFRSF1A and TNF itself, confirming the expected activated complex20.
Figure 3: Use of Virotrap for unbiased interactome analysis
Lysis conditions used in AP–MS strategies are critical for the preservation of protein complexes. A multitude of lysis conditions have been described, culminating in a recent report where protein complex stability was assessed under 96 lysis/purification protocols5. Moreover, the authors suggest to optimize the conditions for every complex, implying an important workload for researchers embarking on protein complex analysis using classical AP–MS. As lysis results in a profound change of the subcellular context and significantly alters the concentration of proteins, loss of complex integrity during a classical AP–MS protocol can be expected. A clear evolution towards ‘lysis-independent’ approaches in the co-complex analysis field is evident with the introduction of BioID6 and APEX25 where proximal proteins, including proteins residing in the complex, are labelled with biotin by an enzymatic activity fused to a bait protein. A side-by-side comparison between classical AP–MS and BioID showed overlapping and unique candidate binding proteins for both approaches7, 8, supporting the notion that complementary methods are needed to provide a comprehensive view on protein complexes. This has also been clearly demonstrated for binary approaches15 and is a logical consequence of the heterogenic nature underlying PPIs (binding mechanism, requirement for posttranslational modifications, location, affinity and so on).
In this report, we explore an alternative, yet complementary method to isolate protein complexes without interfering with cellular integrity. By trapping protein complexes in the protective environment of a virus-like shell, the intact complexes are preserved during the purification process. This constitutes a new concept in co-complex analysis wherein complex stability is physically guaranteed by a protective, physical structure. A comparison of our Virotrap approach with AP–MS shows complementary data, with specific false positives and false negatives for both methods (Supplementary Fig. 7).
The current implementation of the Virotrap platform implies the use of a GAG-bait construct resulting in considerable expression of the bait protein. Different strategies are currently pursued to reduce bait expression including co-expression of a native GAG protein together with the GAG-bait protein, not only reducing bait expression but also creating more ‘space’ in the particles potentially accommodating larger bait protein complexes. Nevertheless, the presence of the bait on the forming GAG scaffold creates an intracellular affinity matrix (comparable to the early in vitro affinity columns for purification of interaction partners from lysates26) that has the potential to compete with endogenous complexes by avidity effects. This avidity effect is a powerful mechanism that aids in the recruitment of cyclophilin to GAG27, a well-known weak interaction (Kd=16 μM (ref. 28)) detectable as a background association in the Virotrap system. Although background binding may be increased by elevated bait expression, weaker associations are readily detectable (for example, MAL—MYD88-binding study; Fig. 2c).
The size of Virotrap particles (around 145 nm) suggests limitations in the size of the protein complex that can be accommodated in the particles. Further experimentation is required to define the maximum size of proteins or the number of protein complexes that can be trapped inside the particles.
….
In conclusion, Virotrap captures significant parts of known interactomes and reveals new interactions. This cell lysis-free approach purifies protein complexes under native conditions and thus provides a powerful method to complement AP–MS or other PPI data. Future improvements of the system include strategies to reduce bait expression to more physiological levels and application of advanced data analysis options to filter out background. These developments can further aid in the deployment of Virotrap as a powerful extension of the current co-complex technology arsenal.
New Autism Blood Biomarker Identified
Researchers at UT Southwestern Medical Center have identified a blood biomarker that may aid in earlier diagnosis of children with autism spectrum disorder, or ASD
In a recent edition of Scientific Reports, UT Southwestern researchers reported on the identification of a blood biomarker that could distinguish the majority of ASD study participants versus a control group of similar age range. In addition, the biomarker was significantly correlated with the level of communication impairment, suggesting that the blood test may give insight into ASD severity.
“Numerous investigators have long sought a biomarker for ASD,” said Dr. Dwight German, study senior author and Professor of Psychiatry at UT Southwestern. “The blood biomarker reported here along with others we are testing can represent a useful test with over 80 percent accuracy in identifying ASD.”
ASD1 – was 66 percent accurate in diagnosing ASD. When combined with thyroid stimulating hormone level measurements, the ASD1-binding biomarker was 73 percent accurate at diagnosis
A Search for Blood Biomarkers for Autism: Peptoids
Autism spectrum disorder (ASD) is a neurodevelopmental disorder characterized by impairments in social interaction and communication, and restricted, repetitive patterns of behavior. In order to identify individuals with ASD and initiate interventions at the earliest possible age, biomarkers for the disorder are desirable. Research findings have identified widespread changes in the immune system in children with autism, at both systemic and cellular levels. In an attempt to find candidate antibody biomarkers for ASD, highly complex libraries of peptoids (oligo-N-substituted glycines) were screened for compounds that preferentially bind IgG from boys with ASD over typically developing (TD) boys. Unexpectedly, many peptoids were identified that preferentially bound IgG from TD boys. One of these peptoids was studied further and found to bind significantly higher levels (>2-fold) of the IgG1 subtype in serum from TD boys (n = 60) compared to ASD boys (n = 74), as well as compared to older adult males (n = 53). Together these data suggest that ASD boys have reduced levels (>50%) of an IgG1 antibody, which resembles the level found normally with advanced age. In this discovery study, the ASD1 peptoid was 66% accurate in predicting ASD.
….
Peptoid libraries have been used previously to search for autoantibodies for neurodegenerative diseases19 and for systemic lupus erythematosus (SLE)21. In the case of SLE, peptoids were identified that could identify subjects with the disease and related syndromes with moderate sensitivity (70%) and excellent specificity (97.5%). Peptoids were used to measure IgG levels from both healthy subjects and SLE patients. Binding to the SLE-peptoid was significantly higher in SLE patients vs. healthy controls. The IgG bound to the SLE-peptoid was found to react with several autoantigens, suggesting that the peptoids are capable of interacting with multiple, structurally similar molecules. These data indicate that IgG binding to peptoids can identify subjects with high levels of pathogenic autoantibodies vs. a single antibody.
In the present study, the ASD1 peptoid binds significantly lower levels of IgG1 in ASD males vs. TD males. This finding suggests that the ASD1 peptoid recognizes antibody(-ies) of an IgG1 subtype that is (are) significantly lower in abundance in the ASD males vs. TD males. Although a previous study14 has demonstrated lower levels of plasma IgG in ASD vs. TD children, here, we additionally quantified serum IgG levels in our individuals and found no difference in IgG between the two groups (data not shown). Furthermore, our IgG levels did not correlate with ASD1 binding levels, indicating that ASD1 does not bind IgG generically, and that the peptoid’s ability to differentiate between ASD and TD males is related to a specific antibody(-ies).
ASD subjects underwent a diagnostic evaluation using the ADOS and ADI-R, and application of the DSM-IV criteria prior to study inclusion. Only those subjects with a diagnosis of Autistic Disorder were included in the study. The ADOS is a semi-structured observation of a child’s behavior that allows examiners to observe the three core domains of ASD symptoms: reciprocal social interaction, communication, and restricted and repetitive behaviors1. When ADOS subdomain scores were compared with peptoid binding, the only significant relationship was with Social Interaction. However, the positive correlation would suggest that lower peptoid binding is associated with better social interaction, not poorer social interaction as anticipated.
The ADI-R is a structured parental interview that measures the core features of ASD symptoms in the areas of reciprocal social interaction, communication and language, and patterns of behavior. Of the three ADI-R subdomains, only the Communication domain was related to ASD1 peptoid binding, and this correlation was negative suggesting that low peptoid binding is associated with greater communication problems. These latter data are similar to the findings of Heuer et al.14 who found that children with autism with low levels of plasma IgG have high scores on the Aberrant Behavior Checklist (p < 0.0001). Thus, peptoid binding to IgG1 may be useful as a severity marker for ASD allowing for further characterization of individuals, but further research is needed.
It is interesting that in serum samples from older men, the ASD1 binding is similar to that in the ASD boys. This is consistent with the observation that with aging there is a reduction in the strength of the immune system, and the changes are gender-specific25. Recent studies using parabiosis26, in which blood from young mice reverse age-related impairments in cognitive function and synaptic plasticity in old mice, reveal that blood constituents from young subjects may contain important substances for maintaining neuronal functions. Work is in progress to identify the antibody/antibodies that are differentially binding to the ASD1 peptoid, which appear as a single band on the electrophoresis gel (Fig. 4).
……..
The ADI-R is a structured parental interview that measures the core features of ASD symptoms in the areas of reciprocal social interaction, communication and language, and patterns of behavior. Of the three ADI-R subdomains, only the Communication domain was related to ASD1 peptoid binding, and this correlation was negative suggesting that low peptoid binding is associated with greater communication problems. These latter data are similar to the findings of Heuer et al.14 who found that children with autism with low levels of plasma IgG have high scores on the Aberrant Behavior Checklist (p < 0.0001). Thus, peptoid binding to IgG1 may be useful as a severity marker for ASD allowing for further characterization of individuals, but further research is needed.
Titration of IgG binding to ASD1 using serum pooled from 10 TD males and 10 ASD males demonstrates ASD1’s ability to differentiate between the two groups. (B)Detecting IgG1 subclass instead of total IgG amplifies this differentiation. (C) IgG1 binding of individual ASD (n=74) and TD (n=60) male serum samples (1:100 dilution) to ASD1 significantly differs with TD>ASD. In addition, IgG1 binding of older adult male (AM) serum samples (n=53) to ASD1 is significantly lower than TD males, and not different from ASD males. The three groups were compared with a Kruskal-Wallis ANOVA, H = 10.1781, p<0.006. **p<0.005. Error bars show SEM. (D) Receiver-operating characteristic curve for ASD1’s ability to discriminate between ASD and TD males.
Association between peptoid binding and ADOS and ADI-R subdomains
Higher scores in any domain on the ADOS and ADI-R are indicative of more abnormal behaviors and/or symptoms. Among ADOS subdomains, there was no significant relationship between Communication and peptoid binding (z = 0.04, p = 0.966), Communication + Social interaction (z = 1.53, p = 0.127), or Stereotyped Behaviors and Restrictive Interests (SBRI) (z = 0.46, p = 0.647). Higher scores on the Social Interaction domain were significantly associated with higher peptoid binding (z = 2.04, p = 0.041).
Among ADI-R subdomains, higher scores on the Communication domain were associated with lower levels of peptoid binding (z = −2.28, p = 0.023). There was not a significant relationship between Social Interaction (z = 0.07, p = 0.941) or Restrictive/Repetitive Stereotyped Behaviors (z = −1.40, p = 0.162) and peptoid binding.
Computational Model Finds New Protein-Protein Interactions
Researchers at University of Pittsburgh have discovered 500 new protein-protein interactions (PPIs) associated with genes linked to schizophrenia.
Using a computational model they developed, researchers at the University of Pittsburgh School of Medicine have discovered more than 500 new protein-protein interactions (PPIs) associated with genes linked to schizophrenia. The findings, published online in npj Schizophrenia, a Nature Publishing Group journal, could lead to greater understanding of the biological underpinnings of this mental illness, as well as point the way to treatments.
There have been many genome-wide association studies (GWAS) that have identified gene variants associated with an increased risk for schizophrenia, but in most cases there is little known about the proteins that these genes make, what they do and how they interact, said senior investigator Madhavi Ganapathiraju, Ph.D., assistant professor of biomedical informatics, Pitt School of Medicine.
“GWAS studies and other research efforts have shown us what genes might be relevant in schizophrenia,” she said. “What we have done is the next step. We are trying to understand how these genes relate to each other, which could show us the biological pathways that are important in the disease.”
Each gene makes proteins and proteins typically interact with each other in a biological process. Information about interacting partners can shed light on the role of a gene that has not been studied, revealing pathways and biological processes associated with the disease and also its relation to other complex diseases.
Dr. Ganapathiraju’s team developed a computational model called High-Precision Protein Interaction Prediction (HiPPIP) and applied it to discover PPIs of schizophrenia-linked genes identified through GWAS, as well as historically known risk genes. They found 504 never-before known PPIs, and noted also that while schizophrenia-linked genes identified historically and through GWAS had little overlap, the model showed they shared more than 100 common interactors.
“We can infer what the protein might do by checking out the company it keeps,” Dr. Ganapathiraju explained. “For example, if I know you have many friends who play hockey, it could mean that you are involved in hockey, too. Similarly, if we see that an unknown protein interacts with multiple proteins involved in neural signaling, for example, there is a high likelihood that the unknown entity also is involved in the same.”
Dr. Ganapathiraju and colleagues have drawn such inferences on protein function based on the PPIs of proteins, and made their findings available on a website Schizo-Pi. This information can be used by biologists to explore the schizophrenia interactome with the aim of understanding more about the disease or developing new treatment drugs.
Schizophrenia interactome with 504 novel protein–protein interactions
(GWAS) have revealed the role of rare and common genetic variants, but the functional effects of the risk variants remain to be understood. Protein interactome-based studies can facilitate the study of molecular mechanisms by which the risk genes relate to schizophrenia (SZ) genesis, but protein–protein interactions (PPIs) are unknown for many of the liability genes. We developed a computational model to discover PPIs, which is found to be highly accurate according to computational evaluations and experimental validations of selected PPIs. We present here, 365 novel PPIs of liability genes identified by the SZ Working Group of the Psychiatric Genomics Consortium (PGC). Seventeen genes that had no previously known interactions have 57 novel interactions by our method. Among the new interactors are 19 drug targets that are targeted by 130 drugs. In addition, we computed 147 novel PPIs of 25 candidate genes investigated in the pre-GWAS era. While there is little overlap between the GWAS genes and the pre-GWAS genes, the interactomes reveal that they largely belong to the same pathways, thus reconciling the apparent disparities between the GWAS and prior gene association studies. The interactome including 504 novel PPIs overall, could motivate other systems biology studies and trials with repurposed drugs. The PPIs are made available on a webserver, called Schizo-Pi at http://severus.dbmi.pitt.edu/schizo-pi with advanced search capabilities.
Schizophrenia (SZ) is a common, potentially severe psychiatric disorder that afflicts all populations.1 Gene mapping studies suggest that SZ is a complex disorder, with a cumulative impact of variable genetic effects coupled with environmental factors.2 As many as 38 genome-wide association studies (GWAS) have been reported on SZ out of a total of 1,750 GWAS publications on 1,087 traits or diseases reported in the GWAS catalog maintained by the National Human Genome Research Institute of USA3 (as of April 2015), revealing the common variants associated with SZ.4 The SZ Working Group of the Psychiatric Genomics Consortium (PGC) identified 108 genetic loci that likely confer risk for SZ.5 While the role of genetics has been clearly validated by this study, the functional impact of the risk variants is not well-understood.6,7 Several of the genes implicated by the GWAS have unknown functions and could participate in possibly hitherto unknown pathways.8 Further, there is little or no overlap between the genes identified through GWAS and ‘candidate genes’ proposed in the pre-GWAS era.9
Interactome-based studies can be useful in discovering the functional associations of genes. For example,disrupted in schizophrenia 1 (DISC1), an SZ related candidate gene originally had no known homolog in humans. Although it had well-characterized protein domains such as coiled-coil domains and leucine-zipper domains, its function was unknown.10,11 Once its protein–protein interactions (PPIs) were determined using yeast 2-hybrid technology,12 investigators successfully linked DISC1 to cAMP signaling, axon elongation, and neuronal migration, and accelerated the research pertaining to SZ in general, and DISC1 in particular.13 Typically such studies are carried out on known protein–protein interaction (PPI) networks, or as in the case of DISC1, when there is a specific gene of interest, its PPIs are determined by methods such as yeast 2-hybrid technology.
Knowledge of human PPI networks is thus valuable for accelerating discovery of protein function, and indeed, biomedical research in general. However, of the hundreds of thousands of biophysical PPIs thought to exist in the human interactome,14,15 <100,000 are known today (Human Protein Reference Database, HPRD16 and BioGRID17 databases). Gold standard experimental methods for the determination of all the PPIs in human interactome are time-consuming, expensive and may not even be feasible, as about 250 million pairs of proteins would need to be tested overall; high-throughput methods such as yeast 2-hybrid have important limitations for whole interactome determination as they have a low recall of 23% (i.e., remaining 77% of true interactions need to be determined by other means), and a low precision (i.e., the screens have to be repeated multiple times to achieve high selectivity).18,19Computational methods are therefore necessary to complete the interactome expeditiously. Algorithms have begun emerging to predict PPIs using statistical machine learning on the characteristics of the proteins, but these algorithms are employed predominantly to study yeast. Two significant computational predictions have been reported for human interactome; although they have had high false positive rates, these methods have laid the foundation for computational prediction of human PPIs.20,21
We have created a new PPI prediction model called High-Confidence Protein–Protein Interaction Prediction (HiPPIP) model. Novel interactions predicted with this model are making translational impact. For example, we discovered a PPI between OASL and DDX58, which on validation showed that an increased expression of OASL could boost innate immunity to combat influenza by activating the RIG-I pathway.22 Also, the interactome of the genes associated with congenital heart disease showed that the disease morphogenesis has a close connection with the structure and function of cilia.23Here, we describe the HiPPIP model and its application to SZ genes to construct the SZ interactome. After computational evaluations and experimental validations of selected novel PPIs, we present here 504 highly confident novel PPIs in the SZ interactome, shedding new light onto several uncharacterized genes that are associated with SZ.
We developed a computational model called HiPPIP to predict PPIs (see Methods and Supplementary File 1). The model has been evaluated by computational methods and experimental validations and is found to be highly accurate. Evaluations on a held-out test data showed a precision of 97.5% and a recall of 5%. 5% recall out of 150,000 to 600,000 estimated number of interactions in the human interactome corresponds to 7,500–30,000 novel PPIs in the whole interactome. Note that, it is likely that the real precision would be higher than 97.5% because in this test data, randomly paired proteins are treated as non-interacting protein pairs, whereas some of them may actually be interacting pairs with a small probability; thus, some of the pairs that are treated as false positives in test set are likely to be true but hitherto unknown interactions. In Figure 1a, we show the precision versus recall of our method on ‘hub proteins’ where we considered all pairs that received a score >0.5 by HiPPIP to be novel interactions. In Figure 1b, we show the number of true positives versus false positives observed in hub proteins. Both these figures also show our method to be superior in comparison to the prediction of membrane-receptor interactome by Qi et al’s.24 True positives versus false positives are also shown for individual hub proteins by our method in Figure 1cand by Qi et al’s.23 in Figure 1d. These evaluations showed that our predictions contain mostly true positives. Unlike in other domains where ranked lists are commonly used such as information retrieval, in PPI prediction the ‘false positives’ may actually be unlabeled instances that are indeed true interactions that are not yet discovered. In fact, such unlabeled pairs predicted as interactors of the hub gene HMGB1 (namely, the pairs HMGB1-KL and HMGB1-FLT1) were validated by experimental methods and found to be true PPIs (See the Figures e–g inSupplementary File 3). Thus, we concluded that the protein pairs that received a score of ⩾0.5 are highly confident to be true interactions. The pairs that receive a score less than but close to 0.5 (i.e., in the range of 0.4–0.5) may also contain several true PPIs; however, we cannot confidently say that all in this range are true PPIs. Only the PPIs predicted with a score >0.5 are included in the interactome.
Computational evaluation of predicted protein–protein interactions on hub proteins: (a) precision recall curve. (b) True positive versus false positives in ranked lists of hub type membrane receptors for our method and that by Qi et al. True positives versus false positives are shown for individual membrane receptors by our method in (c) and by Qi et al. in (d). Thick line is the average, which is also the same as shown in (b). Note:x-axis is recall in (a), whereas it is number of false positives in (b–d). The range of y-axis is observed by varying the threshold from 1.0–0 in (a), and to 0.5 in (b–d).
SZ interactome
By applying HiPPIP to the GWAS genes and Historic (pre-GWAS) genes, we predicted over 500 high confidence new PPIs adding to about 1400 previously known PPIs.
Schizophrenia interactome: network view of the schizophrenia interactome is shown as a graph, where genes are shown as nodes and PPIs as edges connecting the nodes. Schizophrenia-associated genes are shown as dark blue nodes, novel interactors as red color nodes and known interactors as blue color nodes. The source of the schizophrenia genes is indicated by its label font, where Historic genes are shown italicized, GWAS genes are shown in bold, and the one gene that is common to both is shown in italicized and bold. For clarity, the source is also indicated by the shape of the node (triangular for GWAS and square for Historic and hexagonal for both). Symbols are shown only for the schizophrenia-associated genes; actual interactions may be accessed on the web. Red edges are the novel interactions, whereas blue edges are known interactions. GWAS, genome-wide association studies of schizophrenia; PPI, protein–protein interaction.
We have made the known and novel interactions of all SZ-associated genes available on a webserver called Schizo-Pi, at the addresshttp://severus.dbmi.pitt.edu/schizo-pi. This webserver is similar to Wiki-Pi33 which presents comprehensive annotations of both participating proteins of a PPI side-by-side. The difference between Wiki-Pi which we developed earlier, and Schizo-Pi, is the inclusion of novel predicted interactions of the SZ genes into the latter.
Despite the many advances in biomedical research, identifying the molecular mechanisms underlying the disease is still challenging. Studies based on protein interactions were proven to be valuable in identifying novel gene associations that could shed new light on disease pathology.35 The interactome including more than 500 novel PPIs will help to identify pathways and biological processes associated with the disease and also its relation to other complex diseases. It also helps identify potential drugs that could be repurposed to use for SZ treatment.
Functional and pathway enrichment in SZ interactome
When a gene of interest has little known information, functions of its interacting partners serve as a starting point to hypothesize its own function. We computed statistically significant enrichment of GO biological process terms among the interacting partners of each of the genes using BinGO36 (see online at http://severus.dbmi.pitt.edu/schizo-pi).
Protein aggregation and aggregate toxicity: new insights into protein folding, misfolding diseases and biological evolution
Massimo Stefani · Christopher M. Dobson
Abstract The deposition of proteins in the form of amyloid fibrils and plaques is the characteristic feature of more than 20 degenerative conditions affecting either the central nervous system or a variety of peripheral tissues. As these conditions include Alzheimer’s, Parkinson’s and the prion diseases, several forms of fatal systemic amyloidosis, and at least one condition associated with medical intervention (haemodialysis), they are of enormous importance in the context of present-day human health and welfare. Much remains to be learned about the mechanism by which the proteins associated with these diseases aggregate and form amyloid structures, and how the latter affect the functions of the organs with which they are associated. A great deal of information concerning these diseases has emerged, however, during the past 5 years, much of it causing a number of fundamental assumptions about the amyloid diseases to be reexamined. For example, it is now apparent that the ability to form amyloid structures is not an unusual feature of the small number of proteins associated with these diseases but is instead a general property of polypeptide chains. It has also been found recently that aggregates of proteins not associated with amyloid diseases can impair the ability of cells to function to a similar extent as aggregates of proteins linked with specific neurodegenerative conditions. Moreover, the mature amyloid fibrils or plaques appear to be substantially less toxic than the prefibrillar aggregates that are their precursors. The toxicity of these early aggregates appears to result from an intrinsic ability to impair fundamental cellular processes by interacting with cellular membranes, causing oxidative stress and increases in free Ca2+ that eventually lead to apoptotic or necrotic cell death. The ‘new view’ of these diseases also suggests that other degenerative conditions could have similar underlying origins to those of the amyloidoses. In addition, cellular protection mechanisms, such as molecular chaperones and the protein degradation machinery, appear to be crucial in the prevention of disease in normally functioning living organisms. It also suggests some intriguing new factors that could be of great significance in the evolution of biological molecules and the mechanisms that regulate their behaviour.
The genetic information within a cell encodes not only the specific structures and functions of proteins but also the way these structures are attained through the process known as protein folding. In recent years many of the underlying features of the fundamental mechanism of this complex process and the manner in which it is regulated in living systems have emerged from a combination of experimental and theoretical studies [1]. The knowledge gained from these studies has also raised a host of interesting issues. It has become apparent, for example, that the folding and unfolding of proteins is associated with a whole range of cellular processes from the trafficking of molecules to specific organelles to the regulation of the cell cycle and the immune response. Such observations led to the inevitable conclusion that the failure to fold correctly, or to remain correctly folded, gives rise to many different types of biological malfunctions and hence to many different forms of disease [2]. In addition, it has been recognised recently that a large number of eukaryotic genes code for proteins that appear to be ‘natively unfolded’, and that proteins can adopt, under certain circumstances, highly organised multi-molecular assemblies whose structures are not specifically encoded in the amino acid sequence. Both these observations have raised challenging questions about one of the most fundamental principles of biology: the close relationship between the sequence, structure and function of proteins, as we discuss below [3].
It is well established that proteins that are ‘misfolded’, i.e. that are not in their functionally relevant conformation, are devoid of normal biological activity. In addition, they often aggregate and/or interact inappropriately with other cellular components leading to impairment of cell viability and eventually to cell death. Many diseases, often known as misfolding or conformational diseases, ultimately result from the presence in a living system of protein molecules with structures that are ‘incorrect’, i.e. that differ from those in normally functioning organisms [4]. Such diseases include conditions in which a specific protein, or protein complex, fails to fold correctly (e.g. cystic fibrosis, Marfan syndrome, amyotonic lateral sclerosis) or is not sufficiently stable to perform its normal function (e.g. many forms of cancer). They also include conditions in which aberrant folding behaviour results in the failure of a protein to be correctly trafficked (e.g. familial hypercholesterolaemia, α1-antitrypsin deficiency, and some forms of retinitis pigmentosa) [4]. The tendency of proteins to aggregate, often to give species extremely intractable to dissolution and refolding, is of course also well known in other circumstances. Examples include the formation of inclusion bodies during overexpression of heterologous proteins in bacteria and the precipitation of proteins during laboratory purification procedures. Indeed, protein aggregation is well established as one of the major difficulties associated with the production and handling of proteins in the biotechnology and pharmaceutical industries [5].
Considerable attention is presently focused on a group of protein folding diseases known as amyloidoses. In these diseases specific peptides or proteins fail to fold or to remain correctly folded and then aggregate (often with other components) so as to give rise to ‘amyloid’ deposits in tissue. Amyloid structures can be recognised because they possess a series of specific tinctorial and biophysical characteristics that reflect a common core structure based on the presence of highly organised βsheets [6]. The deposits in strictly defined amyloidoses are extracellular and can often be observed as thread-like fibrillar structures, sometimes assembled further into larger aggregates or plaques. These diseases include a range of sporadic, familial or transmissible degenerative diseases, some of which affect the brain and the central nervous system (e.g. Alzheimer’s and Creutzfeldt-Jakob diseases), while others involve peripheral tissues and organs such as the liver, heart and spleen (e.g. systemic amyloidoses and type II diabetes) [7, 8]. In other forms of amyloidosis, such as primary or secondary systemic amyloidoses, proteinaceous deposits are found in skeletal tissue and joints (e.g. haemodialysis-related amyloidosis) as well as in several organs (e.g. heart and kidney). Yet other components such as collagen, glycosaminoglycans and proteins (e.g. serum amyloid protein) are often present in the deposits protecting them against degradation [9, 10, 11]. Similar deposits to those in the amyloidoses are, however, found intracellularly in other diseases; these can be localised either in the cytoplasm, in the form of specialised aggregates known as aggresomes or as Lewy or Russell bodies or in the nucleus (see below).
The presence in tissue of proteinaceous deposits is a hallmark of all these diseases, suggesting a causative link between aggregate formation and pathological symptoms (often known as the amyloid hypothesis) [7, 8, 12]. At the present time the link between amyloid formation and disease is widely accepted on the basis of a large number of biochemical and genetic studies. The specific nature of the pathogenic species, and the molecular basis of their ability to damage cells, are however, the subject of intense debate [13, 14, 15, 16, 17, 18, 19, 20]. In neurodegenerative disorders it is very likely that the impairment of cellular function follows directly from the interactions of the aggregated proteins with cellular components [21, 22]. In the systemic non-neurological diseases, however, it is widely believed that the accumulation in vital organs of large amounts of amyloid deposits can by itself cause at least some of the clinical symptoms [23]. It is quite possible, however, that there are other more specific effects of aggregates on biochemical processes even in these diseases. The presence of extracellular or intracellular aggregates of a specific polypeptide molecule is a characteristic of all the 20 or so recognised amyloid diseases. The polypeptides involved include full length proteins (e.g. lysozyme or immunoglobulin light chains), biological peptides (amylin, atrial natriuretic factor) and fragments of larger proteins produced as a result of specific processing (e.g. the Alzheimer βpeptide) or of more general degradation [e.g. poly(Q) stretches cleaved from proteins with poly(Q) extensions such as huntingtin, ataxins and the androgen receptor]. The peptides and proteins associated with known amyloid diseases are listed in Table 1. In some cases the proteins involved have wild type sequences, as in sporadic forms of the diseases, but in other cases these are variants resulting from genetic mutations associated with familial forms of the diseases. In some cases both sporadic and familial diseases are associated with a given protein; in this case the mutational variants are usually associated with early-onset forms of the disease. In the case of the neurodegenerative diseases associated with the prion protein some forms of the diseases are transmissible. The existence of familial forms of a number of amyloid diseases has provided significant clues to the origins of the pathologies. For example, there are increasingly strong links between the age at onset of familial forms of disease and the effects of the mutations involved on the propensity of the affected proteins to aggregate in vitro. Such findings also support the link between the process of aggregation and the clinical manifestations of disease [24, 25].
The presence in cells of misfolded or aggregated proteins triggers a complex biological response. In the cytosol, this is referred to as the ‘heat shock response’ and in the endoplasmic reticulum (ER) it is known as the ‘unfolded protein response’. These responses lead to the expression, among others, of the genes for heat shock proteins (Hsp, or molecular chaperone proteins) and proteins involved in the ubiquitin-proteasome pathway [26]. The evolution of such complex biochemical machinery testifies to the fact that it is necessary for cells to isolate and clear rapidly and efficiently any unfolded or incorrectly folded protein as soon as it appears. In itself this fact suggests that these species could have a generally adverse effect on cellular components and cell viability. Indeed, it was a major step forward in understanding many aspects of cell biology when it was recognised that proteins previously associated only with stress, such as heat shock, are in fact crucial in the normal functioning of living systems. This advance, for example, led to the discovery of the role of molecular chaperones in protein folding and in the normal ‘housekeeping’ processes that are inherent in healthy cells [27, 28]. More recently a number of degenerative diseases, both neurological and systemic, have been linked to, or shown to be affected by, impairment of the ubiquitin-proteasome pathway (Table 2). The diseases are primarily associated with a reduction in either the expression or the biological activity of Hsps, ubiquitin, ubiquitinating or deubiquitinating enzymes and the proteasome itself, as we show below [29, 30, 31, 32], or even to the failure of the quality control mechanisms that ensure proper maturation of proteins in the ER. The latter normally leads to degradation of a significant proportion of polypeptide chains before they have attained their native conformations through retrograde translocation to the cytosol [33, 34].
….
It is now well established that the molecular basis of protein aggregation into amyloid structures involves the existence of ‘misfolded’ forms of proteins, i.e. proteins that are not in the structures in which they normally function in vivo or of fragments of proteins resulting from degradation processes that are inherently unable to fold [4, 7, 8, 36]. Aggregation is one of the common consequences of a polypeptide chain failing to reach or maintain its functional three-dimensional structure. Such events can be associated with specific mutations, misprocessing phenomena, aberrant interactions with metal ions, changes in environmental conditions, such as pH or temperature, or chemical modification (oxidation, proteolysis). Perturbations in the conformational properties of the polypeptide chain resulting from such phenomena may affect equilibrium 1 in Fig. 1 increasing the population of partially unfolded, or misfolded, species that are much more aggregation-prone than the native state.
Fig. 1 Overview of the possible fates of a newly synthesised polypeptide chain. The equilibrium ① between the partially folded molecules and the natively folded ones is usually strongly in favour of the latter except as a result of specific mutations, chemical modifications or partially destabilising solution conditions. The increased equilibrium populations of molecules in the partially or completely unfolded ensemble of structures are usually degraded by the proteasome; when this clearance mechanism is impaired, such species often form disordered aggregates or shift equilibrium ② towards the nucleation of pre-fibrillar assemblies that eventually grow into mature fibrils (equilibrium ③). DANGER! indicates that pre-fibrillar aggregates in most cases display much higher toxicity than mature fibrils. Heat shock proteins (Hsp) can suppress the appearance of pre-fibrillar assemblies by minimising the population of the partially folded molecules by assisting in the correct folding of the nascent chain and the unfolded protein response target incorrectly folded proteins for degradation.
……
Little is known at present about the detailed arrangement of the polypeptide chains themselves within amyloid fibrils, either those parts involved in the core βstrands or in regions that connect the various β-strands. Recent data suggest that the sheets are relatively untwisted and may in some cases at least exist in quite specific supersecondary structure motifs such as β-helices [6, 40] or the recently proposed µ-helix [41]. It seems possible that there may be significant differences in the way the strands are assembled depending on characteristics of the polypeptide chain involved [6, 42]. Factors including length, sequence (and in some cases the presence of disulphide bonds or post-translational modifications such as glycosylation) may be important in determining details of the structures. Several recent papers report structural models for amyloid fibrils containing different polypeptide chains, including the Aβ40 peptide, insulin and fragments of the prion protein, based on data from such techniques as cryo-electron microscopy and solid-state magnetic resonance spectroscopy [43, 44]. These models have much in common and do indeed appear to reflect the fact that the structures of different fibrils are likely to be variations on a common theme [40]. It is also emerging that there may be some common and highly organised assemblies of amyloid protofilaments that are not simply extended threads or ribbons. It is clear, for example, that in some cases large closed loops can be formed [45, 46, 47], and there may be specific types of relatively small spherical or ‘doughnut’ shaped structures that can result in at least some circumstances (see below).
…..
The similarity of some early amyloid aggregates with the pores resulting from oligomerisation of bacterial toxins and pore-forming eukaryotic proteins (see below) also suggest that the basic mechanism of protein aggregation into amyloid structures may not only be associated with diseases but in some cases could result in species with functional significance. Recent evidence indicates that a variety of micro-organisms may exploit the controlled aggregation of specific proteins (or their precursors) to generate functional structures. Examples include bacterial curli [52] and proteins of the interior fibre cells of mammalian ocular lenses, whose β-sheet arrays seem to be organised in an amyloid-like supramolecular order [53]. In this case the inherent stability of amyloid-like protein structure may contribute to the long-term structural integrity and transparency of the lens. Recently it has been hypothesised that amyloid-like aggregates of serum amyloid A found in secondary amyloidoses following chronic inflammatory diseases protect the host against bacterial infections by inducing lysis of bacterial cells [54]. One particularly interesting example is a ‘misfolded’ form of the milk protein α-lactalbumin that is formed at low pH and trapped by the presence of specific lipid molecules [55]. This form of the protein has been reported to trigger apoptosis selectively in tumour cells providing evidence for its importance in protecting infants from certain types of cancer [55]. ….
Amyloid formation is a generic property of polypeptide chains ….
It is clear that the presence of different side chains can influence the details of amyloid structures, particularly the assembly of protofibrils, and that they give rise to the variations on the common structural theme discussed above. More fundamentally, the composition and sequence of a peptide or protein affects profoundly its propensity to form amyloid structures under given conditions (see below).
Because the formation of stable protein aggregates of amyloid type does not normally occur in vivo under physiological conditions, it is likely that the proteins encoded in the genomes of living organisms are endowed with structural adaptations that mitigate against aggregation under these conditions. A recent survey involving a large number of structures of β-proteins highlights several strategies through which natural proteins avoid intermolecular association of β-strands in their native states [65]. Other surveys of protein databases indicate that nature disfavours sequences of alternating polar and nonpolar residues, as well as clusters of several consecutive hydrophobic residues, both of which enhance the tendency of a protein to aggregate prior to becoming completely folded [66, 67].
……
Precursors of amyloid fibrils can be toxic to cells
It was generally assumed until recently that the proteinaceous aggregates most toxic to cells are likely to be mature amyloid fibrils, the form of aggregates that have been commonly detected in pathological deposits. It therefore appeared probable that the pathogenic features underlying amyloid diseases are a consequence of the interaction with cells of extracellular deposits of aggregated material. As well as forming the basis for understanding the fundamental causes of these diseases, this scenario stimulated the exploration of therapeutic approaches to amyloidoses that focused mainly on the search for molecules able to impair the growth and deposition of fibrillar forms of aggregated proteins. ….
Structural basis and molecular features of amyloid toxicity
The presence of toxic aggregates inside or outside cells can impair a number of cell functions that ultimately lead to cell death by an apoptotic mechanism [95, 96]. Recent research suggests, however, that in most cases initial perturbations to fundamental cellular processes underlie the impairment of cell function induced by aggregates of disease-associated polypeptides. Many pieces of data point to a central role of modifications to the intracellular redox status and free Ca2+ levels in cells exposed to toxic aggregates [45, 89, 97, 98, 99, 100, 101]. A modification of the intracellular redox status in such cells is associated with a sharp increase in the quantity of reactive oxygen species (ROS) that is reminiscent of the oxidative burst by which leukocytes destroy invading foreign cells after phagocytosis. In addition, changes have been observed in reactive nitrogen species, lipid peroxidation, deregulation of NO metabolism [97], protein nitrosylation [102] and upregulation of heme oxygenase-1, a specific marker of oxidative stress [103]. ….
Results have recently been reported concerning the toxicity towards cultured cells of aggregates of poly(Q) peptides which argues against a disease mechanism based on specific toxic features of the aggregates. These results indicate that there is a close relationship between the toxicity of proteins with poly(Q) extensions and their nuclear localisation. In addition they support the hypotheses that the toxicity of poly(Q) aggregates can be a consequence of altered interactions with nuclear coactivator or corepressor molecules including p53, CBP, Sp1 and TAF130 or of the interaction with transcription factors and nuclear coactivators, such as CBP, endowed with short poly(Q) stretches ([95] and references therein)…..
Concluding remarks
The data reported in the past few years strongly suggest that the conversion of normally soluble proteins into amyloid fibrils and the toxicity of small aggregates appearing during the early stages of the formation of the latter are common or generic features of polypeptide chains. Moreover, the molecular basis of this toxicity also appears to display common features between the different systems that have so far been studied. The ability of many, perhaps all, natural polypeptides to ‘misfold’ and convert into toxic aggregates under suitable conditions suggests that one of the most important driving forces in the evolution of proteins must have been the negative selection against sequence changes that increase the tendency of a polypeptide chain to aggregate. Nevertheless, as protein folding is a stochastic process, and no such process can be completely infallible, misfolded proteins or protein folding intermediates in equilibrium with the natively folded molecules must continuously form within cells. Thus mechanisms to deal with such species must have co-evolved with proteins. Indeed, it is clear that misfolding, and the associated tendency to aggregate, is kept under control by molecular chaperones, which render the resulting species harmless assisting in their refolding, or triggering their degradation by the cellular clearance machinery [166, 167, 168, 169, 170, 171, 172, 173, 175, 177, 178].
Misfolded and aggregated species are likely to owe their toxicity to the exposure on their surfaces of regions of proteins that are buried in the interior of the structures of the correctly folded native states. The exposure of large patches of hydrophobic groups is likely to be particularly significant as such patches favour the interaction of the misfolded species with cell membranes [44, 83, 89, 90, 91, 93]. Interactions of this type are likely to lead to the impairment of the function and integrity of the membranes involved, giving rise to a loss of regulation of the intracellular ion balance and redox status and eventually to cell death. In addition, misfolded proteins undoubtedly interact inappropriately with other cellular components, potentially giving rise to the impairment of a range of other biological processes. Under some conditions the intracellular content of aggregated species may increase directly, due to an enhanced propensity of incompletely folded or misfolded species to aggregate within the cell itself. This could occur as the result of the expression of mutational variants of proteins with decreased stability or cooperativity or with an intrinsically higher propensity to aggregate. It could also occur as a result of the overproduction of some types of protein, for example, because of other genetic factors or other disease conditions, or because of perturbations to the cellular environment that generate conditions favouring aggregation, such as heat shock or oxidative stress. Finally, the accumulation of misfolded or aggregated proteins could arise from the chaperone and clearance mechanisms becoming overwhelmed as a result of specific mutant phenotypes or of the general effects of ageing [173, 174].
The topics discussed in this review not only provide a great deal of evidence for the ‘new view’ that proteins have an intrinsic capability of misfolding and forming structures such as amyloid fibrils but also suggest that the role of molecular chaperones is even more important than was thought in the past. The role of these ubiquitous proteins in enhancing the efficiency of protein folding is well established [185]. It could well be that they are at least as important in controlling the harmful effects of misfolded or aggregated proteins as in enhancing the yield of functional molecules.
Nutritional Status is Associated with Faster Cognitive Decline and Worse Functional Impairment in the Progression of Dementia: The Cache County Dementia Progression Study1
Nutritional status may be a modifiable factor in the progression of dementia. We examined the association of nutritional status and rate of cognitive and functional decline in a U.S. population-based sample. Study design was an observational longitudinal study with annual follow-ups up to 6 years of 292 persons with dementia (72% Alzheimer’s disease, 56% female) in Cache County, UT using the Mini-Mental State Exam (MMSE), Clinical Dementia Rating Sum of Boxes (CDR-sb), and modified Mini Nutritional Assessment (mMNA). mMNA scores declined by approximately 0.50 points/year, suggesting increasing risk for malnutrition. Lower mMNA score predicted faster rate of decline on the MMSE at earlier follow-up times, but slower decline at later follow-up times, whereas higher mMNA scores had the opposite pattern (mMNA by time β= 0.22, p = 0.017; mMNA by time2 β= –0.04, p = 0.04). Lower mMNA score was associated with greater impairment on the CDR-sb over the course of dementia (β= 0.35, p < 0.001). Assessment of malnutrition may be useful in predicting rates of progression in dementia and may provide a target for clinical intervention.
Shared Genetic Risk Factors for Late-Life Depression and Alzheimer’s Disease
Background: Considerable evidence has been reported for the comorbidity between late-life depression (LLD) and Alzheimer’s disease (AD), both of which are very common in the general elderly population and represent a large burden on the health of the elderly. The pathophysiological mechanisms underlying the link between LLD and AD are poorly understood. Because both LLD and AD can be heritable and are influenced by multiple risk genes, shared genetic risk factors between LLD and AD may exist. Objective: The objective is to review the existing evidence for genetic risk factors that are common to LLD and AD and to outline the biological substrates proposed to mediate this association. Methods: A literature review was performed. Results: Genetic polymorphisms of brain-derived neurotrophic factor, apolipoprotein E, interleukin 1-beta, and methylenetetrahydrofolate reductase have been demonstrated to confer increased risk to both LLD and AD by studies examining either LLD or AD patients. These results contribute to the understanding of pathophysiological mechanisms that are common to both of these disorders, including deficits in nerve growth factors, inflammatory changes, and dysregulation mechanisms involving lipoprotein and folate. Other conflicting results have also been reviewed, and few studies have investigated the effects of the described polymorphisms on both LLD and AD. Conclusion: The findings suggest that common genetic pathways may underlie LLD and AD comorbidity. Studies to evaluate the genetic relationship between LLD and AD may provide insights into the molecular mechanisms that trigger disease progression as the population ages.
Association of Vitamin B12, Folate, and Sulfur Amino Acids With Brain Magnetic Resonance Imaging Measures in Older Adults: A Longitudinal Population-Based Study
Importance Vitamin B12, folate, and sulfur amino acids may be modifiable risk factors for structural brain changes that precede clinical dementia.
Objective To investigate the association of circulating levels of vitamin B12, red blood cell folate, and sulfur amino acids with the rate of total brain volume loss and the change in white matter hyperintensity volume as measured by fluid-attenuated inversion recovery in older adults.
Design, Setting, and Participants The magnetic resonance imaging subsample of the Swedish National Study on Aging and Care in Kungsholmen, a population-based longitudinal study in Stockholm, Sweden, was conducted in 501 participants aged 60 years or older who were free of dementia at baseline. A total of 299 participants underwent repeated structural brain magnetic resonance imaging scans from September 17, 2001, to December 17, 2009.
Main Outcomes and Measures The rate of brain tissue volume loss and the progression of total white matter hyperintensity volume.
Results In the multi-adjusted linear mixed models, among 501 participants (300 women [59.9%]; mean [SD] age, 70.9 [9.1] years), higher baseline vitamin B12 and holotranscobalamin levels were associated with a decreased rate of total brain volume loss during the study period: for each increase of 1 SD, β (SE) was 0.048 (0.013) for vitamin B12 (P < .001) and 0.040 (0.013) for holotranscobalamin (P = .002). Increased total homocysteine levels were associated with faster rates of total brain volume loss in the whole sample (β [SE] per 1-SD increase, –0.035 [0.015]; P = .02) and with the progression of white matter hyperintensity among participants with systolic blood pressure greater than 140 mm Hg (β [SE] per 1-SD increase, 0.000019 [0.00001]; P = .047). No longitudinal associations were found for red blood cell folate and other sulfur amino acids.
Conclusions and Relevance This study suggests that both vitamin B12 and total homocysteine concentrations may be related to accelerated aging of the brain. Randomized clinical trials are needed to determine the importance of vitamin B12supplementation on slowing brain aging in older adults.
Notes from Kurzweill
This vitamin stops the aging process in organs, say Swiss researchers
A potential breakthrough for regenerative medicine, pending further studies
Improved muscle stem cell numbers and muscle function in NR-treated aged mice: Newly regenerated muscle fibers 7 days after muscle damage in aged mice (left: control group; right: fed NR). (Scale bar = 50 μm). (credit: Hongbo Zhang et al./Science) http://www.kurzweilai.net/images/improved-muscle-fibers.png
EPFL researchers have restored the ability of mice organs to regenerate and extend life by simply administering nicotinamide riboside (NR) to them.
NR has been shown in previous studies to be effective in boosting metabolism and treating a number of degenerative diseases. Now, an article by PhD student Hongbo Zhang published in Science also describes the restorative effects of NR on the functioning of stem cells for regenerating organs.
As in all mammals, as mice age, the regenerative capacity of certain organs (such as the liver and kidneys) and muscles (including the heart) diminishes. Their ability to repair them following an injury is also affected. This leads to many of the disorders typical of aging.
Mitochondria —> stem cells —> organs
To understand how the regeneration process deteriorates with age, Zhang teamed up with colleagues from ETH Zurich, the University of Zurich, and universities in Canada and Brazil. By using several biomarkers, they were able to identify the molecular chain that regulates how mitochondria — the “powerhouse” of the cell — function and how they change with age. “We were able to show for the first time that their ability to function properly was important for stem cells,” said Auwerx.
Under normal conditions, these stem cells, reacting to signals sent by the body, regenerate damaged organs by producing new specific cells. At least in young bodies. “We demonstrated that fatigue in stem cells was one of the main causes of poor regeneration or even degeneration in certain tissues or organs,” said Zhang.
How to revitalize stem cells
Which is why the researchers wanted to “revitalize” stem cells in the muscles of elderly mice. And they did so by precisely targeting the molecules that help the mitochondria to function properly. “We gave nicotinamide riboside to 2-year-old mice, which is an advanced age for them,” said Zhang.
“This substance, which is close to vitamin B3, is a precursor of NAD+, a molecule that plays a key role in mitochondrial activity. And our results are extremely promising: muscular regeneration is much better in mice that received NR, and they lived longer than the mice that didn’t get it.”
Parallel studies have revealed a comparable effect on stem cells of the brain and skin. “This work could have very important implications in the field of regenerative medicine,” said Auwerx. This work on the aging process also has potential for treating diseases that can affect — and be fatal — in young people, like muscular dystrophy (myopathy).
So far, no negative side effects have been observed following the use of NR, even at high doses. But while it appears to boost the functioning of all cells, it could include pathological ones, so further in-depth studies are required.
Abstract of NAD+ repletion improves mitochondrial and stem cell function and enhances life span in mice
Adult stem cells (SCs) are essential for tissue maintenance and regeneration yet are susceptible to senescence during aging. We demonstrate the importance of the amount of the oxidized form of cellular nicotinamide adenine dinucleotide (NAD+) and its impact on mitochondrial activity as a pivotal switch to modulate muscle SC (MuSC) senescence. Treatment with the NAD+ precursor nicotinamide riboside (NR) induced the mitochondrial unfolded protein response (UPRmt) and synthesis of prohibitin proteins, and this rejuvenated MuSCs in aged mice. NR also prevented MuSC senescence in the Mdx mouse model of muscular dystrophy. We furthermore demonstrate that NR delays senescence of neural SCs (NSCs) and melanocyte SCs (McSCs), and increased mouse lifespan. Strategies that conserve cellular NAD+ may reprogram dysfunctional SCs and improve lifespan in mammals.
Discriminating the gene target of a distal regulatory element from other nearby transcribed genes is a challenging problem with the potential to illuminate the causal underpinnings of complex diseases. We present TargetFinder, a computational method that reconstructs regulatory landscapes from diverse features along the genome. The resulting models accurately predict individual enhancer–promoter interactions across multiple cell lines with a false discovery rate up to 15 times smaller than that obtained using the closest gene. By evaluating the genomic features driving this accuracy, we uncover interactions between structural proteins, transcription factors, epigenetic modifications, and transcription that together distinguish interacting from non-interacting enhancer–promoter pairs. Most of this signature is not proximal to the enhancers and promoters but instead decorates the looping DNA. We conclude that complex but consistent combinations of marks on the one-dimensional genome encode the three-dimensional structure of fine-scale regulatory interactions.
Gene Editing with CRISPR gets Crisper, Volume 2 (Volume Two: Latest in Genomics Methodologies for Therapeutics: Gene Editing, NGS and BioInformatics, Simulations and the Genome Ontology), Part 2: CRISPR for Gene Editing and DNA Repair
Gene Editing with CRISPR gets Crisper
Curators: Larry H. Bernstein, MD, FCAP and Aviva Lev-Ari, PhD, RN
CRISPR Moves from Butchery to Surgery
More Genomes Are Going Under the CRISPR Knife, So Surgical Standards Are Rising
The Dharmacon subsidary of GE Healthcare provides the Edit-R Lentiviral Gene Engineering platform. It is based on the natural S. pyrogenes system, but unlike that system, which uses a single guide RNA (sgRNA), the platform uses two component RNAs, a gene-specific CRISPR RNA (crRNA) and a universal trans-activating crRNA (tracrRNA). Once hybridized to the universal tracrRNA (blue), the crRNA (green) directs the Cas9 nuclease to a specific genomic region to induce a double- strand break.
Scientists recently convened at the CRISPR Precision Gene Editing Congress, held in Boston, to discuss the new technology. As with any new technique, scientists have discovered that CRISPR comes with its own set of challenges, and the Congress focused its discussion around improving specificity, efficiency, and delivery.
In the naturally occurring system, CRISPR-Cas9 works like a self-vaccination in the bacterial immune system by targeting and cleaving viral DNA sequences stored from previous encounters with invading phages. The endogenous system uses two RNA elements, CRISPR RNA (crRNA) and trans-activating RNA (tracrRNA), which come together and guide the Cas9 nuclease to the target DNA.
Early publications that demonstrated CRISPR gene editing in mammalian cells combined the crRNA and tracrRNA sequences to form one long transcript called asingle-guide RNA (sgRNA). However, an alternative approach is being explored by scientists at the Dharmacon subsidiary of GE Healthcare. These scientists have a system that mimics the endogenous system through a synthetic two-component approach thatpreserves individual crRNA and tracrRNA. The tracrRNA is universal to any gene target or species; the crRNA contains the information needed to target the gene of interest.
Predesigned Guide RNAs
In contrast to sgRNAs, which are generated through either in vitro transcription of a DNA template or a plasmid-based expression system, synthetic crRNA and tracrRNA eliminate the need for additional cloning and purification steps. The efficacy of guide RNA (gRNA), whether delivered as a sgRNA or individual crRNA and tracrRNA, depends not only on DNA binding, but also on the generation of an indel that will deliver the coup de grâce to gene function.
“Almost all of the gRNAs were able to create a break in genomic DNA,” said Louise Baskin, senior product manager at Dharmacon. “But there was a very wide range in efficiency and in creating functional protein knock-outs.”
To remove the guesswork from gRNA design, Dharmacon developed an algorithm to predict gene knockout efficiency using wet-lab data. They also incorporated specificity as a component of their algorithm, using a much more comprehensive alignment tool to predict potential off-target effects caused by mismatches and bulges often missed by other alignment tools. Customers can enter their target gene to access predesigned gRNAs as either two-component RNAs or lentiviral sgRNA vectors for multiple applications.
“We put time and effort into our algorithm to ensure that our guide RNAs are not only functional but also highly specific,” asserts Baskin. “As a result, customers don’t have to do any design work.”
MilliporeSigma’s CRISPR Epigenetic Activator is based on fusion of a nuclease-deficient Cas9 (dCas9) to the catalytic histone acetyltransferase (HAT) core domain of the human E1A-associated protein p300. This technology allows researchers to target specific DNA regions or gene sequences. Researchers can localize epigenetic changes to their target of interest and see the effects of those changes in gene expression.
Knockout experiments are a powerful tool for analyzing gene function. However, for researchers who want to introduce DNA into the genome, guide design, donor DNA selection, and Cas9 activity are paramount to successful DNA integration.MilliporeSigma offers two formats for donor DNA: double-stranded DNA (dsDNA) plasmids and single-stranded DNA (ssDNA) oligonucleotides. The most appropriate format depends on cell type and length of the donor DNA. “There are some cell types that have immune responses to dsDNA,” said Gregory Davis, Ph.D., R&D manager, MilliporeSigma.
The ssDNA format can save researchers time and money, but it has a limited carrying capacity of approximately 120 base pairs.In addition to selecting an appropriate donor DNA format, controlling where, how, and when the Cas9 enzyme cuts can affect gene-editing efficiency. Scientists are playing tug-of-war, trying to pull cells toward the preferred homology-directed repair (HDR) and away from the less favored nonhomologous end joining (NHEJ) repair mechanism.One method to achieve this modifies the Cas9 enzyme to generate a nickase that cuts only one DNA strand instead of creating a double-strand break. Accordingly, MilliporeSigma has created a Cas9 paired-nickase system that promotes HDR, while also limiting off-target effects and increasing the number of sequences available for site-dependent gene modifications, such as disease-associated single nucleotide polymorphisms (SNPs).“The best thing you can do is to cut as close to the SNP as possible,” advised Dr. Davis. “As you move the double-stranded break away from the site of mutation you get an exponential drop in the frequency of recombination.”
Ribonucleo-protein Complexes
Another strategy to improve gene-editing efficiency, developed by Thermo Fisher, involves combining purified Cas9 protein with gRNA to generate a stable ribonucleoprotein (RNP) complex. In contrast to plasmid- or mRNA-based formats, which require transcription and/or translation, the Cas9 RNP complex cuts DNA immediately after entering the cell. Rapid clearance of the complex from the cell helps to minimize off-target effects, and, unlike a viral vector, the transient complex does not introduce foreign DNA sequences into the genome.
To deliver their Cas9 RNP complex to cells, Thermo Fisher has developed a lipofectamine transfection reagent called CRISPRMAX. “We went back to the drawing board with our delivery, screened a bunch of components, and got a brand-new, fully optimized lipid nanoparticle formulation,” explained Jon Chesnut, Ph.D., the company’s senior director of synthetic biology R&D. “The formulation is specifically designed for delivering the RNP to cells more efficiently.”
Besides the reagent and the formulation, Thermo Fisher has also developed a range of gene-editing tools. For example, it has introduced the Neon® transfection system for delivering DNA, RNA, or protein into cells via electroporation. Dr. Chesnut emphasized the company’s focus on simplifying complex workflows by optimizing protocols and pairing everything with the appropriate up- and downstream reagents.
From Mammalian Cells to Microbes
One of the first sources of CRISPR technology was the Feng Zhang laboratory at the Broad Institute, which counted among its first licensees a company called GenScript. This company offers a gene-editing service called GenCRISPR™ to establish mammalian cell lines with CRISPR-derived gene knockouts.
“There are a lot of challenges with mammalian cells, and each cell line has its own set of issues,” said Laura Geuss, a marketing specialist at GenScript. “We try to offer a variety of packages that can help customers who have difficult-to-work-with cells.” These packages include both viral-based and transient transfection techniques.
However, the most distinctive service offered by GenScript is its microbial genome-editing service for bacteria (Escherichia coli) and yeast (Saccharomyces cerevisiae). The company’s strategy for gene editing in bacteria can enable seamless knockins, knockouts, or gene replacements by combining CRISPR with lambda red recombineering. Traditionally one of the most effective methods for gene editing in microbes, recombineering allows editing without restriction enzymes through in vivo homologous recombination mediated by a phage-based recombination system such as lambda red.
On its own, lambda red technology cannot target multiple genes, but when paired with CRISPR, it allows the editing of multiple genes with greater efficiency than is possible with CRISPR alone, as the lambda red proteins help repair double-strand breaks in E. coli. The ability to knockout different gene combinations makes Genscript’s microbial editing service particularly well suited for the optimization of metabolic pathways.
Pooled and Arrayed Library Strategies
Scientists are using CRISPR technology for applications such as metabolic engineering and drug development. Yet another application area benefitting from CRISPR technology is cancer research. Here, the use of pooled CRISPR libraries is becoming commonplace. Pooled CRISPR libraries can help detect mutations that affect drug resistance, and they can aid in patient stratification and clinical trial design.
Pooled screening uses proliferation or viability as a phenotype to assess how genetic alterations, resulting from the application of a pooled CRISPR library, affect cell growth and death in the presence of a therapeutic compound. The enrichment or depletion of different gRNA populations is quantified using deep sequencing to identify the genomic edits that result in changes to cell viability.
MilliporeSigma provides pooled CRISPR libraries ranging from the whole human genome to smaller custom pools for these gene-function experiments. For pharmaceutical and biotech companies, Horizon Discovery offers a pooled screening service, ResponderSCREEN, which provides a whole-genome pooled screen to identify genes that confer sensitivity or resistance to a compound. This service is comprehensive, taking clients from experimental design all the way through to suggestions for follow-up studies.
Horizon Discovery maintains a Research Biotech business unit that is focused on target discovery and enabling translational medicine in oncology. “Our internal backbone gives us the ability to provide expert advice demonstrated by results,” said Jon Moore, Ph.D., the company’s CSO.
In contrast to a pooled screen, where thousands of gRNA are combined in one tube, an arrayed screen applies one gRNA per well, removing the need for deep sequencing and broadening the options for different endpoint assays. To establish and distribute a whole-genome arrayed lentiviral CRISPR library, MilliporeSigma partnered with the Wellcome Trust Sanger Institute. “This is the first and only arrayed CRISPR library in the world,” declared Shawn Shafer, Ph.D., functional genomics market segment manager, MilliporeSigma. “We were really proud to partner with Sanger on this.”
Pooled and arrayed screens are powerful tools for studying gene function. The appropriate platform for an experiment, however, will be determined by the desired endpoint assay.
The QX200 Droplet Digital PCR System from Bio-Rad Laboratories can provide researchers with an absolute measure of target DNA molecules for EvaGreen or probe-based digital PCR applications. The system, which can provide rapid, low-cost, ultra-sensitive quantification of both NHEJ- and HDR-editing events, consists of two instruments, the QX200 Droplet Generator and the QX200 Droplet Reader, and their associated consumables.
Finally, one last challenge for CRISPR lies in the detection and quantification of changes made to the genome post-editing. Conventional methods for detecting these alterations include gel methods and next-generation sequencing. While gel methods lack sensitivity and scalability, next-generation sequencing is costly and requires intensive bioinformatics.
To address this gap, Bio-Rad Laboratories developed a set of assay strategies to enable sensitive and precise edit detection with its Droplet Digital PCR (ddPCR) technology. The platform is designed to enable absolute quantification of nucleic acids with high sensitivity, high precision, and short turnaround time through massive droplet partitioning of samples.
Using a validated assay, a typical ddPCR experiment takes about five to six hours to complete. The ddPCR platform enables detection of rare mutations, and publications have reported detection of precise edits at a frequency of <0.05%, and of NHEJ-derived indels at a frequency as low as 0.1%. In addition to quantifying precise edits, indels, and computationally predicted off-target mutations, ddPCR can also be used to characterize the consequences of edits at the RNA level.
According to a recently published Science paper, the laboratory of Charles A. Gersbach, Ph.D., at Duke University used ddPCR in a study of muscle function in a mouse model of Duchenne muscular dystrophy. Specifically, ddPCR was used to assess the efficiency of CRISPR-Cas9 in removing the mutated exon 23 from the dystrophin gene. (Exon 23 deletion by CRISPR-Cas9 resulted in expression of the modified dystrophin gene and significant enhancement of muscle force.)
Quantitative ddPCR showed that exon 23 was deleted in ~2% of all alleles from the whole-muscle lysate. Further ddPCR studies found that 59% of mRNA transcripts reflected the deletion.
“There’s an overarching idea that the genome-editing field is moving extremely quickly, and for good reason,” asserted Jennifer Berman, Ph.D., staff scientist, Bio-Rad Laboratories. “There’s a lot of exciting work to be done, but detection and quantification of edits can be a bottleneck for researchers.”
The gene-editing field is moving quickly, and new innovations are finding their way into the laboratory as researchers lay the foundation for precise, well-controlled gene editing with CRISPR.
Researchers utilized a systems biology approach to develop new methods to assess drug sensitivity in cells. [The Institute for Systems Biology]
Understanding how cells respond and proliferate in the presence of anticancer compounds has been the foundation of drug discovery ideology for decades. Now, a new study from scientists at Vanderbilt University casts significant suspicion on the primary method used to test compounds for anticancer activity in cells—instilling doubt on methods employed by the entire scientific enterprise and pharmaceutical industry to discover new cancer drugs.
“More than 90% of candidate cancer drugs fail in late-stage clinical trials, costing hundreds of millions of dollars,” explained co-senior author Vito Quaranta, M.D., director of the Quantitative Systems Biology Center at Vanderbilt. “The flawed in vitro drug discovery metric may not be the only responsible factor, but it may be worth pursuing an estimate of its impact.”
The Vanderbilt investigators have developed what they believe to be a new metric for evaluating a compound’s effect on cell proliferation—called the DIP (drug-induced proliferation) rate—that overcomes the flawed bias in the traditional method.
The findings from this study were published recently in Nature Methods in an article entitled “An Unbiased Metric of Antiproliferative Drug Effect In Vitro.”
For more than three decades, researchers have evaluated the ability of a compound to kill cells by adding the compound in vitro and counting how many cells are alive after 72 hours. Yet, proliferation assays that measure cell number at a single time point don’t take into account the bias introduced by exponential cell proliferation, even in the presence of the drug.
“Cells are not uniform, they all proliferate exponentially, but at different rates,” Dr. Quaranta noted. “At 72 hours, some cells will have doubled three times and others will not have doubled at all.”
Dr. Quaranta added that drugs don’t all behave the same way on every cell line—for example, a drug might have an immediate effect on one cell line and a delayed effect on another.
The research team decided to take a systems biology approach, a mixture of experimentation and mathematical modeling, to demonstrate the time-dependent bias in static proliferation assays and to develop the time-independent DIP rate metric.
“Systems biology is what really makes the difference here,” Dr. Quaranta remarked. “It’s about understanding cells—and life—as dynamic systems.”This new study is of particular importance in light of recent international efforts to generate data sets that include the responses of thousands of cell lines to hundreds of compounds. Using the
Cancer Cell Line Encyclopedia (CCLE) and
Genomics of Drug Sensitivity in Cancer (GDSC) databases
will allow drug discovery scientists to include drug response data along with genomic and proteomic data that detail each cell line’s molecular makeup.
“The idea is to look for statistical correlations—these particular cell lines with this particular makeup are sensitive to these types of compounds—to use these large databases as discovery tools for new therapeutic targets in cancer,” Dr. Quaranta stated. “If the metric by which you’ve evaluated the drug sensitivity of the cells is wrong, your statistical correlations are basically no good.”
The Vanderbilt team evaluated the responses from four different melanoma cell lines to the drug vemurafenib, currently used to treat melanoma, with the standard metric—used for the CCLE and GDSC databases—and with the DIP rate. In one cell line, they found a glaring disagreement between the two metrics.
“The static metric says that the cell line is very sensitive to vemurafenib. However, our analysis shows this is not the case,” said co-lead study author Leonard Harris, Ph.D., a systems biology postdoctoral fellow at Vanderbilt. “A brief period of drug sensitivity, quickly followed by rebound, fools the static metric, but not the DIP rate.”
Dr. Quaranta added that the findings “suggest we should expect melanoma tumors treated with this drug to come back, and that’s what has happened, puzzling investigators. DIP rate analyses may help solve this conundrum, leading to better treatment strategies.”
The researchers noted that using the DIP rate is possible because of advances in automation, robotics, microscopy, and image processing. Moreover, the DIP rate metric offers another advantage—it can reveal which drugs are truly cytotoxic (cell killing), rather than merely cytostatic (cell growth inhibiting). Although cytostatic drugs may initially have promising therapeutic effects, they may leave tumor cells alive that then have the potential to cause the cancer to recur.
The Vanderbilt team is currently in the process of identifying commercial entities that can further refine the software and make it widely available to the research community to inform drug discovery.
An unbiased metric of antiproliferative drug effect in vitro
In vitro cell proliferation assays are widely used in pharmacology, molecular biology, and drug discovery. Using theoretical modeling and experimentation, we show that current metrics of antiproliferative small molecule effect suffer from time-dependent bias, leading to inaccurate assessments of parameters such as drug potency and efficacy. We propose the drug-induced proliferation (DIP) rate, the slope of the line on a plot of cell population doublings versus time, as an alternative, time-independent metric.
Researchers develop a technique to direct chromosome recombination with CRISPR/Cas9, allowing high-resolution genetic mapping of phenotypic traits in yeast.
Researchers used CRISPR/Cas9 to make a targeted double-strand break (DSB) in one arm of a yeast chromosome labeled with a green fluorescent protein (GFP) gene. A within-cell mechanism called homologous repair (HR) mends the broken arm using its homolog, resulting in a recombined region from the site of the break to the chromosome tip. When this cell divides by mitosis, each daughter cell will contain a homozygous section in an outcome known as “loss of heterozygosity” (LOH). One of the daughter cells is detectable because, due to complete loss of the GFP gene, it will no longer be fluorescent.REPRINTED WITH PERMISSION FROM M.J. SADHU ET AL., SCIENCE
When mapping phenotypic traits to specific loci, scientists typically rely on the natural recombination of chromosomes during meiotic cell division in order to infer the positions of responsible genes. But recombination events vary with species and chromosome region, giving researchers little control over which areas of the genome are shuffled. Now, a team at the University of California, Los Angeles (UCLA), has found a way around these problems by using CRISPR/Cas9 to direct targeted recombination events during mitotic cell division in yeast. The team described its technique today (May 5) in Science.
“Current methods rely on events that happen naturally during meiosis,” explained study coauthor Leonid Kruglyak of UCLA. “Whatever rate those events occur at, you’re kind of stuck with. Our idea was that using CRISPR, we can generate those events at will, exactly where we want them, in large numbers, and in a way that’s easy for us to pull out the cells in which they happened.”
Generally, researchers use coinheritance of a trait of interest with specific genetic markers—whose positions are known—to figure out what part of the genome is responsible for a given phenotype. But the procedure often requires impractically large numbers of progeny or generations to observe the few cases in which coinheritance happens to be disrupted informatively. What’s more, the resolution of mapping is limited by the length of the smallest sequence shuffled by recombination—and that sequence could include several genes or gene variants.
“Once you get down to that minimal region, you’re done,” said Kruglyak. “You need to switch to other methods to test every gene and every variant in that region, and that can be anywhere from challenging to impossible.”
But programmable, DNA-cutting champion CRISPR/Cas9 offered an alternative. During mitotic—rather than meiotic—cell division, rare, double-strand breaks in one arm of a chromosome preparing to split are sometimes repaired by a mechanism called homologous recombination. This mechanism uses the other chromosome in the homologous pair to replace the sequence from the break down to the end of the broken arm. Normally, such mitotic recombination happens so rarely as to be impractical for mapping purposes. With CRISPR/Cas9, however, the researchers found that they could direct double-strand breaks to any locus along a chromosome of interest (provided it was heterozygous—to ensure that only one of the chromosomes would be cut), thus controlling the sites of recombination.
Combining this technique with a signal of recombination success, such as a green fluorescent protein (GFP) gene at the tip of one chromosome in the pair, allowed the researchers to pick out cells in which recombination had occurred: if the technique failed, both daughter cells produced by mitotic division would be heterozygous, with one copy of the signal gene each. But if it succeeded, one cell would end up with two copies, and the other cell with none—an outcome called loss of heterozygosity.
“If we get loss of heterozygosity . . . half the cells derived after that loss of heterozygosity event won’t have GFP anymore,” study coauthor Meru Sadhu of UCLA explained. “We search for these cells that don’t have GFP out of the general population of cells.” If these non-fluorescent cells with loss of heterozygosity have the same phenotype as the parent for a trait of interest, then CRISPR/Cas9-targeted recombination missed the responsible gene. If the phenotype is affected, however, then the trait must be linked to a locus in the recombined, now-homozygous region, somewhere between the cut site and the GFP gene.
By systematically making cuts using CRISPR/Cas9 along chromosomes in a hybrid, diploid strain ofSaccharomyces cerevisiae yeast, picking out non-fluorescent cells, and then observing the phenotype, the UCLA team demonstrated that it could rapidly identify the phenotypic contribution of specific gene variants. “We can simply walk along the chromosome and at every [variant] position we can ask, does it matter for the trait we’re studying?” explained Kruglyak.
For example, the team showed that manganese sensitivity—a well-defined phenotypic trait in lab yeast—could be pinpointed using this method to a single nucleotide polymorphism (SNP) in a gene encoding the Pmr1 protein (a manganese transporter).
Jason Moffat, a molecular geneticist at the University of Toronto who was not involved in the work, toldThe Scientist that researchers had “dreamed about” exploiting these sorts of mechanisms for mapping purposes, but without CRISPR, such techniques were previously out of reach. Until now, “it hasn’t been so easy to actually make double-stranded breaks on one copy of a pair of chromosomes, and then follow loss of heterozygosity in mitosis,” he said, adding that he hopes to see the approach translated into human cell lines.
Applying the technique beyond yeast will be important, agreed cell and developmental biologist Ethan Bier of the University of California, San Diego, because chromosomal repair varies among organisms. “In yeast, they absolutely demonstrate the power of [this method],” he said. “We’ll just have to see how the technology develops in other systems that are going to be far less suited to the technology than yeast. . . . I would like to see it implemented in another system to show that they can get the same oomph out of it in, say, mammalian somatic cells.”
Kruglyak told The Scientist that work in higher organisms, though planned, is still in early stages; currently, his team is working to apply the technique to map loci responsible for trait differences between—rather than within—yeast species.
“We have a much poorer understanding of the differences across species,” Sadhu explained. “Except for a few specific examples, we’re pretty much in the dark there.”
Linkage and association studies have mapped thousands of genomic regions that contribute to phenotypic variation, but narrowing these regions to the underlying causal genes and variants has proven much more challenging. Resolution of genetic mapping is limited by the recombination rate. We developed a method that uses CRISPR to build mapping panels with targeted recombination events. We tested the method by generating a panel with recombination events spaced along a yeast chromosome arm, mapping trait variation, and then targeting a high density of recombination events to the region of interest. Using this approach, we fine-mapped manganese sensitivity to a single polymorphism in the transporter Pmr1. Targeting recombination events to regions of interest allows us to rapidly and systematically identify causal variants underlying trait differences.
Thank you, David, for the kind words and comments. We agree that the most immediate applications of the CRISPR-based recombination mapping will be in unicellular organisms and cell culture. We also think the method holds a lot of promise for research in multicellular organisms, although we did not mean to imply that it “will be an efficient mapping method for all multicellular organisms”. Every organism will have its own set of constraints as well as experimental tools that will be relevant when adapting a new technique. To best help experts working on these organisms, here are our thoughts on your questions.
You asked about mutagenesis during recombination. We Sanger sequenced 72 of our LOH lines at the recombination site and did not observe any mutations, as described in the supplementary materials. We expect the absence of mutagenesis is because we targeted heterozygous sites where the untargeted allele did not have a usable PAM site; thus, following LOH, the targeted site is no longer present and cutting stops. In your experiments you targeted sites that were homozygous; thus, following recombination, the CRISPR target site persisted, and continued cutting ultimately led to repair by NHEJ and mutagenesis.
As to the more general question of the optimal mapping strategies in different organisms, they will depend on the ease of generating and screening for editing events, the cost and logistics of maintaining and typing many lines, and generation time, among other factors. It sounds like in Drosophila today, your related approach of generating markers with CRISPR, and then enriching for natural recombination events that separate them, is preferable. In yeast, we’ve found the opposite to be the case. As you note, even in Drosophila, our approach may be preferable for regions with low or highly non-uniform recombination rates.
Finally, mapping in sterile interspecies hybrids should be straightforward for unicellular hybrids (of which there are many examples) and for cells cultured from hybrid animals or plants. For studies in hybrid multicellular organisms, we agree that driving mitotic recombination in the early embryo may be the most promising approach. Chimeric individuals with mitotic clones will be sufficient for many traits. Depending on the system, it may in fact be possible to generate diploid individuals with uniform LOH genotype, but this is certainly beyond the scope of our paper. The calculation of the number of lines assumes that the mapping is done in a single step; as you note in your earlier comment, mapping sequentially can reduce this number dramatically.
This is a lovely method and should find wide applicability in many settings, especially for microorganisms and cell lines. However, it is not clear that this approach will be, as implied by the discussion, an efficient mapping method for all multicellular organisms. I have performed similar experiments in Drosophila, focused on meiotic recombination, on a much smaller scale, and found that CRISPR-Cas9 can indeed generate targeted recombination at gRNA target sites. In every case I tested, I found that the recombination event was associated with a deletion at the gRNA site, which is probably unimportant for most mapping efforts, but may be a concern in some specific cases, for example for clinical applications. It would be interesting to know how often mutations occurred at the targeted gRNA site in this study.
The wider issue, however, is whether CRISPR-mediated recombination will be more efficient than other methods of mapping. After careful consideration of all the costs and the time involved in each of the steps for Drosophila, we have decided that targeted meiotic recombination using flanking visible markers will be, in most cases, considerably more efficient than CRISPR-mediated recombination. This is mainly due to the large expense of injecting embryos and the extensive effort and time required to screen injected animals for appropriate events. It is both cheaper and faster to generate markers (with CRISPR) and then perform a large meiotic recombination mapping experiment than it would be to generate the lines required for CRISPR-mediated recombination mapping. It is possible to dramatically reduce costs by, for example, mapping sequentially at finer resolution. But this approach would require much more time than marker-assisted mapping. If someone develops a rapid and cheap method of reliably introducing DNA into Drosophila embryos, then this calculus might change.
However, it is possible to imagine situations where CRISPR-mediated mapping would be preferable, even for Drosophila. For example, some genomic regions display extremely low or highly non-uniform recombination rates. It is possible that CRISPR-mediated mapping could provide a reasonable approach to fine mapping genes in these regions.
The authors also propose the exciting possibility that CRISPR-mediated loss of heterozygosity could be used to map traits in sterile species hybrids. It is not entirely obvious to me how this experiment would proceed and I hope the authors can illuminate me. If we imagine driving a recombination event in the early embryo (with maternal Cas9 from one parent and gRNA from a second parent), then at best we would end up with chimeric individuals carrying mitotic clones. I don’t think one could generate diploid animals where all cells carried the same loss of heterozygosity event. Even if we could, this experiment would require construction of a substantial number of stable transgenic lines expressing gRNAs. Mapping an ~20Mbp chromosome arm to ~10kb would require on the order of two-thousand transgenic lines. Not an undertaking to be taken lightly. It is already possible to perform similar tests (hemizygosity tests) using D. melanogaster deficiency lines in crosses with D. simulans, so perhaps CRISPR-mediated LOH could complement these deficiency screens for fine mapping efforts. But, at the moment, it is not clear to me how to do the experiment.
Pinchas Cohen led a team that identified tiny proteins that appear to play a role in controlling how the body ages. (Photo/Beth Newcomb)
A group of six newly discovered proteins may help to divulge secrets of how we age, potentially unlocking insights into diabetes, Alzheimer’s, cancer and other aging-related diseases.
The tiny proteins appear to play several big roles in our bodies’ cells, from decreasing the amount of damaging free radicals and controlling the rate at which cells die to boosting metabolism and helping tissues throughout the body respond better to insulin. The naturally occurring amounts of each protein decrease with age, leading researchers to believe that they play an important role in the aging process and the onset of diseases linked to older age.
The research team led by Pinchas Cohen, dean of the USC Davis School of Gerontology, identified the tiny proteins for the first time and observed their surprising origin from organelles in the cell called mitochondria and their game-changing roles in metabolism and cell survival. This latest finding builds upon prior research by Cohen and his team that uncovered two significant proteins, humanin and MOTS-c, hormones that appear to have significant roles in metabolism and diseases of aging.
Unlike most other proteins, humanin and MOTS-c are encoded in mitochondria, the structure within cells that produces energy from food, instead of in the cell’s nucleus where most genes are contained.
Key functions
Mitochondria have their own small collection of genes, which were once thought to play only minor roles within cells but now appear to have important functions throughout the body. Cohen’s team used computer analysis to see if the part of the mitochondrial genome that provides the code for humanin was coding for other proteins as well. The analysis uncovered the genes for six new proteins, which were dubbed small humanin-like peptides, or SHLPs, 1 through 6 (the name of this hardworking group of proteins is appropriately pronounced “schlep”).
After identifying the six SHLPs and successfully developing antibodies to test for several of them, the team examined both mouse tissues and human cells to determine their abundance in different organs as well as their functions. The proteins were distributed quite differently among organs, which suggests that the proteins have varying functions based on where they are in the body.
Of particular interest is SHLP 2, Cohen said. The protein appears to have profound insulin-sensitizing, anti-diabetic effects as well as potent neuro-protective activity that may emerge as a strategy to combat Alzheimer’s disease. He added that SHLP 6 is also intriguing, with a unique ability to promote cancer cell death and thus potentially target malignant diseases.
“Together with the previously identified mitochondrial peptides, the newly recognized SHLP family expands the understanding of the mitochondria as an intracellular signaling organelle that communicates with the rest of the body to regulate metabolism and cell fate,” Cohen said. “The findings are an important advance that will be ripe for rapid translation into drug development for diseases of aging.”
The study first appeared online in the journal Aging on April 10. Cohen’s research team included collaborators from the Albert Einstein College of Medicine; the findings have been licensed to the biotechnology company CohBar for possible drug development.
The research was supported by a Glenn Foundation Award and National Institutes of Health grants to Cohen (1P01AG034906, 1R01AG 034430, 1R01GM 090311, 1R01ES 020812) and an Ellison/AFAR postdoctoral fellowship to Kelvin Yen. Study authors Laura Cobb, Changhan Lee, Nir Barzilai and Pinchas Cohen are consultants and stockholders of CohBar Inc.
On a blazingly hot morning this past June, a half-dozen scientists convened in a hotel conference room in suburban Maryland for the dress rehearsal of what they saw as a landmark event in the history of aging research. In a few hours, the group would meet with officials at the U.S. Food and Drug Administration (FDA), a few kilometers away, to pitch an unprecedented clinical trial—nothing less than the first test of a drug to specifically target the process of human aging.
“We think this is a groundbreaking, perhaps paradigm-shifting trial,” said Steven Austad, chairman of biology at the University of Alabama, Birmingham, and scientific director of the American Federation for Aging Research (AFAR). After Austad’s brief introductory remarks, a scientist named Nir Barzilai tuned up his PowerPoint and launched into a practice run of the main presentation.
Barzilai is a former Israeli army medical officer and head of a well-known study of centenarians based at the Albert Einstein College of Medicine in the Bronx, New York. To anyone who has seen the ebullient scientist in his natural laboratory habitat, often in a short-sleeved shirt and always cracking jokes, he looked uncharacteristically kempt in a blue blazer and dress khakis. But his practice run kept hitting a historical speed bump. He had barely begun to explain the rationale for the trial when he mentioned, in passing, “lots of unproven, untested treatments under the category of anti-aging.” His colleagues pounced.
“Nir,” interrupted S. Jay Olshansky, a biodemographer of aging from the University of Illinois, Chicago. The phrase “anti-aging … has an association that is negative.”
“I wouldn’t dignify them by calling them ‘treatments,’” added Michael Pollak, director of cancer prevention at McGill University in Montreal, Canada. “They’re products.”
Barzilai, a 59-year-old with a boyish mop of gray hair, wore a contrite grin. “We know the FDA is concerned about this,” he conceded, and deleted the offensive phrase.
Then he proceeded to lay out the details of an ambitious clinical trial. The group—academics all—wanted to conduct a double-blind study of roughly 3000 elderly people; half would get a placebo and half would get an old (indeed, ancient) drug for type 2 diabetes called metformin, which has been shown to modify aging in some animal studies. Because there is still no accepted biomarker for aging, the drug’s success would be judged by an unusual standard—whether it could delay the development of several diseases whose incidence increases dramatically with age: cardiovascular disease, cancer, and cognitive decline, along with mortality. When it comes to these diseases, Barzilai is fond of saying, “aging is a bigger risk factor than all of the other factors combined.”
But the phrase “anti-aging” kept creeping into the rehearsal, and critics kept jumping in. “Okay,” Barzilai said with a laugh when it came up again. “Third time, the death penalty.”
The group’s paranoia about the term “anti-aging” captured both the audacity of the proposed trial and the cultural challenge of venturing into medical territory historically associated with charlatans and quacks. The metformin initiative, which Barzilai is generally credited with spearheading, is unusual by almost any standard of drug development. The people pushing for the trial are all academics, none from industry (although Barzilai is co-founder of a biotech company, CohBar Inc., that is working to develop drugs targeting age-related diseases). The trial would be sponsored by the nonprofit AFAR, not a pharmaceutical company. No one stood to make money if the drug worked, the scientists all claimed; indeed, metformin is not only generic, costing just a few cents a dose, but belongs to a class of drugs that has been part of the human apothecary for 500 years. Patient safety was unlikely to be an issue; millions of diabetics have taken metformin since the 1960s, and its generally mild side effects are well-known.
Finally, the metformin group insisted they didn’t need a cent of federal money to proceed (although they do intend to ask for some). Nor did they need formal approval from FDA to proceed. But they very much wanted the agency’s blessing. By recognizing the merit of such a trial, Barzilai believes, FDA would make aging itself a legitimate target for drug development.
By the time the scientists were done, the rehearsal—which was being filmed for a television documentary—had the feel of a pep rally. They spoke with unguarded optimism. “What we’re talking about here,” Olshansky said, “is a fundamental sea change in how we look at aging and disease.” To Austad, it is “the key, potentially, to saving the health care system.”
As the group piled into a van for the drive to FDA headquarters, there was more talk about setting precedents and opening doors. So it was a little disconcerting when Austad led the delegation up to the main entrance of FDA—and couldn’t get the door open. ……
Mitochondrial Peptides Found in a Preclinical Study Seen to Control Cell Metabolism
CohBar, a developer of mitochondria-based therapeutics, announced that preclinical research by its academic collaborators has found small humanin-like peptides (SHLPs) that can control metabolism and cell survival. The findings have implications for age-related diseases such as Alzheimer’s and cancer.
Researchers discovered the SHLPs by examining the genome of mitochondria with the help of a bioinformatics approach, which identified six peptides. The team then verified the presence of the factors and explored their function in laboratory animals.
CohBar, who have the exclusive license to develop SHLPs into therapeutics, works closely with its academic partners to explore the peptides in preclinical models.
While it was previously believed that mitochondria only have 37 genes, research has revealed that the mitochondrial genome is far more versatile, potentially harboring a multitude of new genes, which can encode peptides acting as cellular signaling factors. The peptides, it has turned out, have shown neuroprotective and anti-inflammatory effects, and act to protect cells in disease-modifying ways in preclinical models of aging.
CohBar’s goal is to bring these peptides to the market as therapies for age-related diseases, such as obesity, type 2 diabetes, cancer, atherosclerosis and neurodegenerative disorders.
“Together with the previously described mitochondrial-derived peptides humanin and MOTS-c, the SHLP family expands our understanding of the role that these peptides play in intracellular signaling throughout the body to regulate both metabolism and cell survival,” Pinchas Cohen, dean of the USC Leonard Davis School of Gerontology, founder and director of CohBar, and the study’s senior author, said in a press release. “These findings further illustrate the enormous potential that mitochondria-based therapeutics could have on treating age-associated diseases like Alzheimer’s and cancer.”
“The pre-clinical evidence continues to confirm that these peptides represent a new class of naturally occurring metabolic regulators,” added Simon Allen, CohBar’s CEO. “They form the foundation of our pipeline of first-in-class treatments for age-related diseases, and we are committed to rapidly advancing them through pre-clinical and clinical activities as we move forward.”
Naturally occurring mitochondrial-derived peptides are age-dependent regulators of apoptosis, insulin sensitivity, and inflammatory markers
Laura J. Cobb1,5, Changhan Lee2, Jialin Xiao2, Kelvin Yen2, Richard G. Wong2, Hiromi K. Nakamura1, ….., Derek M. Huffman4, Junxiang Wan2, Radhika Muzumdar3, Nir Barzilai4 , and Pinchas Cohen2 http://www.impactaging.com/papers/v8/n4/full/100943.html
Mitochondria are key players in aging and in the pathogenesis of age-related diseases. Recent mitochondrial transcriptome analyses revealed the existence of multiple small mRNAs transcribed from mitochondrial DNA (mtDNA). Humanin (HN), a peptide encoded in the mtDNA 16S ribosomal RNA region, is a neuroprotective factor. An in silico search revealed six additional peptides in the same region of mtDNA as humanin; we named these peptides small humanin-like peptides (SHLPs). We identified the functional roles for these peptides and the potential mechanisms of action. The SHLPs differed in their ability to regulate cell viability in vitro. We focused on SHLP2 and SHLP3 because they shared similar protective effects with HN. Specifically, they significantly reduced apoptosis and the generation of reactive oxygen species, and improved mitochondrial metabolism in vitro. SHLP2 and SHLP3 also enhanced 3T3-L1 pre-adipocyte differentiation. Systemic hyperinsulinemic-euglycemic clamp studies showed that intracerebrally infused SHLP2 increased glucose uptake and suppressed hepatic glucose production, suggesting that it functions as an insulin sensitizer both peripherally and centrally. Similar to HN, the levels of circulating SHLP2 were found to decrease with age. These results suggest that mitochondria play critical roles in metabolism and survival through the synthesis of mitochondrial peptides, and provide new insights into mitochondrial biology with relevance to aging and human biology.
Human mitochondrial DNA (mtDNA) is a double-stranded, circular molecule of 16,569 bp and contains 37 genes encoding 13 proteins, 22 tRNAs, and 2 rRNAs. Recent mitochondrial transcriptome analyses revealed the existence of small RNAs derived from mtDNA [1]. In 2001, Nishimoto and colleagues identified humanin (HN), a 24-amino-acid peptide encoded from the 16S ribosomal RNA (rRNA) region of mtDNA. HN is a potent neuroprotective factor capable of antagonizing Alzheimer’s disease (AD)-related cellular insults [2]. HN is a component of a novel retrograde signaling pathway from the mitochondria to the nucleus, which is distinct from mitochondrial signaling pathways, such as the SIRT4-AMPK pathway [3]. HN-dependent cellular protection is mediated in part by interacting with and antagonizing pro-apoptotic Bax-related peptides [4] and IGFBP-3 (IGF binding protein 3) [5].
Because of their involvement in energy production and free radical generation, mitochondria likely play a major role in aging and age-related diseases [6–8]. In fact, improvement of mitochondrial function has been shown to ameliorate age-related memory loss in aged mice [9]. Recent studies have shown that HN levels decrease with age, suggesting that HN could play a role in aging and age-related diseases, such as Alzheimer’s disease (AD), atherosclerosis, and diabetes. Along with lower HN levels in the hypothalamus, skeletal muscle, and cortex of older rodents, the circulating levels of HN were found to decline with age in both humans and mice [10]. Notably, circulating HN levels were found to be (i) significantly higher in long-lived Ames dwarf mice but lower in short-lived growth hormone (GH) transgenic mice, (ii) significantly higher in a GH-deficient cohort of patients with Laron syndrome, and (iii) reduced in mice and humans treated with GH or IGF-1 (insulin-like growth factor 1) [11]. Age-dependent declines in the circulating HN levels may be due to higher levels of reactive oxygen species (ROS) that contribute to atherosclerosis development. Using mouse models of atherosclerosis, it was found that HN-treated mice had a reduced disease burden and significant health improvements [12,13]. In addition, HN improved insulin sensitivity, suggesting clinical potential for mitochondrial peptides in diseases of aging [10]. The discovery of HN represents a unique addition to the spectrum of roles that mitochondria play in the cell [14,15]. A second mitochondrial-derived peptide (MDP), MOTS-c (mitochondrial open reading frame of the 12S rRNA-c), has also been shown to have metabolic effects on muscle and may also play a role in aging [16].
We further investigated mtDNA for the presence of other MDPs. Recent technological advances have led to the identification of small open reading frames (sORFs) in the nuclear genomes ofDrosophila[17,18] and mammals [19,20]. Therefore, we attempted to identify novel sORFs using the following approaches: 1) in silico identification of potential sORFs; 2) determination of mRNA expression levels; 3) development of specific antibodies against these novel peptides to allow for peptide detection in cells, organs, and plasma; 4) elucidating the actions of these peptides by performing cell-based assays for mitochondrial function, signaling, viability, and differentiation; and 5) delivering these peptides in vivo to determine their systemic metabolic effects. Focusing on the 16S rRNA region of the mtDNA where the humanin gene is located, we identified six sORFs and named them small humanin-like peptides (SHLPs) 1-6. While surveying the biological effects of SHLPs, we found that SHLP2 and SHLP3 were cytoprotective; therefore, we investigated their effects on apoptosis and metabolism in greater detail. Further, we showed that circulating SHLP2 levels declined with age, similar to HN, suggesting that SHLP2 is involved in aging and age-related disease progression.
SHLP2 and SHLP3 regulate the expression of metabolic and inflammatory markers
Epidemiological studies have demonstrated that increased levels of mediators of inflammation and acute-phase reactants, such as fibrinogen, C-reactive protein (CRP), and IL-6, correlate with the incidence of type 2 diabetes mellitus (T2DM) [34–36]. In humans, anti-inflammatory drugs, such as aspirin and sodium salicylate, reduce fasting plasma glucose levels and ameliorate the symptoms of T2DM. In addition, anti-diabetic drugs, such as fibrates [37] and thiazolidinediones [38], have been found to lower some markers of inflammation. SHLP2 increased the levels of leptin, which is known to improve insulin sensitivity, but had no effect on the levels of the pro-inflammatory cytokines IL-6 and MCP-1. SHLP3 significantly increased the leptin levels, but also elevated IL-6 and MCP-1 levels, which could explain the lack of an in vivo insulin-sensitizing effect of SHLP3. The mechanism by which SHLPs regulate the expression of metabolic and inflammatory markers remains unclear and needs to be further investigated. Furthermore, SHLPs have different effects on inflammatory marker expression, suggesting differential regulation and function of individual SHLPs.
SHLP2 in aging
Mitochondria have been implicated in increased lifespan in several life-extending treatments [39,40]; however, it is not known whether the relationship is correlative or causative [40]. Additionally, it is well known that hormone levels change with aging. For example, levels of aldosterone, calcitonin, growth hormone, and IGF-I decrease with age. Circulating HN levels decline with age in humans and rodents, specifically in the hypothalamus and skeletal muscle of older rats. These changes parallel increases in the incidence of age-associated diseases such as AD and T2DM. The decline in circulating SHLP2 levels with age (Fig. 6), the anti-oxidative stress function of SHLP2 (Fig. 3C), and its neuroprotective effect (Fig. 6B) indicate that SHLP2 has a role in the regulation of aging and age-related diseases.
Conclusion
By analyzing the mitochondrial transcriptome, we found that sORFs from mitochondrial DNA encode functional peptides. We identified many mRNA transcripts within 13 protein-coding mitochondrial genes [1]. Such previously underappreciated sORFs have also been described in the nuclear genome [41]. The MDPs we describe here may represent retrograde communication signals from the mitochondria to the nucleus and may explain important aspects of mitochondrial biology that are implicated in health and longevity.
Larry, John Walker is working on mt proteins dynamics. His rotor – stator mechanism in ATPase synthase, a ‘complex’ that biologist accepted as energy generator is likely wrong. I was suppose to have met him in Germany few years ago. Energy in biological systems has nothing to do with heat. Heat is an outcome of a reaction, meaning that IR spectra accordingly to wave theory is a source of information memorized in water interference with carbon open systems within protein and glyo-proteins complexes as well as genome space-time outcomes. Physically speaking from a pure perspective of science ATP is highly unstable form of phosphate ‘chains’. It cannot hold energy, it is actually in contrary, it is like a resonator, trapping negativity, thus functioning as space propeller by expanding carbon skeleton of protein ‘machines’ Now, we don’t know what is ‘aging’ in a pure physical sense, except that we observe structural changes in what we call complexes. We we know is that proteins are not stationary structures, but highly dynamic forms of matter, seemingly occupying discrete and relative spaces. A piece of mt ATP ase could be discovered in the nucleus as transcription factor. Our notion of operational space in terms of electro dynamics from a motor – stator perspective is now translated toward defining semi conducting and supracoductive strings. The reality of which is so much more fascinating and beautiful as time progresses overally. There are spaces where time does not change, and there are spaces where time walks, and there are spaces, where time flies, and there are spaces where time runs. Amazing, indeed! The story of aging gets a lot deeper that science could even imagine, probably to roots of immortal energy- spaces. We know that matter is transient, that is nearly all living matter, replenishes of about 3 to 7 weeks.
Take a glass full of some kind of liquid, you know the mass of the glass and the mass of the liquid (say wine, beer, water, or milk) You also know to an approximate reality the composition of both. Now lift the glass full of liquid and let it break on a surface of your choice. Depending on the surface pieces of the glass would travel differential from a center projected by the vertical axis of your hand. What technology does today is recollecting those pieces and modelling them to fit in a form again that would resemble a holding device, a glass. The liquid we don’t know exactly how it spilled due the nature of its absorbancy of both surface physics and physical ‘state’ properties. Thus we can say how much approximate energy we have held thinking of m/z as time flight objectives. Each technology can read 1D and approximate the 2D, absolutely lacking computational methodology for 3D dynamic reality. Many scientists confuse space and volume. Volume is a one dimensional characteristic! So is crystalography! BY taking quantum chemical method computing principles following imaginative rules we could approach 2D, however , that is not enough to define 3D. Time we use as a reference frame of clocks we have invented in order to keep track of a sense to observable ‘change’ . But remember, time is absolute and parallel in continuity while energy is discrete , coming in quantum packages, realization of accumulated information. Information is highly redundant we see, so annotating information is an objective to modern days simulations that could predict outcomes of possible parallel realities we call worlds. One could ‘jump’ from one reality to another through guidance of light and water, but what remains unsolved is why people make mistakes, constantly by accusing in name of greed and power , or disobedience of commandments of the Lord!
On Thu, Apr 21, 2016 at 3:41 AM, Leaders in Pharmaceutical Business Intelligence (LPBI) Group wrote:
> larryhbern posted: “New Insights into mtDNA, mitochondrial proteins, > aging, and metabolic control Larry H. Bernstein, MD, FCAP, Curator LPBI > Newly discovered proteins may protect against age-related illnesses The > proteins could play a key role in the ” >
Metabolic features of the cell danger response
– Mitochondria in Health and Disease
The Cell Danger Response (CDR) is defined in terms of an ancient metabolic response to threat.
The CDR encompasses inflammation, innate immunity, oxidative stress, and the ER stress response.
The CDR is maintained by extracellular nucleotide (purinergic) signaling.
Abnormal persistence of the CDR lies at the heart of many chronic diseases.
Antipurinergic therapy (APT) has proven effective in many chronic disorders in animal models
The cell danger response (CDR) is the evolutionarily conserved metabolic response that protects cells and hosts from harm. It is triggered by encounters with chemical, physical, or biological threats that exceed the cellular capacity for homeostasis. The resulting metabolic mismatch between available resources and functional capacity produces a cascade of changes in cellular electron flow, oxygen consumption, redox, membrane fluidity, lipid dynamics, bioenergetics, carbon and sulfur resource allocation, protein folding and aggregation, vitamin availability, metal homeostasis, indole, pterin, 1-carbon and polyamine metabolism, and polymer formation. The first wave of danger signals consists of the release of metabolic intermediates like ATP and ADP, Krebs cycle intermediates, oxygen, and reactive oxygen species (ROS), and is sustained by purinergic signaling. After the danger has been eliminated or neutralized, a choreographed sequence of anti-inflammatory and regenerative pathways is activated to reverse the CDR and to heal. When the CDR persists abnormally, whole body metabolism and the gut microbiome are disturbed, the collective performance of multiple organ systems is impaired, behavior is changed, and chronic disease results. Metabolic memory of past stress encounters is stored in the form of altered mitochondrial and cellular macromolecule content, resulting in an increase in functional reserve capacity through a process known as mitocellular hormesis. The systemic form of the CDR, and its magnified form, the purinergic life-threat response (PLTR), are under direct control by ancient pathways in the brain that are ultimately coordinated by centers in the brainstem. Chemosensory integration of whole body metabolism occurs in the brainstem and is a prerequisite for normal brain, motor, vestibular, sensory, social, and speech development. An understanding of the CDR permits us to reframe old concepts of pathogenesis for a broad array of chronic, developmental, autoimmune, and degenerative disorders. These disorders include autism spectrum disorders (ASD), attention deficit hyperactivity disorder (ADHD), asthma, atopy, gluten and many other food and chemical sensitivity syndromes, emphysema, Tourette’s syndrome, bipolar disorder, schizophrenia, post-traumatic stress disorder (PTSD), chronic traumatic encephalopathy (CTE), traumatic brain injury (TBI), epilepsy, suicidal ideation, organ transplant biology, diabetes, kidney, liver, and heart disease, cancer, Alzheimer and Parkinson disease, and autoimmune disorders like lupus, rheumatoid arthritis, multiple sclerosis, and primary sclerosing cholangitis.
This is a confocal microscopy image of human fibroblasts derived from embryonic stem cells. The nuclei appear in blue, while smaller and more numerous mitochondria appear in red. [Shoukhrat Mitalipov]
Mutations in our mitochondrial DNA tend to be inconspicuous, but they can become more prevalent as we age. They can even vary in frequency from cell to cell. Naturally, some cells will be relatively compromised because they happen to have a higher percentage of mutated mitochondrial DNA. Such cells make a poor basis for stem cell lines. They should be excluded. But how?
To answer this question, a team of scientists scrutinized skin fibroblasts, blood cells, and induced pluripotent stem cells (iPSCs) for mitochondrial genome integrity. When the scientists tested the samples for mitochondrial DNA mutations, the levels of mutations appeared low. But when the scientists sequenced the iPS cell lines, they found higher numbers of mitochondrial DNA mutations, particularly in cells from patients over 60.
The scientists were led by Shoukhrat Mitalipov, Ph.D., director of the Center for Embryonic Cell and Gene Therapy at Oregon Health & Science University, and Taosheng Huang, M.D., a medical geneticist and director of the Mitochondrial Medicine Program at Cincinnati Children’s Hospital. The Mitalipov/Huang-led team also found higher percentages of mitochondria containing mutations within a cell. The higher the load of mutated mitochondrial DNA in a cell, the more compromised the cell’s function.
Since each iPSC line is created from a different cell, each line may contain different types of mitochondrial DNA mutations and mutation loads. To choose the least damaged line, the authors recommend screening multiple lines per patient. “It’s a good idea to check the iPS clones for mitochondrial DNA mutations and make sure you pick a good cell line,” said Dr. Huang.
This recommendation appeared April 14 in the journal Cell Stem Cell, in an article entitled, “Age-Related Accumulation of Somatic Mitochondrial DNA Mutations in Adult-Derived Human iPSCs.” This article holds that mitochondrial genome integrity is a vital readout in assessing the proficiency of patient-derived regenerative products destined for clinical applications.
“We found that pooled skin and blood mtDNA contained low heteroplasmic point mutations, but a panel of ten individual iPSC lines from each tissue or clonally expanded fibroblasts carried an elevated load of heteroplasmic or homoplasmic mutations, suggesting that somatic mutations randomly arise within individual cells but are not detectable in whole tissues,” wrote the article’s authors. “The frequency of mtDNA defects in iPSCs increased with age, and many mutations were nonsynonymous or resided in RNA coding genes and thus can lead to respiratory defects.”
Potential therapies using stem cells hold tremendous promise for treating human disease. However, defects in the mitochondria could undermine the iPS cells’ ability to repair damaged tissue or organs.
“If you want to use iPS cells in a human, you must check for mutations in the mitochondrial genome,” declared Dr. Huang. “Every single cell can be different. Two cells next to each other could have different mutations or different percentages of mutations.”
Prior to the creation of a therapeutic iPS cell line, a collection of cells is taken from the patient. These cells will be tested for mutations. If the tester uses Sanger sequencing, older technology that is not as sensitive as newer next-generation sequencing, any mutation that occurs in less than 20% of the sample will go undetected. But mitochondrial DNA mutations might occur in less than 20% of mitochondria in the pooled cells. As a result, mutation rates have not been well understood. “These mitochondrial mutations are actually hidden,” explained Dr. Mitalipov.
The mitochondrial genome is relatively small, containing just 37 genes, so screening should be feasible using next generation sequencing, Dr. Mitalipov added. “It should be relatively cheap and do-able.”
Dr. Mitalipov also commented on a more general point, the implications of the current study on illuminating the mechanisms of age-related disease: “Pathogenic mutations in our mitochondrial DNA have long been thought to be a driving force in aging and age-onset diseases, though clear evidence was missing. This foundational knowledge of how cells are damaged in the natural process of aging may help to illuminate the role of mutated mitochondria in degenerative disease.”
A new paperfrom Shoukhrat Mitalipov’s lab on stem cell mitochondria points to a pattern whereby induced pluripotent stem (IPS) cells tend to have more problems if they are from older patients.
What does this paper mean for the stem cell field and could it impact more specifically the clinical applications of IPS cells?
The new paper Kang, et al is entitled “Age-Related Accumulation of Somatic Mitochondrial DNA Mutations in Adult-Derived Human iPSCs”.
This paper reminds us of the very important realities that mitochondria are key players in stem cell function and that mitochondria have their own genomes that impact that function. A lot of us don’t think about mitochondria and their genome as often as we should.
The paper came to three major scientific conclusions (this from the Highlights section of the paper and also see the graphical abstract for a visual sense of the results overall):
Human iPSC clones derived from elderly adults show accumulation of mtDNA mutations
Fewer mtDNA mutations are present in ESCs and iPSCs derived from younger adults
Accumulated mtDNA mutations can impact metabolic function in iPSCs
Importantly the team looked at IPS cells derived from both blood and skin cells and found that the former were less likely to have mitochondrial mutations.
This study suggests that those teams producing or working with human IPS cells (hIPSCs) should be screening the different lines for mitochondrial mutations. This excellent piece from Sara Reardon on the Mitalipov paper quotes IPS cell expert Jeanne Loring on this very point:
“It’s one of those things most of us don’t think about,” says Jeanne Loring, a stem-cell biologist at the Scripps Research Institute in La Jolla, California. Her lab is working towards using iPS cells to treat Parkinson’s disease, and Loring now plans to go back and examine the mitochondria in her cell lines. She suspects that it will be fairly easy for researchers to screen cells for use in therapies.”
Mitalipov goes further and suggests that his team’s new findings could support the use of human embryonic stem cells (hESC) derived by somatic cell nuclear transfer (SCNT) which would be expected to have mitochondria with fewer mutations. However, as Loring points out in the Reardon article, SCNT is really difficult to successfully perform and only a few labs in the world can do it at present. In that context, working with hIPSC and adding on the additional layer of mitochondrial DNA mutation screening could be more practical.
New York stem cell researcher Dieter Egli, however, is quoted that hIPSC have other differences with hESC as well such as epigenetic differences and he’s quoted in the Reardon piece, “It’s going to be very hard to find a cell line that’s perfect.”
One might reasonably ask both Egli and oneself, “What is a perfect cell line”?
In the end the best approach for use of human pluripotent stem cells of any kind is going to involve a balance between practicality of production and the potentially positive or negative traits of those cells as determined by rigorous validation screening.
With this new paper we’ve just learned more about another layer of screening that is needed. An interesting question is whether adult stem cells such as mesenchymal stromal/stem cells (MSC) also should be screened for mitochondrial mutations. They are often produced from patients who are getting up there in years. I hope that someone will publish on that too.
As to pluripotent cells, I expect that sometimes the best lines, meaning those most perfect for a given clinical application, will be hIPSC (autologous or allogeneic in some instances) and in other cases they may be hESC made from leftover IVF embryos. If SCNT-derived hESC can be more widely produced in an affordable manner and they pass validation as well then those (sometimes called NT-hESC) may also come into play clinically. So far that hasn’t happened for the SCNT cells, but it may over time. …..
Age-Related Accumulation of Somatic Mitochondrial DNA Mutations in Adult-Derived Human iPSCs
In Brief Mitalipov, Huang, and colleagues show that human iPSCs derived from older adults carry more mitochondrial DNA mutations than those derived from younger individuals. Defects in metabolic function caused by mtDNA mutations suggest careful screening of hiPSC clones for mutational load before clinical application.
Highlights
Human iPSC clones derived from elderly adults show accumulation of mtDNA mutations
Fewer mtDNA mutations are present in ESCs and iPSCs derived from younger adults
Accumulated mtDNA mutations can impact metabolic function in iPSCs
The genetic integrity of iPSCs is an important consideration for therapeutic application. In this study, we examine the accumulation of somatic mitochondrial genome (mtDNA) mutations in skin fibroblasts, blood, and iPSCs derived from young and elderly subjects (24–72 years). We found that pooled skin and blood mtDNA contained low heteroplasmic point mutations, but a panel of ten individual iPSC lines from each tissue or clonally expanded fibroblasts carried an elevated load of heteroplasmic or homoplasmic mutations, suggesting that somatic mutations randomly arise within individual cells but are not detectable in whole tissues. The frequency of mtDNA defects in iPSCs increased with age, and many mutations were non-synonymous or resided in RNA coding genes and thus can lead to respiratory defects. Our results highlight a need to monitor mtDNA mutations in iPSCs, especially those generated from older patients, and to examine the metabolic status of iPSCs destined for clinical applications.
Induced pluripotent stem cells (iPSCs) offer an unlimited source for autologous cell replacement therapies to treat age-associated degenerative diseases. Aging is generally characterized by increased DNA damage and genomic instability (Garinis et al., 2008; Lombard et al., 2005); thus, iPSCs derived from elderly subjects may harbor point mutations and larger genomic rearrangements. Indeed, iPSCs display increased chromosome aberrations (Mayshar et al., 2010), subchromosomal copy number variations (CNVs) (Abyzov et al., 2012; Laurent et al., 2011), and exome mutations (Johannesson et al., 2014), compared to natural embryonic stem cell (ESC) counterparts (Ma et al., 2014). The rate of mtDNA mutations is believed to be at least 10- to 20-fold higher than that observed in the nuclear genome (Wallace, 1994), and often both mutated and wild-type mtDNA (heteroplasmy) can coexist in the same cell (Rossignol et al., 2003). Large deletions are most frequently observed mtDNA abnormalities in aged post-mitotic tissues such as brain, heart, and muscle (Bender et al., 2006; Bua et al., 2006; Corral-Debrinski et al., 1992; Cortopassi et al., 1992; Mohamed et al., 2006) and have been implicated in aging and diseases such as Alzheimer’s disease (AD), Parkinson’s disease (PD), amyotrophic lateral sclerosis (ALS), and diabetes (Larsson, 2010; Lin and Beal, 2006; Petersen et al., 2003; Wallace, 2005). In addition, mtDNA point mutations were reported in some tumors and replicating tissues (Chatterjee et al., 2006; Ju et al., 2014; Michikawa et al., 1999; Taylor et al., 2003). However, the extent of mtDNA defects in proliferating peripheral tissues commonly used for iPSC induction, such as skin and blood, is thought to be low and limited to common non-coding variants (Schon et al., 2012; Yao et al., 2015). Accumulation of mtDNA variants in these tissues with age was insignificant (Greaves et al., 2010; Hashizume et al., 2015). Several point mutations were identified in iPSCs generated from the newborn foreskin fibroblasts, although most of these variants were non-coding, common for the general population, and did not affect their metabolic activity (Prigione et al., 2011). Somatic mtDNA mutations may be under-reported secondary to the level of sample interrogation. …..
Figure 2. mtDNA Mutations in Skin Fibroblasts, Blood, and the iPSCs of a 72-YearOld B Subject (A) Sixteen mutations at low heteroplasmy levels were detected in the DNA of PF, while a panel of ten FiPSC lines carried nine mutations, including four that were homoplasmic. Gray rectangles define the mutations shared between PF and FiPSCs. (B) Venn diagram showing only one mutation in FiPSCs shared with PF. (C) All ten FiPSC lines carried between one and five high-heteroplasmy (>15%) mutations. (D) Mutation distribution in whole blood and BiPSCs was similar to that in PF and FiPSCs. Six mutations at low-heteroplasmy levels were observed in blood, while BiPSC lines displayed 21 mutations, including four over the 80% heteroplasmy level. (E) Venn diagram showing four mutations in BiPSCs shared with whole blood and the 17 novel variants. (F) Distribution of mutations in individual BiPSC lines. See also Figures S2 and S3; Table S1; Table S3, sheet 2; and Table S4, sheet 1 ….
Figure 4. Transmission and Distribution of Somatic mtDNA Mutations to iPSCs (A) A total of 112 mtDNA mutations were discovered in parental cells (PF, CF, and blood) from 11 subjects. Of these, 39 variants (35%) were found in corresponding 130 iPSC lines. Among non-transmitted, transmitted, and novel mutations in iPSCs, comparable percentages of variants (68%, 69%, and 79%, respectively) were coding mutations in protein, rRNA, or tRNA genes. This suggests that most pathogenic mutations do not affect iPSC induction. However, certain coding mutations including in ND3, ND4L, and 14 tRNA genes were not detected in iPSCs, suggesting possible pathogenicity. n, the number of mtDNA mutations. Blue font genes were detected in parental cells. (B–D) A total of 80 high heteroplasmic (>15%) variants were detected in the present study in 130 FiPSC or BiPSC lines from 11 subjects. (B) The majority of these variants (76%) were non-synonymous or frame-shift mutations in protein-coding genes or affected rRNA and tRNA genes. (C) More than half of the mutations (56%) were never reported in a database containing whole mtDNA sequences from 26,850 healthy subjects representing the general human population (http://www.mitomap.org/MITOMAP). (D) Most mutations (90%) were never reported in a database containing sequences from healthy subjects with corresponding mtDNA haplotypes. freq., frequent. See also Figure S5 and Tables S3 and S4. ….
sjwilliamspa
Mutations will accumulate over age in mitochondrial DNA, however the current study has the difficulty that the authors could not use patient-age-matched controls, in essence they could only compare induced pluripotent stem cells derived from different patients. This could confound the results but the result with higher frequency of mutation in mtDNA in cells reprogrammed from younger patients is interesting but might limit the ability of autologous regenerative therapy in older patients. However reprogramming, although the method not mentioned here although I am assuming by transfection with lentivirus is a rough procedure, involving multiple dedifferentiation steps. Therefore it is very understandable that cells obtained from elderly patients would respond less favorably to such a rough reprogramming regimen, especially if it produced a higher degree of ROS, which has been shown to alter mtDNA. This is why I feel it is more advantageous to obtain a stem cell population from fat cells and forgo the Oct4, htert, reprogramming with lentiviral vectors.
CRISPR/Cas9, Familial Amyloid Polyneuropathy (FAP) and Neurodegenerative Disease, Volume 2 (Volume Two: Latest in Genomics Methodologies for Therapeutics: Gene Editing, NGS and BioInformatics, Simulations and the Genome Ontology), Part 2: CRISPR for Gene Editing and DNA Repair
CRISPR/Cas9, Familial Amyloid Polyneuropathy ( FAP) and Neurodegenerative Disease
Curator: Larry H. Bernstein, MD, FCAP
CRISPR/Cas9 and Targeted Genome Editing: A New Era in Molecular Biology
The development of efficient and reliable ways to make precise, targeted changes to the genome of living cells is a long-standing goal for biomedical researchers. Recently, a new tool based on a bacterial CRISPR-associated protein-9 nuclease (Cas9) from Streptococcus pyogenes has generated considerable excitement (1). This follows several attempts over the years to manipulate gene function, including homologous recombination (2) and RNA interference (RNAi) (3). RNAi, in particular, became a laboratory staple enabling inexpensive and high-throughput interrogation of gene function (4, 5), but it is hampered by providing only temporary inhibition of gene function and unpredictable off-target effects (6). Other recent approaches to targeted genome modification – zinc-finger nucleases [ZFNs, (7)] and transcription-activator like effector nucleases [TALENs (8)]– enable researchers to generate permanent mutations by introducing doublestranded breaks to activate repair pathways. These approaches are costly and time-consuming to engineer, limiting their widespread use, particularly for large scale, high-throughput studies.
The Biology of Cas9
The functions of CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) and CRISPR-associated (Cas) genes are essential in adaptive immunity in select bacteria and archaea, enabling the organisms to respond to and eliminate invading genetic material. These repeats were initially discovered in the 1980s in E. coli (9), but their function wasn’t confirmed until 2007 by Barrangou and colleagues, who demonstrated that S. thermophilus can acquire resistance against a bacteriophage by integrating a genome fragment of an infectious virus into its CRISPR locus (10).
Three types of CRISPR mechanisms have been identified, of which type II is the most studied. In this case, invading DNA from viruses or plasmids is cut into small fragments and incorporated into a CRISPR locus amidst a series of short repeats (around 20 bps). The loci are transcribed, and transcripts are then processed to generate small RNAs (crRNA – CRISPR RNA), which are used to guide effector endonucleases that target invading DNA based on sequence complementarity (Figure 1) (11).
Figure 1. Cas9 in vivo: Bacterial Adaptive Immunity
In the acquisition phase, foreign DNA is incorporated into the bacterial genome at the CRISPR loci. CRISPR loci is then transcribed and processed into crRNA during crRNA biogenesis. During interference, Cas9 endonuclease complexed with a crRNA and separate tracrRNA cleaves foreign DNA containing a 20-nucleotide crRNA complementary sequence adjacent to the PAM sequence. (Figure not drawn to scale.)
One Cas protein, Cas9 (also known as Csn1), has been shown, through knockdown and rescue experiments to be a key player in certain CRISPR mechanisms (specifically type II CRISPR systems). The type II CRISPR mechanism is unique compared to other CRISPR systems, as only one Cas protein (Cas9) is required for gene silencing (12). In type II systems, Cas9 participates in the processing of crRNAs (12), and is responsible for the destruction of the target DNA (11). Cas9’s function in both of these steps relies on the presence of two nuclease domains, a RuvC-like nuclease domain located at the amino terminus and a HNH-like nuclease domain that resides in the mid-region of the protein (13).
To achieve site-specific DNA recognition and cleavage, Cas9 must be complexed with both a crRNA and a separate trans-activating crRNA (tracrRNA or trRNA), that is partially complementary to the crRNA (11). The tracrRNA is required for crRNA maturation from a primary transcript encoding multiple pre-crRNAs. This occurs in the presence of RNase III and Cas9 (12).
During the destruction of target DNA, the HNH and RuvC-like nuclease domains cut both DNA strands, generating double-stranded breaks (DSBs) at sites defined by a 20-nucleotide target sequence within an associated crRNA transcript (11, 14). The HNH domain cleaves the complementary strand, while the RuvC domain cleaves the noncomplementary strand.
The double-stranded endonuclease activity of Cas9 also requires that a short conserved sequence, (2–5 nts) known as protospacer-associated motif (PAM), follows immediately 3´- of the crRNA complementary sequence (15). In fact, even fully complementary sequences are ignored by Cas9-RNA in the absence of a PAM sequence (16).
Cas9 and CRISPR as a New Tool in Molecular Biology
The simplicity of the type II CRISPR nuclease, with only three required components (Cas9 along with the crRNA and trRNA) makes this system amenable to adaptation for genome editing. This potential was realized in 2012 by the Doudna and Charpentier labs (11). Based on the type II CRISPR system described previously, the authors developed a simplified two-component system by combining trRNA and crRNA into a single synthetic single guide RNA (sgRNA). sgRNAprogrammed Cas9 was shown to be as effective as Cas9 programmed with separate trRNA and crRNA in guiding targeted gene alterations (Figure 2A).
To date, three different variants of the Cas9 nuclease have been adopted in genome-editing protocols. The first is wild-type Cas9, which can site-specifically cleave double-stranded DNA, resulting in the activation of the doublestrand break (DSB) repair machinery. DSBs can be repaired by the cellular Non-Homologous End Joining (NHEJ) pathway (17), resulting in insertions and/or deletions (indels) which disrupt the targeted locus. Alternatively, if a donor template with homology to the targeted locus is supplied, the DSB may be repaired by the homology-directed repair (HDR) pathway allowing for precise replacement mutations to be made (Figure 2A) (17, 18).
Cong and colleagues (1) took the Cas9 system a step further towards increased precision by developing a mutant form, known as Cas9D10A, with only nickase activity. This means it cleaves only one DNA strand, and does not activate NHEJ. Instead, when provided with a homologous repair template, DNA repairs are conducted via the high-fidelity HDR pathway only, resulting in reduced indel mutations (1, 11, 19). Cas9D10A is even more appealing in terms of target specificity when loci are targeted by paired Cas9 complexes designed to generate adjacent DNA nicks (20) (see further details about “paired nickases” in Figure 2B).
The third variant is a nuclease-deficient Cas9 (dCas9, Figure 2C) (21). Mutations H840A in the HNH domain and D10A in the RuvC domain inactivate cleavage activity, but do not prevent DNA binding (11, 22). Therefore, this variant can be used to sequence-specifically target any region of the genome without cleavage. Instead, by fusing with various effector domains, dCas9 can be used either as a gene silencing or activation tool (21, 23–26). Furthermore, it can be used as a visualization tool. For instance, Chen and colleagues used dCas9 fused to Enhanced Green Fluorescent Protein (EGFP) to visualize repetitive DNA sequences with a single sgRNA or nonrepetitive loci using multiple sgRNAs (27).
Wild-type Cas9 nuclease site specifically cleaves double-stranded DNA activating double-strand break repair machinery. In the absence of a homologous repair template non-homologous end joining can result in indels disrupting the target sequence. Alternatively, precise mutations and knock-ins can be made by providing a homologous repair template and exploiting the homology directed repair pathway.
B. Mutated Cas9 makes a site specific single-strand nick. Two sgRNA can be used to introduce a staggered double-stranded break which can then undergo homology directed repair.
C. Nuclease-deficient Cas9 can be fused with various effector domains allowing specific localization. For example, transcriptional activators, repressors, and fluorescent proteins.
Targeting Efficiency and Off-target Mutations
Targeting efficiency, or the percentage of desired mutation achieved, is one of the most important parameters by which to assess a genome-editing tool. The targeting efficiency of Cas9 compares favorably with more established methods, such as TALENs or ZFNs (8). For example, in human cells, custom-designed ZFNs and TALENs could only achieve efficiencies ranging from 1% to 50% (29–31). In contrast, the Cas9 system has been reported to have efficiencies up to >70% in zebrafish (32) and plants (33), and ranging from 2–5% in induced pluripotent stem cells (34). In addition, Zhou and colleagues were able to improve genome targeting up to 78% in one-cell mouse embryos, and achieved effective germline transmission through the use of dual sgRNAs to simultaneously target an individual gene (35).
A widely used method to identify mutations is the T7 Endonuclease I mutation detection assay (36, 37) (Figure 3). This assay detects heteroduplex DNA that results from the annealing of a DNA strand, including desired mutations, with a wildtype DNA strand (37).
Figure 3. T7 Endonuclease I Targeting Efficiency Assay
Genomic DNA is amplified with primers bracketing the modified locus. PCR products are then denatured and re-annealed yielding 3 possible structures. Duplexes containing a mismatch are digested by T7 Endonuclease I. The DNA is then electrophoretically separated and fragment analysis is used to calculate targeting efficiency.
Another important parameter is the incidence of off-target mutations. Such mutations are likely to appear in sites that have differences of only a few nucleotides compared to the original sequence, as long as they are adjacent to a PAM sequence. This occurs as Cas9 can tolerate up to 5 base mismatches within the protospacer region (36) or a single base difference in the PAM sequence (38). Off-target mutations are generally more difficult to detect, requiring whole-genome sequencing to rule them out completely.
Recent improvements to the CRISPR system for reducing off-target mutations have been made through the use of truncated gRNA (truncated within the crRNA-derived sequence) or by adding two extra guanine (G) nucleotides to the 5´ end (28, 37). Another way researchers have attempted to minimize off-target effects is with the use of “paired nickases” (20). This strategy uses D10A Cas9 and two sgRNAs complementary to the adjacent area on opposite strands of the target site (Figure 2B). While this induces DSBs in the target DNA, it is expected to create only single nicks in off-target locations and, therefore, result in minimal off-target mutations.
By leveraging computation to reduce off-target mutations, several groups have developed webbased tools to facilitate the identification of potential CRISPR target sites and assess their potential for off-target cleavage. Examples include the CRISPR Design Tool (38) and the ZiFiT Targeter, Version 4.2 (39, 40).
Applications as a Genome-editing and Genome Targeting Tool
Following its initial demonstration in 2012 (9), the CRISPR/Cas9 system has been widely adopted. This has already been successfully used to target important genes in many cell lines and organisms, including human (34), bacteria (41), zebrafish (32), C. elegans (42), plants (34), Xenopus tropicalis (43), yeast (44), Drosophila (45), monkeys (46), rabbits (47), pigs (42), rats (48) and mice (49). Several groups have now taken advantage of this method to introduce single point mutations (deletions or insertions) in a particular target gene, via a single gRNA (14, 21, 29). Using a pair of gRNA-directed Cas9 nucleases instead, it is also possible to induce large deletions or genomic rearrangements, such as inversions or translocations (50). A recent exciting development is the use of the dCas9 version of the CRISPR/Cas9 system to target protein domains for transcriptional regulation (26, 51, 52), epigenetic modification (25), and microscopic visualization of specific genome loci (27).
The CRISPR/Cas9 system requires only the redesign of the crRNA to change target specificity. This contrasts with other genome editing tools, including zinc finger and TALENs, where redesign of the protein-DNA interface is required. Furthermore, CRISPR/Cas9 enables rapid genome-wide interrogation of gene function by generating large gRNA libraries (51, 53) for genomic screening.
The Future of CRISPR/Cas9
The rapid progress in developing Cas9 into a set of tools for cell and molecular biology research has been remarkable, likely due to the simplicity, high efficiency and versatility of the system. Of the designer nuclease systems currently available for precision genome engineering, the CRISPR/Cas system is by far the most user friendly. It is now also clear that Cas9’s potential reaches beyond DNA cleavage, and its usefulness for genome locus-specific recruitment of proteins will likely only be limited by our imagination.
Scientists urge caution in using new CRISPR technology to treat human genetic disease
The bacterial enzyme Cas9 is the engine of RNA-programmed genome engineering in human cells. (Graphic by Jennifer Doudna/UC Berkeley)
A group of 18 scientists and ethicists today warned that a revolutionary new tool to cut and splice DNA should be used cautiously when attempting to fix human genetic disease, and strongly discouraged any attempts at making changes to the human genome that could be passed on to offspring.
Among the authors of this warning is Jennifer Doudna, the co-inventor of the technology, called CRISPR-Cas9, which is driving a new interest in gene therapy, or “genome engineering.” She and colleagues co-authored a perspective piece that appears in the March 20 issue of Science, based on discussions at a meeting that took place in Napa on Jan. 24. The same issue of Science features a collection of recent research papers, commentary and news articles on CRISPR and its implications. …..
A prudent path forward for genomic engineering and germline gene modification
Scientists today are changing DNA sequences to correct genetic defects in animals as well as cultured tissues generated from stem cells, strategies that could eventually be used to treat human disease. The technology can also be used to engineer animals with genetic diseases mimicking human disease, which could lead to new insights into previously enigmatic disorders.
The CRISPR-Cas9 tool is still being refined to ensure that genetic changes are precisely targeted, Doudna said. Nevertheless, the authors met “… to initiate an informed discussion of the uses of genome engineering technology, and to identify proactively those areas where current action is essential to prepare for future developments. We recommend taking immediate steps toward ensuring that the application of genome engineering technology is performed safely and ethically.”
Amyloid CRISPR Plasmids and si/shRNA Gene Silencers
Santa Cruz Biotechnology, Inc. offers a broad range of gene silencers in the form of siRNAs, shRNA Plasmids and shRNA Lentiviral Particles as well as CRISPR/Cas9 Knockout and CRISPR Double Nickase plasmids. Amyloid gene silencers are available as Amyloid siRNA, Amyloid shRNA Plasmid, Amyloid shRNA Lentiviral Particles and Amyloid CRISPR/Cas9 Knockout plasmids. Amyloid CRISPR/dCas9 Activation Plasmids and CRISPR Lenti Activation Systems for gene activation are also available. Gene silencers and activators are useful for gene studies in combination with antibodies used for protein detection. Amyloid CRISPR Knockout, HDR and Nickase Knockout Plasmids
CRISPR-Cas9-Based Knockout of the Prion Protein and Its Effect on the Proteome
The molecular function of the cellular prion protein (PrPC) and the mechanism by which it may contribute to neurotoxicity in prion diseases and Alzheimer’s disease are only partially understood. Mouse neuroblastoma Neuro2a cells and, more recently, C2C12 myocytes and myotubes have emerged as popular models for investigating the cellular biology of PrP. Mouse epithelial NMuMG cells might become attractive models for studying the possible involvement of PrP in a morphogenetic program underlying epithelial-to-mesenchymal transitions. Here we describe the generation of PrP knockout clones from these cell lines using CRISPR-Cas9 knockout technology. More specifically, knockout clones were generated with two separate guide RNAs targeting recognition sites on opposite strands within the first hundred nucleotides of the Prnp coding sequence. Several PrP knockout clones were isolated and genomic insertions and deletions near the CRISPR-target sites were characterized. Subsequently, deep quantitative global proteome analyses that recorded the relative abundance of>3000 proteins (data deposited to ProteomeXchange Consortium) were undertaken to begin to characterize the molecular consequences of PrP deficiency. The levels of ∼120 proteins were shown to reproducibly correlate with the presence or absence of PrP, with most of these proteins belonging to extracellular components, cell junctions or the cytoskeleton.
Recent advances in genome engineering technologies based on the CRISPR-associated RNA-guided endonuclease Cas9 are enabling the systematic interrogation of mammalian genome function. Analogous to the search function in modern word processors, Cas9 can be guided to specific locations within complex genomes by a short RNA search string. Using this system, DNA sequences within the endogenous genome and their functional outputs are now easily edited or modulated in virtually any organism of choice. Cas9-mediated genetic perturbation is simple and scalable, empowering researchers to elucidate the functional organization of the genome at the systems level and establish causal linkages between genetic variations and biological phenotypes. In this Review, we describe the development and applications of Cas9 for a variety of research or translational applications while highlighting challenges as well as future directions. Derived from a remarkable microbial defense system, Cas9 is driving innovative applications from basic biology to biotechnology and medicine.
The development of recombinant DNA technology in the 1970s marked the beginning of a new era for biology. For the first time, molecular biologists gained the ability to manipulate DNA molecules, making it possible to study genes and harness them to develop novel medicine and biotechnology. Recent advances in genome engineering technologies are sparking a new revolution in biological research. Rather than studying DNA taken out of the context of the genome, researchers can now directly edit or modulate the function of DNA sequences in their endogenous context in virtually any organism of choice, enabling them to elucidate the functional organization of the genome at the systems level, as well as identify causal genetic variations.
Broadly speaking, genome engineering refers to the process of making targeted modifications to the genome, its contexts (e.g., epigenetic marks), or its outputs (e.g., transcripts). The ability to do so easily and efficiently in eukaryotic and especially mammalian cells holds immense promise to transform basic science, biotechnology, and medicine (Figure 1).
For life sciences research, technologies that can delete, insert, and modify the DNA sequences of cells or organisms enable dissecting the function of specific genes and regulatory elements. Multiplexed editing could further allow the interrogation of gene or protein networks at a larger scale. Similarly, manipulating transcriptional regulation or chromatin states at particular loci can reveal how genetic material is organized and utilized within a cell, illuminating relationships between the architecture of the genome and its functions. In biotechnology, precise manipulation of genetic building blocks and regulatory machinery also facilitates the reverse engineering or reconstruction of useful biological systems, for example, by enhancing biofuel production pathways in industrially relevant organisms or by creating infection-resistant crops. Additionally, genome engineering is stimulating a new generation of drug development processes and medical therapeutics. Perturbation of multiple genes simultaneously could model the additive effects that underlie complex polygenic disorders, leading to new drug targets, while genome editing could directly correct harmful mutations in the context of human gene therapy (Tebas et al., 2014).
Eukaryotic genomes contain billions of DNA bases and are difficult to manipulate. One of the breakthroughs in genome manipulation has been the development of gene targeting by homologous recombination (HR), which integrates exogenous repair templates that contain sequence homology to the donor site (Figure 2A) (Capecchi, 1989). HR-mediated targeting has facilitated the generation of knockin and knockout animal models via manipulation of germline competent stem cells, dramatically advancing many areas of biological research. However, although HR-mediated gene targeting produces highly precise alterations, the desired recombination events occur extremely infrequently (1 in 106–109 cells) (Capecchi, 1989), presenting enormous challenges for large-scale applications of gene-targeting experiments.
Genome Editing Technologies Exploit Endogenous DNA Repair Machinery
To overcome these challenges, a series of programmable nuclease-based genome editing technologies have been developed in recent years, enabling targeted and efficient modification of a variety of eukaryotic and particularly mammalian species. Of the current generation of genome editing technologies, the most rapidly developing is the class of RNA-guided endonucleases known as Cas9 from the microbial adaptive immune system CRISPR (clustered regularly interspaced short palindromic repeats), which can be easily targeted to virtually any genomic location of choice by a short RNA guide. Here, we review the development and applications of the CRISPR-associated endonuclease Cas9 as a platform technology for achieving targeted perturbation of endogenous genomic elements and also discuss challenges and future avenues for innovation. ……
Figure 4Natural Mechanisms of Microbial CRISPR Systems in Adaptive Immunity
…… A key turning point came in 2005, when systematic analysis of the spacer sequences separating the individual direct repeats suggested their extrachromosomal and phage-associated origins (Mojica et al., 2005; Pourcel et al., 2005; Bolotin et al., 2005). This insight was tremendously exciting, especially given previous studies showing that CRISPR loci are transcribed (Tang et al., 2002) and that viruses are unable to infect archaeal cells carrying spacers corresponding to their own genomes (Mojica et al., 2005). Together, these findings led to the speculation that CRISPR arrays serve as an immune memory and defense mechanism, and individual spacers facilitate defense against bacteriophage infection by exploiting Watson-Crick base-pairing between nucleic acids (Mojica et al., 2005; Pourcel et al., 2005). Despite these compelling realizations that CRISPR loci might be involved in microbial immunity, the specific mechanism of how the spacers act to mediate viral defense remained a challenging puzzle. Several hypotheses were raised, including thoughts that CRISPR spacers act as small RNA guides to degrade viral transcripts in a RNAi-like mechanism (Makarova et al., 2006) or that CRISPR spacers direct Cas enzymes to cleave viral DNA at spacer-matching regions (Bolotin et al., 2005). …..
As the pace of CRISPR research accelerated, researchers quickly unraveled many details of each type of CRISPR system (Figure 4). Building on an earlier speculation that protospacer adjacent motifs (PAMs) may direct the type II Cas9 nuclease to cleave DNA (Bolotin et al., 2005), Moineau and colleagues highlighted the importance of PAM sequences by demonstrating that PAM mutations in phage genomes circumvented CRISPR interference (Deveau et al., 2008). Additionally, for types I and II, the lack of PAM within the direct repeat sequence within the CRISPR array prevents self-targeting by the CRISPR system. In type III systems, however, mismatches between the 5′ end of the crRNA and the DNA target are required for plasmid interference (Marraffini and Sontheimer, 2010). …..
In 2013, a pair of studies simultaneously showed how to successfully engineer type II CRISPR systems from Streptococcus thermophilus (Cong et al., 2013) andStreptococcus pyogenes (Cong et al., 2013; Mali et al., 2013a) to accomplish genome editing in mammalian cells. Heterologous expression of mature crRNA-tracrRNA hybrids (Cong et al., 2013) as well as sgRNAs (Cong et al., 2013; Mali et al., 2013a) directs Cas9 cleavage within the mammalian cellular genome to stimulate NHEJ or HDR-mediated genome editing. Multiple guide RNAs can also be used to target several genes at once. Since these initial studies, Cas9 has been used by thousands of laboratories for genome editing applications in a variety of experimental model systems (Sander and Joung, 2014). ……
The majority of CRISPR-based technology development has focused on the signature Cas9 nuclease from type II CRISPR systems. However, there remains a wide diversity of CRISPR types and functions. Cas RAMP module (Cmr) proteins identified in Pyrococcus furiosus and Sulfolobus solfataricus (Hale et al., 2012) constitute an RNA-targeting CRISPR immune system, forming a complex guided by small CRISPR RNAs that target and cleave complementary RNA instead of DNA. Cmr protein homologs can be found throughout bacteria and archaea, typically relying on a 5′ site tag sequence on the target-matching crRNA for Cmr-directed cleavage.
Unlike RNAi, which is targeted largely by a 6 nt seed region and to a lesser extent 13 other bases, Cmr crRNAs contain 30–40 nt of target complementarity. Cmr-CRISPR technologies for RNA targeting are thus a promising target for orthogonal engineering and minimal off-target modification. Although the modularity of Cmr systems for RNA-targeting in mammalian cells remains to be investigated, Cmr complexes native to P. furiosus have already been engineered to target novel RNA substrates (Hale et al., 2009, 2012). ……
Although Cas9 has already been widely used as a research tool, a particularly exciting future direction is the development of Cas9 as a therapeutic technology for treating genetic disorders. For a monogenic recessive disorder due to loss-of-function mutations (such as cystic fibrosis, sickle-cell anemia, or Duchenne muscular dystrophy), Cas9 may be used to correct the causative mutation. This has many advantages over traditional methods of gene augmentation that deliver functional genetic copies via viral vector-mediated overexpression—particularly that the newly functional gene is expressed in its natural context. For dominant-negative disorders in which the affected gene is haplosufficient (such as transthyretin-related hereditary amyloidosis or dominant forms of retinitis pigmentosum), it may also be possible to use NHEJ to inactivate the mutated allele to achieve therapeutic benefit. For allele-specific targeting, one could design guide RNAs capable of distinguishing between single-nucleotide polymorphism (SNP) variations in the target gene, such as when the SNP falls within the PAM sequence.
CRISPR/Cas9: a powerful genetic engineering tool for establishing large animal models of neurodegenerative diseases
Zhuchi Tu, Weili Yang, Sen Yan, Xiangyu Guo and Xiao-Jiang Li
Animal models are extremely valuable to help us understand the pathogenesis of neurodegenerative disorders and to find treatments for them. Since large animals are more like humans than rodents, they make good models to identify the important pathological events that may be seen in humans but not in small animals; large animals are also very important for validating effective treatments or confirming therapeutic targets. Due to the lack of embryonic stem cell lines from large animals, it has been difficult to use traditional gene targeting technology to establish large animal models of neurodegenerative diseases. Recently, CRISPR/Cas9 was used successfully to genetically modify genomes in various species. Here we discuss the use of CRISPR/Cas9 technology to establish large animal models that can more faithfully mimic human neurodegenerative diseases.
Neurodegenerative diseases — Alzheimer’s disease(AD),Parkinson’s disease(PD), amyotrophic lateral sclerosis (ALS), Huntington’s disease (HD), and frontotemporal dementia (FTD) — are characterized by age-dependent and selective neurodegeneration. As the life expectancy of humans lengthens, there is a greater prevalence of these neurodegenerative diseases; however, the pathogenesis of most of these neurodegenerative diseases remain unclear, and we lack effective treatments for these important brain disorders.
CRISPR/Cas9, Non-human primates, Neurodegenerative diseases, Animal model
There are a number of excellent reviews covering different types of neurodegenerative diseases and their genetic mouse models [8–12]. Investigations of different mouse models of neurodegenerative diseases have revealed a common pathology shared by these diseases. First, the development of neuropathology and neurological symptoms in genetic mouse models of neurodegenerative diseases is age dependent and progressive. Second, all the mouse models show an accumulation of misfolded or aggregated proteins resulting from the expression of mutant genes. Third, despite the widespread expression of mutant proteins throughout the body and brain, neuronal function appears to be selectively or preferentially affected. All these facts indicate that mouse models of neurodegenerative diseases recapitulate important pathologic features also seen in patients with neurodegenerative diseases.
However, it seems that mouse models can not recapitulate the full range of neuropathology seen in patients with neurodegenerative diseases. Overt neurodegeneration, which is the most important pathological feature in patient brains, is absent in genetic rodent models of AD, PD, and HD. Many rodent models that express transgenic mutant proteins under the control of different promoters do not replicate overt neurodegeneration, which is likely due to their short life spans and the different aging processes of small animals. Also important are the remarkable differences in brain development between rodents and primates. For example, the mouse brain takes 21 days to fully develop, whereas the formation of primate brains requires more than 150 days [13]. The rapid development of the brain in rodents may render neuronal cells resistant to misfolded protein-mediated neurodegeneration. Another difficulty in using rodent models is how to analyze cognitive and emotional abnormalities, which are the early symptoms of most neurodegenerative diseases in humans. Differences in neuronal circuitry, anatomy, and physiology between rodent and primate brains may also account for the behavioral differences between rodent and primate models.
Mitochondrial dynamics–fusion, fission, movement, and mitophagy–in neurodegenerative diseases
Neurons are metabolically active cells with high energy demands at locations distant from the cell body. As a result, these cells are particularly dependent on mitochondrial function, as reflected by the observation that diseases of mitochondrial dysfunction often have a neurodegenerative component. Recent discoveries have highlighted that neurons are reliant particularly on the dynamic properties of mitochondria. Mitochondria are dynamic organelles by several criteria. They engage in repeated cycles of fusion and fission, which serve to intermix the lipids and contents of a population of mitochondria. In addition, mitochondria are actively recruited to subcellular sites, such as the axonal and dendritic processes of neurons. Finally, the quality of a mitochondrial population is maintained through mitophagy, a form of autophagy in which defective mitochondria are selectively degraded. We review the general features of mitochondrial dynamics, incorporating recent findings on mitochondrial fusion, fission, transport and mitophagy. Defects in these key features are associated with neurodegenerative disease. Charcot-Marie-Tooth type 2A, a peripheral neuropathy, and dominant optic atrophy, an inherited optic neuropathy, result from a primary deficiency of mitochondrial fusion. Moreover, several major neurodegenerative diseases—including Parkinson’s, Alzheimer’s and Huntington’s disease—involve disruption of mitochondrial dynamics. Remarkably, in several disease models, the manipulation of mitochondrial fusion or fission can partially rescue disease phenotypes. We review how mitochondrial dynamics is altered in these neurodegenerative diseases and discuss the reciprocal interactions between mitochondrial fusion, fission, transport and mitophagy.
Applications of CRISPR–Cas systems in Neuroscience
Genome-editing tools, and in particular those based on CRISPR–Cas (clustered regularly interspaced short palindromic repeat (CRISPR)–CRISPR-associated protein) systems, are accelerating the pace of biological research and enabling targeted genetic interrogation in almost any organism and cell type. These tools have opened the door to the development of new model systems for studying the complexity of the nervous system, including animal models and stem cell-derived in vitro models. Precise and efficient gene editing using CRISPR–Cas systems has the potential to advance both basic and translational neuroscience research.
Cellular neuroscience, DNA recombination, Genetic engineering, Molecular neuroscience
Figure 3: In vitro applications of Cas9 in human iPSCs.close
a | Evaluation of disease candidate genes from large-population genome-wide association studies (GWASs). Human primary cells, such as neurons, are not easily available and are difficult to expand in culture. By contrast, induced pluripo…
The development of the CRISPR/Cas9 system has made gene editing a relatively simple task. While CRISPR and other gene editing technologies stand to revolutionize biomedical research and offers many promising therapeutic avenues (such as in the treatment of HIV), a great deal of debate exists over whether CRISPR should be used to modify human embryos. As I discussed in my previous Insight article, we lack enough fundamental biological knowledge to enhance many traits like height or intelligence, so we are not near a future with genetically-enhanced super babies. However, scientists have identified a few rare genetic variants that protect against disease. One such protective variant is a mutation in the APP gene that protects against Alzheimer’s disease and cognitive decline in old age. If we can perfect gene editing technologies, is this mutation one that we should be regularly introducing into embryos? In this article, I explore the potential for using gene editing as a way to prevent Alzheimer’s disease in future generations. Alzheimer’s Disease: Medicine’s Greatest Challenge in the 21st Century Can gene editing be the missing piece in the battle against Alzheimer’s? (Source: bostonbiotech.org) I chose to assess the benefit of germline gene editing in the context of Alzheimer’s disease because this disease is one of the biggest challenges medicine faces in the 21st century. Alzheimer’s disease is a chronic neurodegenerative disease responsible for the majority of the cases of dementia in the elderly. The disease symptoms begins with short term memory loss and causes more severe symptoms – problems with language, disorientation, mood swings, behavioral issues – as it progresses, eventually leading to the loss of bodily functions and death. Because of the dementia the disease causes, Alzheimer’s patients require a great deal of care, and the world spends ~1% of its total GDP on caring for those with Alzheimer’s and related disorders. Because the prevalence of the disease increases with age, the situation will worsen as life expectancies around the globe increase: worldwide cases of Alzheimer’s are expected to grow from 35 million today to over 115 million by 2050.
Despite much research, the exact causes of Alzheimer’s disease remains poorly understood. The disease seems to be related to the accumulation of plaques made of amyloid-β peptides that form on the outside of neurons, as well as the formation of tangles of the protein tau inside of neurons. Although many efforts have been made to target amyloid-β or the enzymes involved in its formation, we have so far been unsuccessful at finding any treatment that stops the disease or reverses its progress. Some researchers believe that most attempts at treating Alzheimer’s have failed because, by the time a patient shows symptoms, the disease has already progressed past the point of no return.
While research towards a cure continues, researchers have sought effective ways to prevent Alzheimer’s disease. Although some studies show that mental and physical exercise may lower ones risk of Alzheimer’s disease, approximately 60-80% of the risk for Alzheimer’s disease appears to be genetic. Thus, if we’re serious about prevention, we may have to act at the genetic level. And because the brain is difficult to access surgically for gene therapy in adults, this means using gene editing on embryos.
With the latest CRISPR/Cas9 advance, the exhortation “turn on, tune in, drop out” comes to mind. The CRISPR/Cas9 gene-editing system was already a well-known means of “tuning in” (inserting new genes) and “dropping out” (knocking out genes). But when it came to “turning on” genes, CRISPR/Cas9 had little potency. That is, it had demonstrated only limited success as a way to activate specific genes.
A new CRISPR/Cas9 approach, however, appears capable of activating genes more effectively than older approaches. The new approach may allow scientists to more easily determine the function of individual genes, according to Feng Zhang, Ph.D., a researcher at MIT and the Broad Institute. Dr. Zhang and colleagues report that the new approach permits multiplexed gene activation and rapid, large-scale studies of gene function.
The new technique was introduced in the December 10 online edition of Nature, in an article entitled, “Genome-scale transcriptional activation by an engineered CRISPR-Cas9 complex.” The article describes how Dr. Zhang, along with the University of Tokyo’s Osamu Nureki, Ph.D., and Hiroshi Nishimasu, Ph.D., overhauled the CRISPR/Cas9 system. The research team based their work on their analysis (published earlier this year) of the structure formed when Cas9 binds to the guide RNA and its target DNA. Specifically, the team used the structure’s 3D shape to rationally improve the system.
In previous efforts to revamp CRISPR/Cas9 for gene activation purposes, scientists had tried to attach the activation domains to either end of the Cas9 protein, with limited success. From their structural studies, the MIT team realized that two small loops of the RNA guide poke out from the Cas9 complex and could be better points of attachment because they allow the activation domains to have more flexibility in recruiting transcription machinery.
Using their revamped system, the researchers activated about a dozen genes that had proven difficult or impossible to turn on using the previous generation of Cas9 activators. Each gene showed at least a twofold boost in transcription, and for many genes, the researchers found multiple orders of magnitude increase in activation.
After investigating single-guide RNA targeting rules for effective transcriptional activation, demonstrating multiplexed activation of 10 genes simultaneously, and upregulating long intergenic noncoding RNA transcripts, the research team decided to undertake a large-scale screen. This screen was designed to identify genes that confer resistance to a melanoma drug called PLX-4720.
“We … synthesized a library consisting of 70,290 guides targeting all human RefSeq coding isoforms to screen for genes that, upon activation, confer resistance to a BRAF inhibitor,” wrote the authors of the Nature paper. “The top hits included genes previously shown to be able to confer resistance, and novel candidates were validated using individual [single-guide RNA] and complementary DNA overexpression.”
A gene signature based on the top screening hits, the authors added, correlated with a gene expression signature of BRAF inhibitor resistance in cell lines and patient-derived samples. It was also suggested that large-scale screens such as the one demonstrated in the current study could help researchers discover new cancer drugs that prevent tumors from becoming resistant.
Familial amyloid polyneuropathy type I is an autosomal dominant disorder caused by mutations in the transthyretin (TTR ) gene; however, carriers of the same mutation exhibit variability in penetrance and clinical expression. We analyzed alleles of candidate genes encoding non-fibrillar components of TTR amyloid deposits and a molecule metabolically interacting with TTR [retinol-binding protein (RBP)], for possible associations with age of disease onset and/or susceptibility in a Portuguese population sample with the TTR V30M mutation and unrelated controls. We show that the V30M carriers represent a distinct subset of the Portuguese population. Estimates of genetic distance indicated that the controls and the classical onset group were furthest apart, whereas the late-onset group appeared to differ from both. Importantly, the data also indicate that genetic interactions among the multiple loci evaluated, rather than single-locus effects, are more likely to determine differences in the age of disease onset. Multifactor dimensionality reduction indicated that the best genetic model for classical onset group versus controls involved the APCS gene, whereas for late-onset cases, one APCS variant (APCSv1) and two RBP variants (RBPv1 and RBPv2) are involved. Thus, although the TTR V30M mutation is required for the disease in Portuguese patients, different genetic factors may govern the age of onset, as well as the occurrence of anticipation.
Autosomal dominant disorders may vary in expression even within a given kindred. The basis of this variability is uncertain and can be attributed to epigenetic factors, environment or epistasis. We have studied familial amyloid polyneuropathy (FAP), an autosomal dominant disorder characterized by peripheral sensorimotor and autonomic neuropathy. It exhibits variation in cardiac, renal, gastrointestinal and ocular involvement, as well as age of onset. Over 80 missense mutations in the transthyretin gene (TTR ) result in autosomal dominant disease http://www.ibmc.up.pt/~mjsaraiv/ttrmut.html). The presence of deposits consisting entirely of wild-type TTR molecules in the hearts of 10– 25% of individuals over age 80 reveals its inherent in vivo amyloidogenic potential (1).
FAP was initially described in Portuguese (2) where, until recently, the TTR V30M has been the only pathogenic mutation associated with the disease (3,4). Later reports identified the same mutation in Swedish and Japanese families (5,6). The disorder has since been recognized in other European countries and in North American kindreds in association with V30M, as well as other mutations (7).
TTR V30M produces disease in only 5–10% of Swedish carriers of the allele (8), a much lower degree of penetrance than that seen in Portuguese (80%) (9) or in Japanese with the same mutation. The actual penetrance in Japanese carriers has not been formally established, but appears to resemble that seen in Portuguese. Portuguese and Japanese carriers show considerable variation in the age of clinical onset (10,11). In both populations, the first symptoms had originally been described as typically occurring before age 40 (so-called ‘classical’ or early-onset); however, in recent years, more individuals developing symptoms late in life have been identified (11,12). Hence, present data indicate that the distribution of the age of onset in Portuguese is continuous, but asymmetric with a mean around age 35 and a long tail into the older age group (Fig. 1) (9,13). Further, DNA testing in Portugal has identified asymptomatic carriers over age 70 belonging to a subset of very late-onset kindreds in whose descendants genetic anticipation is frequent. The molecular basis of anticipation in FAP, which is not mediated by trinucleotide repeat expansions in the TTR or any other gene (14), remains elusive.
Variation in penetrance, age of onset and clinical features are hallmarks of many autosomal dominant disorders including the human TTR amyloidoses (7). Some of these clearly reflect specific biological effects of a particular mutation or a class of mutants. However, when such phenotypic variability is seen with a single mutation in the gene encoding the same protein, it suggests an effect of modifying genetic loci and/or environmental factors contributing differentially to the course of disease. We have chosen to examine age of onset as an example of a discrete phenotypic variation in the presence of the particular autosomal dominant disease-associated mutation TTR V30M. Although the role of environmental factors cannot be excluded, the existence of modifier genes involved in TTR amyloidogenesis is an attractive hypothesis to explain the phenotypic variability in FAP. ….
ATTR (TTR amyloid), like all amyloid deposits, contains several molecular components, in addition to the quantitatively dominant fibril-forming amyloid protein, including heparan sulfate proteoglycan 2 (HSPG2 or perlecan), SAP, a plasma glycoprotein of the pentraxin family (encoded by the APCS gene) that undergoes specific calcium-dependent binding to all types of amyloid fibrils, and apolipoprotein E (ApoE), also found in all amyloid deposits (15). The ApoE4 isoform is associated with an increased frequency and earlier onset of Alzheimer’s disease (Ab), the most common form of brain amyloid, whereas the ApoE2 isoform appears to be protective (16). ApoE variants could exert a similar modulatory effect in the onset of FAP, although early studies on a limited number of patients suggested this was not the case (17).
In at least one instance of senile systemic amyloidosis, small amounts of AA-related material were found in TTR deposits (18). These could reflect either a passive co-aggregation or a contributory involvement of protein AA, encoded by the serum amyloid A (SAA ) genes and the main component of secondary (reactive) amyloid fibrils, in the formation of ATTR.
Retinol-binding protein (RBP), the serum carrier of vitamin A, circulates in plasma bound to TTR. Vitamin A-loaded RBP and L-thyroxine, the two natural ligands of TTR, can act alone or synergistically to inhibit the rate and extent of TTR fibrillogenesis in vitro, suggesting that RBP may influence the course of FAP pathology in vivo (19). We have analyzed coding and non-coding sequence polymorphisms in the RBP4 (serum RBP, 10q24), HSPG2 (1p36.1), APCS (1q22), APOE (19q13.2), SAA1 and SAA2 (11p15.1) genes with the goal of identifying chromosomes carrying common and functionally significant variants. At the time these studies were performed, the full human genome sequence was not completed and systematic singlenucleotide polymorphism (SNP) analyses were not available for any of the suspected candidate genes. We identified new SNPs in APCS and RBP4 and utilized polymorphisms in SAA, HSPG2 and APOE that had already been characterized and shown to have potential pathophysiologic significance in other disorders (16,20–22). The genotyping data were analyzed for association with the presence of the V30M amyloidogenic allele (FAP patients versus controls) and with the age of onset (classical- versus late-onset patients). Multilocus analyses were also performed to examine the effects of simultaneous contributions of the six loci for determining the onset of the first symptoms. …..
The potential for different underlying models for classical and late onset is supported by the MDR analysis, which produces two distinct models when comparing each class with the controls. One could view the two onset classes as unique diseases. If this is the case, then the failure to detect a single predictive genetic model is consistent with two related, but different, diseases. This is exactly what would be expected in such a case of genetic heterogeneity (28). Using this approach, a major gene effect can be viewed as a necessary, but not sufficient, condition to explain the course of the disease. Analyzing the cases but omitting from the analysis of phenotype the necessary allele, in this case TTR V30M, can then reveal a variety of important modifiers that are distinct between the phenotypes.
The significant comparisons obtained in our study cohort indicate that the combined effects mainly result from two and three-locus interactions involving all loci except SAA1 and SAA2 for susceptibility to disease. A considerable number of four-site combinations modulate the age of onset with SAA1 appearing in a majority of significant combinations in late-onset disease, perhaps indicating a greater role of the SAA variants in the age of onset of FAP.
The correlation between genotype and phenotype in socalled simple Mendelian disorders is often incomplete, as only a subset of all mutations can reliably predict specific phenotypes (34). This is because non-allelic genetic variations and/or environmental influences underlie these disorders whose phenotypes behave as complex traits. A few examples include the identification of the role of homozygozity for the SAA1.1 allele in conferring the genetic susceptibility to renal amyloidosis in FMF (20) and the association of an insertion/deletion polymorphism in the ACE gene with disease severity in familial hypertrophic cardiomyopathy (35). In these disorders, the phenotypes arise from mutations in MEFV and b-MHC, but are modulated by independently inherited genetic variation. In this report, we show that interactions among multiple genes, whose products are confirmed or putative constituents of ATTR deposits, or metabolically interact with TTR, modulate the onset of the first symptoms and predispose individuals to disease in the presence of the V30M mutation in TTR. The exact nature of the effects identified here requires further study with potential application in the development of genetic screening with prognostic value pertaining to the onset of disease in the TTR V30M carriers.
If the effects of additional single or interacting genes dictate the heterogeneity of phenotype, as reflected in variability of onset and clinical expression (with the same TTR mutation), the products encoded by alleles at such loci could contribute to the process of wild-type TTR deposition in elderly individuals without a mutation (senile systemic amyloidosis), a phenomenon not readily recognized as having a genetic basis because of the insensitivity of family history in the elderly.
Safety and Efficacy of RNAi Therapy for Transthyretin Amyloidosis
Transthyretin amyloidosis is caused by the deposition of hepatocyte-derived transthyretin amyloid in peripheral nerves and the heart. A therapeutic approach mediated by RNA interference (RNAi) could reduce the production of transthyretin.
Methods We identified a potent antitransthyretin small interfering RNA, which was encapsulated in two distinct first- and second-generation formulations of lipid nanoparticles, generating ALN-TTR01 and ALN-TTR02, respectively. Each formulation was studied in a single-dose, placebo-controlled phase 1 trial to assess safety and effect on transthyretin levels. We first evaluated ALN-TTR01 (at doses of 0.01 to 1.0 mg per kilogram of body weight) in 32 patients with transthyretin amyloidosis and then evaluated ALN-TTR02 (at doses of 0.01 to 0.5 mg per kilogram) in 17 healthy volunteers.
Results Rapid, dose-dependent, and durable lowering of transthyretin levels was observed in the two trials. At a dose of 1.0 mg per kilogram, ALN-TTR01 suppressed transthyretin, with a mean reduction at day 7 of 38%, as compared with placebo (P=0.01); levels of mutant and nonmutant forms of transthyretin were lowered to a similar extent. For ALN-TTR02, the mean reductions in transthyretin levels at doses of 0.15 to 0.3 mg per kilogram ranged from 82.3 to 86.8%, with reductions of 56.6 to 67.1% at 28 days (P<0.001 for all comparisons). These reductions were shown to be RNAi mediated. Mild-to-moderate infusion-related reactions occurred in 20.8% and 7.7% of participants receiving ALN-TTR01 and ALN-TTR02, respectively.
ALN-TTR01 and ALN-TTR02 suppressed the production of both mutant and nonmutant forms of transthyretin, establishing proof of concept for RNAi therapy targeting messenger RNA transcribed from a disease-causing gene.
Alnylam May Seek Approval for TTR Amyloidosis Rx in 2017 as Other Programs Advance
Officials from Alnylam Pharmaceuticals last week provided updates on the two drug candidates from the company’s flagship transthyretin-mediated amyloidosis program, stating that the intravenously delivered agent patisiran is proceeding toward a possible market approval in three years, while a subcutaneously administered version called ALN-TTRsc is poised to enter Phase III testing before the end of the year.
Meanwhile, Alnylam is set to advance a handful of preclinical therapies into human studies in short order, including ones for complement-mediated diseases, hypercholesterolemia, and porphyria.
The officials made their comments during a conference call held to discuss Alnylam’s second-quarter financial results.
ATTR is caused by a mutation in the TTR gene, which normally produces a protein that acts as a carrier for retinol binding protein and is characterized by the accumulation of amyloid deposits in various tissues. Alnylam’s drugs are designed to silence both the mutant and wild-type forms of TTR.
Patisiran, which is delivered using lipid nanoparticles developed by Tekmira Pharmaceuticals, is currently in a Phase III study in patients with a form of ATTR called familial amyloid polyneuropathy (FAP) affecting the peripheral nervous system. Running at over 20 sites in nine countries, that study is set to enroll up to 200 patients and compare treatment to placebo based on improvements in neuropathy symptoms.
According to Alnylam Chief Medical Officer Akshay Vaishnaw, Alnylam expects to have final data from the study in two to three years, which would put patisiran on track for a new drug application filing in 2017.
Meanwhile, ALN-TTRsc, which is under development for a version of ATTR that affects cardiac tissue called familial amyloidotic cardiomyopathy (FAC) and uses Alnylam’s proprietary GalNAc conjugate delivery technology, is set to enter Phase III by year-end as Alnylam holds “active discussions” with US and European regulators on the design of that study, CEO John Maraganore noted during the call.
In the interim, Alnylam continues to enroll patients in a pilot Phase II study of ALN-TTRsc, which is designed to test the drug’s efficacy for FAC or senile systemic amyloidosis (SSA), a condition caused by the idiopathic accumulation of wild-type TTR protein in the heart.
Based on “encouraging” data thus far, Vaishnaw said that Alnylam has upped the expected enrollment in this study to 25 patients from 15. Available data from the trial is slated for release in November, he noted, stressing that “any clinical endpoint result needs to be considered exploratory given the small sample size and the very limited duration of treatment of only six weeks” in the trial.
Vaishnaw added that an open-label extension (OLE) study for patients in the ALN-TTRsc study will kick off in the coming weeks, allowing the company to gather long-term dosing tolerability and clinical activity data on the drug.
Enrollment in an OLE study of patisiran has been completed with 27 patients, he said, and, “as of today, with up to nine months of therapy … there have been no study drug discontinuations.” Clinical endpoint data from approximately 20 patients in this study will be presented at the American Neurological Association meeting in October.
As part of its ATTR efforts, Alnylam has also been conducting natural history of disease studies in both FAP and FAC patients. Data from the 283-patient FAP study was presented earlier this year and showed a rapid progression in neuropathy impairment scores and a high correlation of this measurement with disease severity.
During last week’s conference call, Vaishnaw said that clinical endpoint and biomarker data on about 400 patients with either FAC or SSA have already been collected in a nature history study on cardiac ATTR. Maraganore said that these findings would likely be released sometime next year.
The first medication for a rare and often fatal protein misfolding disorder has been approved in Europe. On November 16, the E gave a green light to Pfizer’s Vyndaqel (tafamidis) for treating transthyretin amyloidosis in adult patients with stage 1 polyneuropathy symptoms. [Jeffery Kelly, La Jolla]
The most clinically advanced RNA interference (RNAi) therapeutic achieved a milestone in April when Alnylam Pharmaceuticals in Cambridge, Massachusetts, reported positive results for patisiran, a small interfering RNA (siRNA) oligonucleotide targeting transthyretin for treating familial amyloidotic polyneuropathy (FAP). …
FAP is characterized by the systemic deposition of amyloidogenic variants of the transthyretin protein, especially in the peripheral nervous system, causing a progressive sensory and motor polyneuropathy.
FAP is caused by a mutation of the TTR gene, located on human chromosome 18q12.1-11.2.[5] A replacement of valine by methionine at position 30 (TTR V30M) is the mutation most commonly found in FAP.[1] The variant TTR is mostly produced by the liver.[citation needed] The transthyretin protein is a tetramer. ….
An open mind and collaborative spirit have taken Hans Clevers on a journey from medicine to developmental biology, gastroenterology, cancer, and stem cells.
Ihave had to talk a lot about my science recently and it’s made me think about how science works,” says Hans Clevers. “Scientists are trained to think science is driven by hypotheses, but for [my lab], hypothesis-driven research has never worked. Instead, it has been about trying to be as open-minded as possible—which is not natural for our brains,” adds the Utrecht University molecular genetics professor. “The human mind is such that it tries to prove it’s right, so pursuing a hypothesis can result in disaster. My advice to my own team and others is to not preformulate an answer to a scientific question, but just observe and never be afraid of the unknown. What has worked well for us is to keep an open mind and do the experiments. And find a collaborator if it is outside our niche.”
“One thing I have learned is that hypothesis-driven research tends not to be productive when you are in an unknown territory.”
Clevers entered medical school at Utrecht University in The Netherlands in 1978 while simultaneously pursuing a master’s degree in biology. Drawn to working with people in the clinic, Clevers had a training position in pediatrics lined up after medical school, but then mentors persuaded him to spend an additional year converting the master’s degree to a PhD in immunology. “At the end of that year, looking back, I got more satisfaction from the research than from seeing patients.” Clevers also had an aptitude for benchwork, publishing four papers from his PhD year. “They were all projects I had made up myself. The department didn’t do the kind of research I was doing,” he says. “Now that I look back, it’s surprising that an inexperienced PhD student could come up with a project and publish independently.”
Clevers studied T- and B-cell signaling; he set up assays to visualize calcium ion flux and demonstrated that the ions act as messengers to activate human B cells, signaling through antibodies on the cell surface. “As soon as the experiment worked, I got T cells from the lab next door and did the same experiment. That was my strategy: as soon as something worked, I would apply it elsewhere and didn’t stop just because I was a B-cell biologist and not a T-cell biologist. What I learned then, that I have continued to benefit from, is that a lot of scientists tend to adhere to a niche. They cling to these niches and are not that flexible. You think scientists are, but really most are not.”
Here, Clevers talks about promoting a collaborative spirit in research, the art of doing a pilot experiment, and growing miniature organs in a dish.
Clevers Creates
Re-search? Clevers was born in Eindhoven, in the south of The Netherlands. The town was headquarters to Philips Electronics, where his father worked as a businessman, and his mother took care of Clevers and his three brothers. Clevers did well in school but his passion was sports, especially tennis and field hockey, “a big thing in Holland.” Then in 1975, at age 18, he moved to Utrecht University, where he entered an intensive, biology-focused program. “I knew I wanted to be a biology researcher since I was young. In Dutch, the word for research is ‘onderzoek’ and I knew the English word ‘research’ and had wondered why there was the ‘re’ in the word, because I wanted to search but I didn’t want to do re-search—to find what someone else had already found.”
Opportunity to travel. “I was very disappointed in my biology studies, which were old-fashioned and descriptive,” says Clevers. He thought medicine might be more interesting and enrolled in medical school while still pursuing a master’s degree in biology at Utrecht. For the master’s, Clevers had to do three rotations. He spent a year at the International Laboratory for Research on Animal Diseases (ILRAD) in Nairobi, Kenya, and six months in Bethesda, Maryland, at the National Institutes of Health. “Holland is really small, so everyone travels.” Clevers saw those two rotations more as travel explorations. In Nairobi, he went on safaris and explored the country in Land Rovers borrowed from the institute. While in Maryland in 1980, Clevers—with the consent of his advisor, who thought it was a good idea for him to get a feel for the U.S.—flew to Portland, Oregon, and drove back to Boston with a musician friend along the Canadian border. He met the fiancé of political activist and academic Angela Davis in New York City and even stayed in their empty apartment there.
Life and lab lessons. Back in Holland, Clevers joined Rudolf Eugène Ballieux’s lab at Utrecht University to pursue his PhD, for which he studied immune cell signaling. “I didn’t learn much science from him, but I learned that you always have to create trust and to trust people around you. This became a major theme in my own lab. We don’t distrust journals or reviewers or collaborators. We trust everyone and we share. There will be people who take advantage, but there have only been a few of those. So I learned from Ballieux to give everyone maximum trust and then change this strategy only if they fail that trust. We collaborate easily because we give out everything and we also easily get reagents and tools that we may need. It’s been valuable to me in my career. And it is fun!”
Clevers Concentrates
On a mission. “Once I decided to become a scientist, I knew I needed to train seriously. Up to that point, I was totally self-trained.” From an extensive reading of the immunology literature, Clevers became interested in how T cells recognize antigens, and headed off to spend a postdoc studying the problem in Cox Terhorst’s lab at Dana-Farber Cancer Institute in Boston. “Immunology was young, but it was very exciting and there was a lot to discover. I became a professional scientist there and experienced how tough science is.” In 1988, Clevers cloned and characterized the gene for a component of the T-cell receptor (TCR) called CD3-epsilon, which binds antigen and activates intracellular signaling pathways.
On the fast track in Holland. Clevers returned to Utrecht University in 1989 as a professor of immunology. Within one month of setting up his lab, he had two graduate students and a technician, and the lab had cloned the first T cell–specific transcription factor, which they called TCF-1, in human T cells. When his former thesis advisor retired, Clevers was asked, at age 33, to become head of the immunology department. While the appointment was high-risk for him and for the department, Clevers says, he was chosen because he was good at multitasking and because he got along well with everyone.
Problem-solving strategy. “My strategy in research has always been opportunistic. One thing I have learned is that hypothesis-driven research tends not to be productive when you are in an unknown territory. I think there is an art to doing pilot experiments. So we have always just set up systems in which something happens and then you try and try things until a pattern appears and maybe you formulate a small hypothesis. But as soon as it turns out not to be exactly right, you abandon it. It’s a very open-minded type of research where you question whether what you are seeing is a real phenomenon without spending a year on doing all of the proper controls.”
Trial and error. Clevers’s lab found that while TCF-1 bound to DNA, it did not alter gene expression, despite the researchers’ tinkering with promoter and enhancer assays. “For about five years this was a problem. My first PhD students were leaving and they thought the whole TCF project was a failure,” says Clevers. His lab meanwhile cloned TCF homologs from several model organisms and made many reagents including antibodies against these homologs. To try to figure out the function of TCF-1, the lab performed a two-hybrid screen and identified components of the Wnt signaling pathway as binding partners of TCF-1. “We started to read about Wnt and realized that you study Wnt not in T cells but in frogs and flies, so we rapidly transformed into a developmental biology lab. We showed that we held the key for a major issue in developmental biology, the final protein in the Wnt cascade: TCF-1 binds b-catenin when b-catenin becomes available and activates transcription.” In 1996, Clevers published the mechanism of how the TCF-1 homolog in Xenopus embryos, called XTcf-3, is integrated into the Wnt signaling pathway.
Clevers Catapults
COURTESY OF HANS CLEVERS AND JEROEN HUIJBEN, NYMUS
3DCrypt building and colon cancer.
Clevers next collaborated with Bert Vogelstein’s lab at Johns Hopkins, linking TCF to Wnt signaling in colon cancer. In colon cancer cell lines with mutated forms of the tumor suppressor gene APC, the APC protein can’t rein in b-catenin, which accumulates in the cytoplasm, forms a complex with TCF-4 (later renamed TCF7L2) in the nucleus, and caninitiate colon cancer by changing gene expression. Then, the lab showed that Wnt signaling is necessary for self-renewal of adult stem cells, as mice missing TCF-4 do not have intestinal crypts, the site in the gut where stem cells reside. “This was the first time Wnt was shown to play a role in adults, not just during development, and to be crucial for adult stem cell maintenance,” says Clevers. “Then, when I started thinking about studying the gut, I realized it was by far the best way to study stem cells. And I also realized that almost no one in the world was studying the healthy gut. Almost everyone who researched the gut was studying a disease.” The main advantages of the murine model are rapid cell turnover and the presence of millions of stereotypic crypts throughout the entire intestine.
Against the grain. In 2007, Nick Barker, a senior scientist in the Clevers lab, identified the Wnt target gene Lgr5 as a unique marker of adult stem cells in several epithelial organs, including the intestine, hair follicle, and stomach. In the intestine, the gene codes for a plasma membrane protein on crypt stem cells that enable the intestinal epithelium to self-renew, but can also give rise to adenomas of the gut. Upon making mice with adult stem cell populations tagged with a fluorescent Lgr5-binding marker, the lab helped to overturn assumptions that “stem cells are rare, impossible to find, quiescent, and divide asymmetrically.”
On to organoids. Once the lab could identify adult stem cells within the crypts of the gut, postdoc Toshiro Sato discovered that a single stem cell, in the presence of Matrigel and just three growth factors, could generate a miniature crypt structure—what is now called an organoid. “Toshi is very Japanese and doesn’t always talk much,” says Clevers. “One day I had asked him, while he was at the microscope, if the gut stem cells were growing, and he said, ‘Yes.’ Then I looked under the microscope and saw the beautiful structures and said, ‘Why didn’t you tell me?’ and he said, ‘You didn’t ask.’ For three months he had been growing them!” The lab has since also grown mini-pancreases, -livers, -stomachs, and many other mini-organs.
Tumor Organoids. Clevers showed that organoids can be grown from diseased patients’ samples, a technique that could be used in the future to screen drugs. The lab is also building biobanks of organoidsderived from tumor samples and adjacent normal tissue, which could be especially useful for monitoring responses to chemotherapies. “It’s a similar approach to getting a bacterium cultured to identify which antibiotic to take. The most basic goal is not to give a toxic chemotherapy to a patient who will not respond anyway,” says Clevers. “Tumor organoids grow slower than healthy organoids, which seems counterintuitive, but with cancer cells, often they try to divide and often things go wrong because they don’t have normal numbers of chromosomes and [have] lots of mutations. So, I am not yet convinced that this approach will work for every patient. Sometimes, the tumor organoids may just grow too slowly.”
Selective memory. “When I received the Breakthrough Prize in 2013, I invited everyone who has ever worked with me to Amsterdam, about 100 people, and the lab organized a symposium where many of the researchers gave an account of what they had done in the lab,” says Clevers. “In my experience, my lab has been a straight line from cloning TCF-1 to where we are now. But when you hear them talk it was ‘Hans told me to try this and stop this’ and ‘Half of our knockout mice were never published,’ and I realized that the lab is an endless list of failures,” Clevers recalls. “The one thing we did well is that we would start something and, as soon as it didn’t look very good, we would stop it and try something else. And the few times when we seemed to hit gold, I would regroup my entire lab. We just tried a lot of things, and the 10 percent of what worked, those are the things I remember.”
Greatest Hits
Cloned the first T cell–specific transcription factor, TCF-1, and identified homologous genes in model organisms including the fruit fly, frog, and worm
Found that transcriptional activation by the abundant β-catenin/TCF-4 [TCF7L2] complex drives cancer initiation in colon cells missing the tumor suppressor protein APC
First to extend the role of Wnt signaling from developmental biology to adult stem cells by showing that the two Wnt pathway transcription factors, TCF-1 and TCF-4, are necessary for maintaining the stem cell compartments in the thymus and in the crypt structures of the small intestine, respectively
Identified Lgr5 as an adult stem cell marker of many epithelial stem cells including those of the colon, small intestine, hair follicle, and stomach, and found that Lgr5-expressing crypt cells in the small intestine divide constantly and symmetrically, disproving the common belief that stem cell division is asymmetrical and uncommon
Established a three-dimensional, stable model, the “organoid,” grown from adult stem cells, to study diseased patients’ tissues from the gut, stomach, liver, and prostate
Regenerative Medicine Comes of Age
“Anti-Aging Medicine” Sounds Vaguely Disreputable, So Serious Scientists Prefer to Speak of “Regenerative Medicine”
Induced pluripotent stem cells (iPSCs) and genome-editing techniques have facilitated manipulation of living organisms in innumerable ways at the cellular and genetic levels, respectively, and will underpin many aspects of regenerative medicine as it continues to evolve.
An attitudinal change is also occurring. Experts in regenerative medicine have increasingly begun to embrace the view that comprehensively repairing the damage of aging is a practical and feasible goal.
A notable proponent of this view is Aubrey de Grey, Ph.D., a biomedical gerontologist who has pioneered an regenerative medicine approach called Strategies for Engineered Negligible Senescence (SENS). He works to “develop, promote, and ensure widespread access to regenerative medicine solutions to the disabilities and diseases of aging” as CSO and co-founder of the SENS Research Foundation. He is also the editor-in-chief of Rejuvenation Research, published by Mary Ann Liebert.
Dr. de Grey points out that stem cell treatments for age-related conditions such as Parkinson’s are already in clinical trials, and immune therapies to remove molecular waste products in the extracellular space, such as amyloid in Alzheimer’s, have succeeded in such trials. Recently, there has been progress in animal models in removing toxic cells that the body is failing to kill. The most encouraging work is in cancer immunotherapy, which is rapidly advancing after decades in the doldrums.
Many damage-repair strategies are at an early stage of research. Although these strategies look promising, they are handicapped by a lack of funding. If that does not change soon, the scientific community is at risk of failing to capitalize on the relevant technological advances.
Regenerative medicine has moved beyond boutique applications. In degenerative disease, cells lose their function or suffer elimination because they harbor genetic defects. iPSC therapies have the potential to be curative, replacing the defective cells and eliminating symptoms in their entirety. One of the biggest hurdles to commercialization of iPSC therapies is manufacturing.
Building Stem Cell Factories
Cellular Dynamics International (CDI) has been developing clinically compatible induced pluripotent stem cells (iPSCs) and iPSC-derived human retinal pigment epithelial (RPE) cells. CDI’s MyCell Retinal Pigment Epithelial Cells are part of a possible therapy for macular degeneration. They can be grown on bioengineered, nanofibrous scaffolds, and then the RPE cell–enriched scaffolds can be transplanted into patients’ eyes. In this pseudo-colored image, RPE cells are shown growing over the nanofibers. Each cell has thousands of “tongue” and “rod” protrusions that could naturally support rod and cone cells in the eye.
“Now that an infrastructure is being developed to make unlimited cells for the tools business, new opportunities are being created. These cells can be employed in a therapeutic context, and they can be used to understand the efficacy and safety of drugs,” asserts Chris Parker, executive vice president and CBO, Cellular Dynamics International (CDI). “CDI has the capability to make a lot of cells from a single iPSC line that represents one person (a capability termed scale-up) as well as the capability to do it in parallel for multiple individuals (a capability termed scale-out).”
Minimally manipulated adult stem cells have progressed relatively quickly to the clinic. In this scenario, cells are taken out of the body, expanded unchanged, then reintroduced. More preclinical rigor applies to potential iPSC therapy. In this case, hematopoietic blood cells are used to make stem cells, which are manufactured into the cell type of interest before reintroduction. Preclinical tests must demonstrate that iPSC-derived cells perform as intended, are safe, and possess little or no off-target activity.
For example, CDI developed a Parkinsonian model in which iPSC-derived dopaminergic neurons were introduced to primates. The model showed engraftment and enervation, and it appeared to be free of proliferative stem cells.
“You will see iPSCs first used in clinical trials as a surrogate to understand efficacy and safety,” notes Mr. Parker. “In an ongoing drug-repurposing trial with GlaxoSmithKline and Harvard University, iPSC-derived motor neurons will be produced from patients with amyotrophic lateral sclerosis and tested in parallel with the drug.” CDI has three cell-therapy programs in their commercialization pipeline focusing on macular degeneration, Parkinson’s disease, and postmyocardial infarction.
Keeping an Eye on Aging Eyes
The California Project to Cure Blindness is evaluating a stem cell–based treatment strategy for age-related macular degeneration. The strategy involves growing retinal pigment epithelium (RPE) cells on a biostable, synthetic scaffold, then implanting the RPE cell–enriched scaffold to replace RPE cells that are dying or dysfunctional. One of the project’s directors, Dennis Clegg, Ph.D., a researcher at the University of California, Santa Barbara, provided this image, which shows stem cell–derived RPE cells. Cell borders are green, and nuclei are red.
The eye has multiple advantages over other organ systems for regenerative medicine. Advanced surgical methods can access the back of the eye, noninvasive imaging methods can follow the transplanted cells, good outcome parameters exist, and relatively few cells are needed.
These advantages have attracted many groups to tackle ocular disease, in particular age-related macular degeneration, the leading cause of blindness in the elderly in the United States. Most cases of age-related macular degeneration are thought to be due to the death or dysfunction of cells in the retinal pigment epithelium (RPE). RPE cells are crucial support cells for the rods, cones, and photoreceptors. When RPE cells stop working or die, the photoreceptors die and a vision deficit results.
A regenerated and restored RPE might prevent the irreversible loss of photoreceptors, possibly via the the transplantation of functionally polarized RPE monolayers derived from human embryonic stem cells. This approach is being explored by the California Project to Cure Blindness, a collaborative effort involving the University of Southern California (USC), the University of California, Santa Barbara (UCSB), the California Institute of Technology, City of Hope, and Regenerative Patch Technologies.
The project, which is funded by the California Institute of Regenerative Medicine (CIRM), started in 2010, and an IND was filed early 2015. Clinical trial recruitment has begun.
One of the project’s leaders is Dennis Clegg, Ph.D., Wilcox Family Chair in BioMedicine, UCSB. His laboratory developed the protocol to turn undifferentiated H9 embryonic stem cells into a homogenous population of RPE cells.
“These are not easy experiments,” remarks Dr. Clegg. “Figuring out the biology and how to make the cell of interest is a challenge that everyone in regenerative medicine faces. About 100,000 RPE cells will be grown as a sheet on a 3 × 5 mm biostable, synthetic scaffold, and then implanted in the patients to replace the cells that are dying or dysfunctional. The idea is to preserve the photoreceptors and to halt disease progression.”
Moving therapies such as this RPE treatment from concept to clinic is a huge team effort and requires various kinds of expertise. Besides benefitting from Dr. Clegg’s contribution, the RPE project incorporates the work of Mark Humayun, M.D., Ph.D., co-director of the USC Eye Institute and director of the USC Institute for Biomedical Therapeutics and recipient of the National Medal of Technology and Innovation, and David Hinton, Ph.D., a researcher at USC who has studied how actvated RPE cells can alter the local retinal microenvironment.
Scientists at the Swiss Federal Institute of Technology (ETH) in Zurich have found an exciting new use for the cells that reside in the undesirable flabby tissue—creating pancreatic beta cells. The ETH researchers extracted stem cells from a 50-year-old test subject’s fatty tissue and reprogrammed them into mature, insulin-producing beta cells.
The findings from this study were published recently in Nature Communications in an article entitled “A Programmable Synthetic Lineage-Control Network That Differentiates Human IPSCs into Glucose-Sensitive Insulin-Secreting Beta-Like Cells.”
The investigators added a highly complex synthetic network of genes to the stem cells to recreate precisely the key growth factors involved in this maturation process. Central to the process were the growth factors Ngn3, Pdx1, and MafA; the researchers found that concentrations of these factors change during the differentiation process.
For instance, MafA is not present at the start of maturation. Only on day 4, in the final maturation step, does it appear, its concentration rising steeply and then remaining at a high level. The changes in the concentrations of Ngn3 and Pdx1, however, are very complex: while the concentration of Ngn3 rises and then falls again, the level of Pdx1 rises at the beginning and toward the end of maturation.
Senior study author Martin Fussenegger, Ph.D., professor of biotechnology and bioengineering at ETH Zurich’s department of biosystems science and engineering stressed that it was essential to reproduce these natural processes as closely as possible to produce functioning beta cells, stating that “the timing and the quantities of these growth factors are extremely important.”
The ETH researchers believe that their work is a real breakthrough, in that a synthetic gene network has been used successfully to achieve genetic reprogramming that delivers beta cells. Until now, scientists have controlled such stem cell differentiation processes by adding various chemicals and proteins exogenously.
“It’s not only really hard to add just the right quantities of these components at just the right time, but it’s also inefficient and impossible to scale up,” Dr. Fussenegger noted.
While the beta cells not only looked very similar to their natural counterparts—containing dark spots known as granules that store insulin—the artificial beta cells also functioned in a very similar manner. However, the researchers admit that more work needs to be done to increase the insulin output.
“At the present time, the quantities of insulin they secrete are not as great as with natural beta cells,” Dr. Fussenegger stated. Yet, the key point is that the researchers have for the first time succeeded in reproducing the entire natural process chain, from stem cell to differentiated beta cell.
In future, the ETH scientists’ novel technique might make it possible to implant new functional beta cells in diabetes sufferers that are made from their adipose tissue. While beta cells have been transplanted in the past, this has always required subsequent suppression of the recipient’s immune system—as with any transplant of donor organs or tissue.
“With our beta cells, there would likely be no need for this action since we can make them using endogenous cell material taken from the patient’s own body,” Dr. Fussenegger said. “This is why our work is of such interest in the treatment of diabetes.”
A programmable synthetic lineage-control network that differentiates human IPSCs into glucose-sensitive insulin-secreting beta-like cells
Synthetic biology has advanced the design of standardized transcription control devices that programme cellular behaviour. By coupling synthetic signalling cascade- and transcription factor-based gene switches with reverse and differential sensitivity to the licensed food additive vanillic acid, we designed a synthetic lineage-control network combining vanillic acid-triggered mutually exclusive expression switches for the transcription factors Ngn3 (neurogenin 3; OFF-ON-OFF) and Pdx1 (pancreatic and duodenal homeobox 1; ON-OFF-ON) with the concomitant induction of MafA (V-maf musculoaponeurotic fibrosarcoma oncogene homologue A; OFF-ON). This designer network consisting of different network topologies orchestrating the timely control of transgenic and genomic Ngn3, Pdx1 and MafA variants is able to programme human induced pluripotent stem cells (hIPSCs)-derived pancreatic progenitor cells into glucose-sensitive insulin-secreting beta-like cells, whose glucose-stimulated insulin-release dynamics are comparable to human pancreatic islets. Synthetic lineage-control networks may provide the missing link to genetically programme somatic cells into autologous cell phenotypes for regenerative medicine.
Cell-fate decisions during development are regulated by various mechanisms, including morphogen gradients, regulated activation and silencing of key transcription factors, microRNAs, epigenetic modification and lateral inhibition. The latter implies that the decision of one cell to adopt a specific phenotype is associated with the inhibition of neighbouring cells to enter the same developmental path. In mammals, insights into the role of key transcription factors that control development of highly specialized organs like the pancreas were derived from experiments in mice, especially various genetically modified animals1, 2, 3, 4. Normal development of the pancreas requires the activation of pancreatic duodenal homeobox protein (Pdx1) in pre-patterned cells of the endoderm. Inactivating mutations of Pdx1 are associated with pancreas agenesis in mouse and humans5, 6. A similar cell fate decision occurs later with the activation of Ngn3 that is required for the development of all endocrine cells in the pancreas7. Absence of Ngn3 is associated with the loss of pancreatic endocrine cells, whereas the activation of Ngn3 not only allows the differentiation of endocrine cells but also induces lateral inhibition of neighbouring cells—via Delta-Notch pathway—to enter the same pancreatic endocrine cell fate8. This Ngn3-mediated cell-switch occurs at a specific time point and for a short period of time in mice9. Thereafter, it is silenced and becomes almost undetectable in postnatal pancreatic islets. Conversely, Pdx1-positive Ngn3-positive cells reduce Pdx1 expression, as Ngn3-positive cells are Pdx1 negative10. They re-express Pdx1, however, as they go on their path towards glucose-sensitive insulin-secreting cells with parallel induction of MafA that is required for proper differentiation and maturation of pancreatic beta cells11. Data supporting these expression dynamics are derived from mice experiments1, 11, 12. A synthetic gene-switch governing cell fate decision in human induced pluripotent stem cells (hIPSCs) could facilitate the differentiation of glucose-sensitive insulin-secreting cells.
In recent years, synthetic biology has significantly advanced the rational design of synthetic gene networks that can interface with host metabolism, correct physiological disturbances13 and provide treatment strategies for a variety of metabolic disorders, including gouty arthritis14, obesity15 and type-2 diabetes16. Currently, synthetic biology principles may provide the componentry and gene network topologies for the assembly of synthetic lineage-control networks that can programme cell-fate decisions and provide targeted differentiation of stem cells into terminally differentiated somatic cells. Synthetic lineage-control networks may therefore provide the missing link between human pluripotent stem cells17 and their true impact on regenerative medicine18, 19, 20. The use of autologous stem cells in regenerative medicine holds great promise for curing many diseases, including type-1 diabetes mellitus (T1DM), which is characterized by the autoimmune destruction of insulin-producing pancreatic beta cells, thus making patients dependent on exogenous insulin to control their blood glucose21, 22. Although insulin therapy has changed the prospects and survival of T1DM patients, these patients still suffer from diabetic complications arising from the lack of physiological insulin secretion and excessive glucose levels23. The replacement of the pancreatic beta cells either by pancreas transplantation or by transplantation of pancreatic islets has been shown to normalize blood glucose and even improve existing complications of diabetes24. However, insulin independence 5 years after islet transplantation can only be achieved in up to 55% of the patients even when using the latest generation of immune suppression strategies25, 26. Transplantation of human islets or the entire pancreas has allowed T1DM patients to become somewhat insulin independent, which provides a proof-of-concept for beta-cell replacement therapies27, 28. However, because of the shortage of donor pancreases and islets, as well as the significant risk associated with transplantation and life-long immunosuppression, the rational differentiation of stem cells into functional beta-cells remains an attractive alternative29, 30. Nevertheless, a definitive cure for T1DM should address both the beta-cell deficit and the autoimmune response to cells that express insulin. Any beta-cell mimetic should be able to store large amounts of insulin and secrete it on demand, as in response to glucose stimulation29, 31. The most effective protocols for the in vitro generation of bonafide insulin-secreting beta-like cells that are suitable for transplantation have been the result of sophisticated trial-and-error studies elaborating timely addition of complex growth factor and small-molecule compound cocktails to human pancreatic progenitor cells32, 33, 34. The differentiation of pancreatic progenitor cells to beta-like cells is the most challenging part as current protocols provide inconsistent results and limited success in programming pancreatic progenitor cells into glucose-sensitive insulin-secreting beta-like cells35, 36, 37. One of the reasons for these observations could be the heterogeneity in endocrine differentiation and maturation towards a beta cell phenotype. Here we show that a synthetic lineage-control network programming the dynamic expression of the transcription factors Ngn3, Pdx1 and MafA enables the differentiation of hIPSC-derived pancreatic progenitor cells to glucose-sensitive insulin-secreting beta-like cells (Supplementary Fig. 1).
The differentiation pathway from pancreatic progenitor cells to glucose-sensitive insulin-secreting pancreatic beta-cells combines the transient mutually exclusive expression switches of Ngn3 (OFF-ON-OFF) and Pdx1 (ON-OFF-ON) with the concomitant induction of MafA (OFF-ON) expression10,11. Since independent control of the pancreatic transcription factors Ngn3, Pdx1 and MafA by different antibiotic transgene control systems responsive to tetracycline, erythromycin and pristinamycin did not result in the desired differential control dynamics (Supplementary Fig. 2), we have designed a vanillic acid-programmable synthetic lineage-control network that programmes hIPSC-derived pancreatic progenitor cells to specifically differentiate into glucose-sensitive insulin-secreting beta-like cells in a seamless and self-sufficient manner. The timely coordination of mutually exclusive Ngn3 and Pdx1 expression with MafA induction requires the trigger-controlled execution of a complex genetic programme that orchestrates two overlapping antagonistic band-pass filter expression profiles (OFF-ON-OFF and ON-OFF-ON), a positive band-pass filter for Ngn3 (OFF-ON-OFF) and a negative band-pass filter, also known as band-stop filter, for Pdx1 (ON-OFF-ON), the ramp-up expression phase of which is linked to a graded induction of MafA (OFF-ON).
The core of the synthetic lineage-control network consists of two transgene control devices that are sensitive to the food component and licensed food additive vanillic acid. These devices are a synthetic vanillic acid-inducible (ON-type) signalling cascade that is gradually induced by increasing the vanillic acid concentration and a vanillic acid-repressible (OFF-type) gene switch that is repressed in a vanillic acid dose-dependent manner (Fig. 1a,b). The designer cascade consists of the vanillic acid-sensitive mammalian olfactory receptor MOR9-1, which sequentially activates the G protein Sα (GSα) and adenylyl cyclase to produce a cyclic AMP (cAMP) second messenger surge38 that is rewired via the cAMP-responsive protein kinase A-mediated phospho-activation of the cAMP-response element-binding protein 1 (CREB1) to the induction of synthetic promoters (PCRE) containing CREB1-specific cAMP response elements (CRE; Fig. 1a). The co-transfection of pCI-MOR9-1 (PhCMV-MOR9-1-pASV40) and pCK53 (PCRE-SEAP-pASV40) into human mesenchymal stem cells (hMSC-TERT) confirmed the vanillic acid-adjustable secreted alkaline phosphatase (SEAP) induction of the designer cascade (>10nM vanillic acid; Fig. 1a). The vanillic acid-repressible gene switch consists of the vanillic acid-dependent transactivator (VanA1), which binds and activates vanillic acid-responsive promoters (for example, P1VanO2) at low and medium vanillic acid levels (<2μM). At high vanillic acid concentrations (>2μM), VanA1 dissociates from P1VanO2, which results in the dose-dependent repression of transgene expression39 (Fig. 1b). The co-transfection of pMG250 (PSV40-VanA1-pASV40) and pMG252 (P1VanO2-SEAP-pASV40) into hMSC-TERT corroborated the fine-tuning of the vanillic acid-repressible SEAP expression (Fig. 1b).
Figure 1: Design of a vanillic acid-responsive positive band-pass filter providing an OFF-ON-OFF expression profile.
a) Vanillic acid-inducible transgene expression. The constitutively expressed vanillic acid-sensitive olfactory G protein-coupled receptor MOR9-1 (pCI-MOR9-1; PhCMV-MOR9-1-pA) senses extracellular vanillic acid levels and triggers G protein (Gs)-mediated activation of the membrane-bound adenylyl cyclase (AC) that converts ATP into cyclic AMP (cAMP). The resulting intracellular cAMP surge activates PKA (protein kinase A), whose catalytic subunits translocate into the nucleus to phosphorylate cAMP response element-binding protein 1 (CREB1). Activated CREB1 binds to synthetic promoters (PCRE) containing cAMP-response elements (CRE) and induces PCRE-driven expression of human placental secreted alkaline phosphatase (SEAP; pCK53, PCRE-SEAP-pA). Co-transfection of pCI-MOR9-1 and pCK53 into human mesenchymal stem cells (hMSC-TERT) grown for 48h in the presence of increasing vanillic acid concentrations results in a dose-inducible SEAP expression profile. (b) Vanillic acid-repressible transgene expression. The constitutively expressed, vanillic acid-dependent transactivator VanA1(pMG250, PSV40-VanA1-pA, VanA1, VanR-VP16) binds and activates the chimeric promoter P1VanO2 (pMG252, P1VanO2-SEAP-pA) in the absence of vanillic acid. In the presence of increasing vanillic acid concentrations, VanA1 is released from P1VanO2, and transgene expression is shut down. Co-transfection of pMG250 and pMG252 into hMSC-TERT grown for 48h in the presence of increasing vanillic acid concentrations results in a dose-repressible SEAP expression profile. (c) Positive band-pass expression filter. Serial interconnection of the synthetic vanillic acid-inducible signalling cascade (a) with the vanillic acid-repressible transcription factor-based gene switch (b) by PCRE-mediated expression of VanA1 (pSP1, PCRE-VanA1-pA) results in a two-level feed-forward cascade. Owing to the opposing responsiveness and differential sensitivity to vanillic acid, this synthetic gene network programmes SEAP expression with a positive band-pass filter profile (OFF-ON-OFF) as vanillic acid levels are increased. Medium vanillic acid levels activate MOR9-1, which induces PCRE-driven VanA1 expression. VanA1remains active and triggers P1VanO2-mediated SEAP expression in feed-forward manner, which increases to maximum levels. At high vanillic acid concentrations, MOR9-1 maintains PCRE-driven VanA1 expression, but the transactivator dissociates from P1VanO2, which shuts SEAP expression down. Co-transfection of pCI-MOR9-1, pSP1 and pMG252 into hMSC-TERT grown for 48h in the presence of increasing vanillic acid concentrations programmes SEAP expression with a positive band-pass profile (OFF-ON-OFF). Data are the means±s.d. of triplicate experiments (n=9).
The opposing responsiveness and differential sensitivity of the control devices to vanillic acid are essential to programme band-pass filter expression profiles. Upon daisy-chaining the designer cascade (pCI-MOR9-1; PhCMV-MOR9-1-pASV40; pSP1, PCRE-VanA1-pASV40) and the gene switch (pSP1, PCRE-VanA1-pASV40; pMG252, P1VanO2-SEAP-pASV40) in the same cell, the network executes a band-pass filter SEAP expression profile when exposed to increasing concentrations of vanillic acid (Fig. 1c). Medium vanillic acid levels (10nM to 2μM) activate MOR9-1, which induces PCRE-driven VanA1 expression. VanA1 remains active within this concentration range and, in a feed-forward amplifier manner, triggers P1VanO2-mediated SEAP expression, which gradually increases to maximum levels (Fig. 1c). At high vanillic acid concentrations (2μM to 400μM), MOR9-1 maintains PCRE-driven VanA1 expression, but the transactivator is inactivated and dissociates from P1VanO2, which results in the gradual shutdown of SEAP expression (Fig. 1c).
For the design of the vanillic acid-programmable synthetic lineage-control network, constitutive MOR9-1 expression and PCRE-driven VanA1 expression were combined with pSP12 (pASV40-Ngn3cm←P3VanO2mFT-miR30Pdx1g-shRNA-pASV40) for endocrine specification and pSP17(PCREm-Pdx1cm-2A-MafAcm-pASV40) for maturation of developing beta-cells (Fig. 2a,b). ThepSP12-encoded expression unit enables the VanA1-controlled induction of the optimized bidirectional vanillic acid-responsive promoter (P3VanO2) that drives expression of a codon-modified Ngn3cm, the nucleic acid sequence of which is distinct from its genomic counterpart (Ngn3g) to allow for quantitative reverse transcription–PCR (qRT–PCR)-based discrimination. In the opposite direction, P3VanO2 transcribes miR30Pdx1g-shRNA, which exclusively targets genomicPdx1 (Pdx1g) transcripts for RNA interference-based destruction and is linked to the production of a blue-to-red medium fluorescent timer40 (mFT) for precise visualization of the unit’s expression dynamics in situ.pSP17 contains a dicistronic expression unit in which the modified high-tightness and lower-sensitivity PCREm promoter (see below) drives co-cistronic expression of Pdx1cm andMafAcm, which are codon-modified versions producing native transcription factors that specifically differ from their genomic counterparts (Pdx1g, MafAg) in their nucleic acid sequence. After individual validation of the vanillic acid-controlled expression and functionality of all network components (Supplementary Figs 2–9), the lineage-control network was ready to be transfected into hIPSC-derived pancreatic progenitor cells. These cells are characterized by high expression of Pdx1g and Nkx6.1 levels and the absence of Ngn3g and MafAg production32, 33, 34 (day 0:Supplementary Figs 10–16).
(a) Schematic of the synthetic lineage-control network. The constitutively expressed, vanillic acid-sensitive olfactory G protein-coupled receptor MOR9-1 (pCI-MOR9-1; PhCMV-MOR9-1-pA) senses extracellular vanillic acid levels and triggers a synthetic signalling cascade, inducing PCRE-driven expression of the transcription factor VanA1 (pSP1, PCRE-VanA1-pA). At medium vanillic acid concentrations (purple arrows), VanA1 binds and activates the bidirectional vanillic acid-responsive promoter P3VanO2 (pSP12, pA-Ngn3cm←P3VanO2mFT-miR30Pdx1g-shRNA-pA), which drives the induction of codon-modified Neurogenin 3 (Ngn3cm) as well as the coexpression of both the blue-to-red medium fluorescent timer (mFT) for precise visualization of the unit’s expression dynamics and miR30pdx1g-shRNA (a small hairpin RNA programming the exclusive destruction of genomic pancreatic and duodenal homeobox 1 (Pdx1g) transcripts). Consequently, Ngn3cm levels switch from low to high (OFF-to-ON), and Pdx1g levels toggle from high to low (ON-to-OFF). In addition, Ngn3cm triggers the transcription of Ngn3g from its genomic promoter, which initiates a positive-feedback loop. At high vanillic acid levels (orange arrows), VanA1 is inactivated, and both Ngn3cm and miR30pdx1g-shRNA are shut down. At the same time, the MOR9-1-driven signalling cascade induces the modified high-tightness and lower-sensitivity PCREm promoter that drives the co-cistronic expression of the codon-modified variants of Pdx1 (Pdx1cm) and V-maf musculoaponeurotic fibrosarcoma oncogene homologue A (MafAcm; pSP17, PCREm-Pdx1cm-2A-MafAcm-pA). Consequently, Pdx1cm and MafAcm become fully induced. As Pdx1cm expression ramps up, it initiates a positive-feedback loop by inducing the genomic counterparts Pdx1g and MafAg. Importantly, Pdx1cm levels are not affected by miR30Pdx1g-shRNA because the latter is specific for genomic Pdx1g transcripts and because the positive feedback loop-mediated amplification of Pdx1gexpression becomes active only after the shutdown of miR30Pdx1g-shRNA. Overall, the synthetic lineage-control network provides vanillic acid-programmable, transient, mutually exclusive expression switches for Ngn3 (OFF-ON-OFF) and Pdx1 (ON-OFF-ON) as well as the concomitant induction of MafA (OFF-ON) expression, which can be followed in real time (Supplementary Movies 1 and 2). (b) Schematic illustrating the individual differentiation steps from human IPSCs towards beta-like cells. The colours match the cell phenotypes reached during the individual differentiation stages programmed by the lineage-control network shown in a.
Following the co-transfection of pCI-MOR9-1 (PhCMV-MOR9-1-pASV40), pSP1 (PCRE-VanA1-pASV40), pSP12 (pASV40-Ngn3cm←P3VanO2mFT-miR30Pdx1g-shRNA-pASV40) and pSP17(PCREm-Pdx1cm-2A-MafAcm-pASV40) into hIPSC-derived pancreatic progenitor cells, the synthetic lineage-control network should override random endogenous differentiation activities and execute the pancreatic beta-cell-specific differentiation programme in a vanillic acid remote-controlled manner. To confirm that the lineage-control network operates as programmed, we cultivated network-containing and pEGFP-N1-transfected (negative-control) cells for 4 days at medium (2μM) and then 7 days at high (400μM) vanillic acid concentrations and profiled the differential expression dynamics of all of the network components and their genomic counterparts as well as the interrelated transcription factors and hormones in both whole populations and individual cells at days 0, 4, 11 and 14 (Figs 2 and 3 and Supplementary Figs 11–17).
Figure 3: Dynamics of the lineage-control network.
(a,b) Quantitative RT–PCR-based expression profiling of the pancreatic transcription factors Ngn3cm/g, Pdx1cm/g and MafAcm/g in hIPSC-derived pancreatic progenitor cells containing the synthetic lineage-control network at days 4 and 11. Data are the means±s.d. of triplicate experiments (n=9). (c–g) Immunocytochemistry of pancreatic transcription factors Ngn3cm/g, Pdx1cm/g and MafAcm/g in hIPSC-derived pancreatic progenitor cells containing the synthetic lineage-control network at days 4 and 11. hIPSC-derived pancreatic progenitor cells were co-transfected with the lineage-control vectors pCI-MOR9-1 (PhCMV-MOR9-1-pA), pSP1 (PCRE-VanA1-pA), pSP12 (pA-Ngn3cm←P3VanO2mFT-miR30Pdx1g-shRNA-pA) and pSP17 (PCREm-Pdx1cm-2A-MafAcm) and immunocytochemically stained for (c) VanA1 and Pdx1 (day 4), (d) VanA1 and Ngn3 (day 4), (e) VanA1 and Pdx1 (day 11), (f) MafA and Pdx1 (day 11) as well as (g) VanA1 and insulin (C-peptide) (day 11). The cells staining positive for VanA1 are containing the lineage-control network. DAPI, 4′,6-diamidino-2-phenylindole. Scale bar, 100μm.
…….
Multicellular organisms, including humans, consist of a highly structured assembly of a multitude of specialized cell phenotypes that originate from the same zygote and have traversed a preprogrammed multifactorial developmental plan that orchestrates sequential differentiation steps with high precision in space and time19, 51. Because of the complexity of terminally differentiated cells, the function of damaged tissues can for most medical indications only be restored via the transplantation of donor material, which is in chronically short supply52.
Despite significant progress in regenerative medicine and the availability of stem cells, the design of protocols that replicate natural differentiation programmes and provide fully functional cell mimetics remains challenging29, 53. For example, efforts to generate beta-cells from human embryonic stem cells (hESCs) have led to reliable protocols involving the sequential administration of growth factors (activin A, bone morphogenetic protein 4 (BMP-4), basic fibroblast growth factor (bFGF), FGF-10, Noggin, vascular endothelial growth factor (VEGF) and Wnt3A) and small-molecule compounds (cyclopamine, forskolin, indolactam V, IDE1, IDE2, nicotinamide, retinoic acid, SB−431542 and γ-secretase inhibitor) that modulate differentiation-specific signalling pathways31, 54, 55. In vitro differentiation of hESC-derived pancreatic progenitor cells into beta-like cells is more challenging and has been achieved recently by a complex media formulation with chemicals and growth factors32, 33, 34.
hIPSCs have become a promising alternative to hESCs; however, their use remains restricted in many countries56. Most hIPSCs used for directed differentiation studies were derived from a juvenescent cell source that is expected to show a higher degree of differentiation potential compared with older donors that typically have a higher need for medical interventions37, 57, 58. We previously succeeded in producing mRNA-reprogrammed hIPSCs from adipose tissue-derived mesenchymal stem cells of a 50-year-old donor, demonstrating that the reprogramming of cells from a donor of advanced age is possible in principle59.
Recent studies applying similar hESC-based differentiation protocols to hIPSCs have produced cells that release insulin in response to high glucose32, 33, 34. This observation suggests that functional beta-like cells can eventually be derived from hIPSCs32, 33. In our hands, the growth-factor/chemical-based technique for differentiating human IPSCs resulted in beta-like cells with poor glucose responsiveness. Recent studies have revealed significant variability in the lineage specification propensity of different hIPSC lines35, 60 and substantial differences in the expression profiles of key transcription factors in hIPSC-derived beta-like cells33. Therefore, the growth-factor/chemical-based protocols may require further optimization and need to be customized for specific hIPSC lines35. Synthetic lineage-control networks providing precise dynamic control of transcription factor expression may overcome the challenges associated with the programming of beta-like cells from different hIPSC lines.
Rather than exposing hIPSCs to a refined compound cocktail that triggers the desired differentiation in a fraction of the stem cell population, we chose to design a synthetic lineage-control network to enable single input-programmable differentiation of hIPSC-derived pancreatic progenitor cells into glucose-sensitive insulin-secreting beta-like cells. In contrast with the use of growth-factor/chemical-based cocktails, synthetic lineage-control networks are expected to (i) be more economical because of in situ production of the required transcription factors, (ii) enable simultaneous control of ectopic and chromosomally encoded transcription factor variants, (iii) tap into endogenous pathways and not be limited to cell-surface input, (iv) display improved reversibility that is not dependent on the removal of exogenous growth factors via culture media replacement, (v) provide lateral inhibition, thereby reducing the random differentiation of neighbouring cells and (vi) enable trigger-programmable and (vii) precise differential transcription factor expression switches.
The synthetic lineage-control network that precisely replicates the endogenous relative expression dynamics of the transcription factors Pdx-1, Ngn3 and MafA required the design of a new network topology that interconnects a synthetic signalling cascade and a gene switch with differential and opposing sensitivity to the food additive vanillic acid. This differentiation device provides different band-pass filter, time-delay and feed-forward amplifier topologies that interface with endogenous positive-feedback loops to orchestrate the timely expression and repression of heterologous and chromosomally encoded Ngn3, Pdx1 and MafA variants. The temporary nature of the engineering intervention, which consists of transient transfection of the genetic lineage-control components in the absence of any selection, is expected to avoid stable modification of host chromosomes and alleviate potential safety concerns. In addition, the resulting beta-cell mass could be encapsulated inside vascularized microcontainers28, a proven containment strategy in prototypic cell-based therapies currently being tested in animal models of prominent human diseases14, 15, 16, 61, 62 as well as in human clinical trials28.
The hIPSC-derived beta-like cells resulting from this trigger-induced synthetic lineage-control network exhibited glucose-stimulated insulin-release dynamics and capacity matching the human physiological range and transcriptional profiling, flow cytometric analysis and electron microscopy corroborated the lineage-controlled stem cells reached a mature beta-cell phenotype. In principle, the combination of hIPSCs derived from the adipose tissue of a 50-year-old donor59 with a synthetic lineage-control network programming glucose-sensitive insulin-secreting beta-like cells closes the design cycle of regenerative medicine63. However, hIPSCs that are derived from T1DM patients, differentiated into beta-like cells and transplanted back into the donor would still be targeted by the immune system, as demonstrated in the transplantation of segmental pancreatic grafts from identical twins64. Therefore, any beta-cell-replacement therapy will require complementary modulation of the immune system either via drugs30, 65, engineering or cell-based approaches66, 67 or packaging inside vascularizing, semi-permeable immunoprotective microcontainers28.
Capitalizing on the design principles of synthetic biology, we have successfully constructed and validated a synthetic lineage-control network that replicates the differential expression dynamics of critical transcription factors and mimicks the native differentiation pathway to programme hIPSC-derived pancreatic progenitor cells into glucose-sensitive insulin-secreting beta-like cells that compare with human pancreatic islets at a high level. The design of input-triggered synthetic lineage-control networks that execute a preprogrammed sequential differentiation agenda coordinating the timely induction and repression of multiple genes could provide a new impetus for the advancement of developmental biology and regenerative medicine.
Other related articles published in this Open Access Online Scientific Journal include the following:
Adipocyte Derived Stroma Cells: Their Usage in Regenerative Medicine and Reprogramming into Pancreatic Beta-Like Cells
Genomics and epigenetics link to DNA structure, Volume 2 (Volume Two: Latest in Genomics Methodologies for Therapeutics: Gene Editing, NGS and BioInformatics, Simulations and the Genome Ontology), Part 1: Next Generation Sequencing (NGS)
Genomics and epigenetics link to DNA structure
Larry H. Bernstein, MD, FCAP, Curator
LPBI
Sequence and Epigenetic Factors Determine Overall DNA Structure
Atomic-level simulations show electrostatic forces between each atom. [Alek Aksimentiev, University of Illinois at Urbana-Champaign]
The traditionally held hypothesis about the highly ordered organization of DNA describes the interaction of various proteins with DNA sequences to mediate the dynamic structure of the molecule. However, recent evidence has emerged that stretches of homologous DNA sequences can associate preferentially with one another, even in the absence of proteins.
Researchers at the University of Illinois Center for the Physics of Living Cells, Johns Hopkins University, and Ulsan National Institute of Science and Technology (UNIST) in South Korea found that DNA molecules interact directly with one another in ways that are dependent on the sequence of the DNA and epigenetic factors, such as methylation.
The researchers described evidence they found for sequence-dependent attractive interactions between double-stranded DNA molecules that neither involve intermolecular strand exchange nor are mediated by DNA-binding proteins.
“DNA molecules tend to repel each other in water, but in the presence of special types of cations, they can attract each other just like nuclei pulling each other by sharing electrons in between,” explained lead study author Hajin Kim, Ph.D., assistant professor of biophysics at UNIST. “Our study suggests that the attractive force strongly depends on the nucleic acid sequence and also the epigenetic modifications.”
The investigators used atomic-level supercomputer simulations to measure the forces between a pair of double-stranded DNA helices and proposed that the distribution of methyl groups on the DNA was the key to regulating this sequence-dependent attraction. To verify their findings experimentally, the scientists were able to observe a single pair of DNA molecules within nanoscale bubbles.
“Here we combine molecular dynamics simulations with single-molecule fluorescence resonance energy transfer experiments to examine the interactions between duplex DNA in the presence of spermine, a biological polycation,” the authors wrote. “We find that AT-rich DNA duplexes associate more strongly than GC-rich duplexes, regardless of the sequence homology. Methyl groups of thymine act as a steric block, relocating spermine from major grooves to interhelical regions, thereby increasing DNA–DNA attraction.”
The findings from this study were published recently in Nature Communications in an article entitled “Direct Evidence for Sequence-Dependent Attraction Between Double-Stranded DNA Controlled by Methylation.”
After conducting numerous further simulations, the research team concluded that direct DNA–DNA interactions could play a central role in how chromosomes are organized in the cell and which ones are expanded or folded up compactly, ultimately determining functions of different cell types or regulating the cell cycle.
“Biophysics is a fascinating subject that explores the fundamental principles behind a variety of biological processes and life phenomena,” Dr. Kim noted. “Our study requires cross-disciplinary efforts from physicists, biologists, chemists, and engineering scientists and we pursue the diversity of scientific disciplines within the group.”
Dr. Kim concluded by stating that “in our lab, we try to unravel the mysteries within human cells based on the principles of physics and the mechanisms of biology. In the long run, we are seeking for ways to prevent chronic illnesses and diseases associated with aging.”
Direct evidence for sequence-dependent attraction between double-stranded DNA controlled by methylation
Jejoong Yoo, Hajin Kim, Aleksei Aksimentiev, and Taekjip Ha Nature Communications 7 11045 (2016) DOI:10.1038/ncomms11045BibTex
Although proteins mediate highly ordered DNA organization in vivo, theoretical studies suggest that homologous DNA duplexes can preferentially associate with one another even in the absence of proteins. Here we combine molecular dynamics simulations with single-molecule fluorescence resonance energy transfer experiments to examine the interactions between duplex DNA in the presence of spermine, a biological polycation. We find that AT-rich DNA duplexes associate more strongly than GC-rich duplexes, regardless of the sequence homology. Methyl groups of thymine acts as a steric block, relocating spermine from major grooves to interhelical regions, thereby increasing DNA–DNA attraction. Indeed, methylation of cytosines makes attraction between GC-rich DNA as strong as that between AT-rich DNA. Recent genome-wide chromosome organization studies showed that remote contact frequencies are higher for AT-rich and methylated DNA, suggesting that direct DNA–DNA interactions that we report here may play a role in the chromosome organization and gene regulation.
Formation of a DNA double helix occurs through Watson–Crick pairing mediated by the complementary hydrogen bond patterns of the two DNA strands and base stacking. Interactions between double-stranded (ds)DNA molecules in typical experimental conditions containing mono- and divalent cations are repulsive1, but can turn attractive in the presence of high-valence cations2. Theoretical studies have identified the ion–ion correlation effect as a possible microscopic mechanism of the DNA condensation phenomena3, 4, 5. Theoretical investigations have also suggested that sequence-specific attractive forces might exist between two homologous fragments of dsDNA6, and this ‘homology recognition’ hypothesis was supported by in vitro atomic force microscopy7 and in vivo point mutation assays8. However, the systems used in these measurements were too complex to rule out other possible causes such as Watson–Crick strand exchange between partially melted DNA or protein-mediated association of DNA.
Here we present direct evidence for sequence-dependent attractive interactions between dsDNA molecules that neither involve intermolecular strand exchange nor are mediated by proteins. Further, we find that the sequence-dependent attraction is controlled not by homology—contradictory to the ‘homology recognition’ hypothesis6—but by a methylation pattern. Unlike the previous in vitro study that used monovalent (Na+) or divalent (Mg2+) cations7, we presumed that for the sequence-dependent attractive interactions to operate polyamines would have to be present. Polyamine is a biological polycation present at a millimolar concentration in most eukaryotic cells and essential for cell growth and proliferation9, 10. Polyamines are also known to condense DNA in a concentration-dependent manner2, 11. In this study, we use spermine4+(Sm4+) that contains four positively charged amine groups per molecule.
Sequence dependence of DNA–DNA forces
To characterize the molecular mechanisms of DNA–DNA attraction mediated by polyamines, we performed molecular dynamics (MD) simulations where two effectively infinite parallel dsDNA molecules, 20 base pairs (bp) each in a periodic unit cell, were restrained to maintain a prescribed inter-DNA distance; the DNA molecules were free to rotate about their axes. The two DNA molecules were submerged in 100mM aqueous solution of NaCl that also contained 20 Sm4+molecules; thus, the total charge of Sm4+, 80 e, was equal in magnitude to the total charge of DNA (2 × 2 × 20 e, two unit charges per base pair; Fig. 1a). Repeating such simulations at various inter-DNA distances and applying weighted histogram analysis12 yielded the change in the interaction free energy (ΔG) as a function of the DNA–DNA distance (Fig. 1b,c). In a broad agreement with previous experimental findings13, ΔG had a minimum, ΔGmin, at the inter-DNA distance of 25−30Å for all sequences examined, indeed showing that two duplex DNA molecules can attract each other. The free energy of inter-duplex attraction was at least an order of magnitude smaller than the Watson–Crick interaction free energy of the same length DNA duplex. A minimum of ΔG was not observed in the absence of polyamines, for example, when divalent or monovalent ions were used instead14, 15.
Figure 1: Polyamine-mediated DNA sequence recognition observed in MD simulations and smFRET experiments.
(a) Set-up of MD simulations. A pair of parallel 20-bp dsDNA duplexes is surrounded by aqueous solution (semi-transparent surface) containing 20 Sm4+ molecules (which compensates exactly the charge of DNA) and 100mM NaCl. Under periodic boundary conditions, the DNA molecules are effectively infinite. A harmonic potential (not shown) is applied to maintain the prescribed distance between the dsDNA molecules. (b,c) Interaction free energy of the two DNA helices as a function of the DNA–DNA distance for repeat-sequence DNA fragments (b) and DNA homopolymers (c). (d) Schematic of experimental design. A pair of 120-bp dsDNA labelled with a Cy3/Cy5 FRET pair was encapsulated in a ~200-nm diameter lipid vesicle; the vesicles were immobilized on a quartz slide through biotin–neutravidin binding. Sm4+ molecules added after immobilization penetrated into the porous vesicles. The fluorescence signals were measured using a total internal reflection microscope. (e) Typical fluorescence signals indicative of DNA–DNA binding. Brief jumps in the FRET signal indicate binding events. (f) The fraction of traces exhibiting binding events at different Sm4+ concentrations for AT-rich, GC-rich, AT nonhomologous and CpG-methylated DNA pairs. The sequence of the CpG-methylated DNA specifies the methylation sites (CG sequence, orange), restriction sites (BstUI, triangle) and primer region (underlined). The degree of attractive interaction for the AT nonhomologous and CpG-methylated DNA pairs was similar to that of the AT-rich pair. All measurements were done at [NaCl]=50mM and T=25°C. (g) Design of the hybrid DNA constructs: 40-bp AT-rich and 40-bp GC-rich regions were flanked by 20-bp common primers. The two labelling configurations permit distinguishing parallel from anti-parallel orientation of the DNA. (h) The fraction of traces exhibiting binding events as a function of NaCl concentration at fixed concentration of Sm4+ (1mM). The fraction is significantly higher for parallel orientation of the DNA fragments.
Unexpectedly, we found that DNA sequence has a profound impact on the strength of attractive interaction. The absolute value of ΔG at minimum relative to the value at maximum separation, |ΔGmin|, showed a clearly rank-ordered dependence on the DNA sequence: |ΔGmin| of (A)20>|ΔGmin| of (AT)10>|ΔGmin| of (GC)10>|ΔGmin| of (G)20. Two trends can be noted. First, AT-rich sequences attract each other more strongly than GC-rich sequences16. For example, |ΔGmin| of (AT)10 (1.5kcalmol−1 per turn) is about twice |ΔGmin| of (GC)10 (0.8kcalmol−1 per turn) (Fig. 1b). Second, duplexes having identical AT content but different partitioning of the nucleotides between the strands (that is, (A)20 versus (AT)10 or (G)20 versus (GC)10) exhibit statistically significant differences (~0.3kcalmol−1 per turn) in the value of |ΔGmin|.
To validate the findings of MD simulations, we performed single-molecule fluorescence resonance energy transfer (smFRET)17 experiments of vesicle-encapsulated DNA molecules. Equimolar mixture of donor- and acceptor-labelled 120-bp dsDNA molecules was encapsulated in sub-micron size, porous lipid vesicles18 so that we could observe and quantitate rare binding events between a pair of dsDNA molecules without triggering large-scale DNA condensation2. Our DNA constructs were long enough to ensure dsDNA–dsDNA binding that is stable on the timescale of an smFRET measurement, but shorter than the DNA’s persistence length (~150bp (ref. 19)) to avoid intramolecular condensation20. The vesicles were immobilized on a polymer-passivated surface, and fluorescence signals from individual vesicles containing one donor and one acceptor were selectively analysed (Fig. 1d). Binding of two dsDNA molecules brings their fluorescent labels in close proximity, increasing the FRET efficiency (Fig. 1e).
FRET signals from individual vesicles were diverse. Sporadic binding events were observed in some vesicles, while others exhibited stable binding; traces indicative of frequent conformational transitions were also observed (Supplementary Fig. 1A). Such diverse behaviours could be expected from non-specific interactions of two large biomolecules having structural degrees of freedom. No binding events were observed in the absence of Sm4+ (Supplementary Fig. 1B) or when no DNA molecules were present. To quantitatively assess the propensity of forming a bound state, we chose to use the fraction of single-molecule traces that showed any binding events within the observation time of 2min (Methods). This binding fraction for the pair of AT-rich dsDNAs (AT1, 100% AT in the middle 80-bp section of the 120-bp construct) reached a maximum at ~2mM Sm4+(Fig. 1f), which is consistent with the results of previous experimental studies2, 3. In accordance with the prediction of our MD simulations, GC-rich dsDNAs (GC1, 75% GC in the middle 80bp) showed much lower binding fraction at all Sm4+ concentrations (Fig. 1b,c). Regardless of the DNA sequence, the binding fraction reduced back to zero at high Sm4+ concentrations, likely due to the resolubilization of now positively charged DNA–Sm4+ complexes2, 3, 13.
Because the donor and acceptor fluorophores were attached to the same sequence of DNA, it remained possible that the sequence homology between the donor-labelled DNA and the acceptor-labelled DNA was necessary for their interaction6. To test this possibility, we designed another AT-rich DNA construct AT2 by scrambling the central 80-bp section of AT1 to remove the sequence homology (Supplementary Table 1). The fraction of binding traces for this nonhomologous pair of donor-labelled AT1 and acceptor-labelled AT2 was comparable to that for the homologous AT-rich pair (donor-labelled AT1 and acceptor-labelled AT1) at all Sm4+ concentrations tested (Fig. 1f). Furthermore, this data set rules out the possibility that the higher binding fraction observed experimentally for the AT-rich constructs was caused by inter-duplex Watson–Crick base pairing of the partially melted constructs.
Next, we designed a DNA construct named ATGC, containing, in its middle section, a 40-bp AT-rich segment followed by a 40-bp GC-rich segment (Fig. 1g). By attaching the acceptor to the end of either the AT-rich or GC-rich segments, we could compare the likelihood of observing the parallel binding mode that brings the two AT-rich segments together and the anti-parallel binding mode. Measurements at 1mM Sm4+ and 25 or 50mM NaCl indicated a preference for the parallel binding mode by ~30% (Fig. 1h). Therefore, AT content can modulate DNA–DNA interactions even in a complex sequence context. Note that increasing the concentration of NaCl while keeping the concentration of Sm4+ constant enhances competition between Na+ and Sm4+ counterions, which reduces the concentration of Sm4+ near DNA and hence the frequency of dsDNA–dsDNA binding events (Supplementary Fig. 2).
Methylation determines the strength of DNA–DNA attraction
Analysis of the MD simulations revealed the molecular mechanism of the polyamine-mediated sequence-dependent attraction (Fig. 2). In the case of the AT-rich fragments, the bulky methyl group of thymine base blocks Sm4+ binding to the N7 nitrogen atom of adenine, which is the cation-binding hotspot21, 22. As a result, Sm4+ is not found in the major grooves of the AT-rich duplexes and resides mostly near the DNA backbone (Fig. 2a,d). Such relocated Sm4+ molecules bridge the two DNA duplexes better, accounting for the stronger attraction16, 23, 24, 25. In contrast, significant amount of Sm4+ is adsorbed to the major groove of the GC-rich helices that lacks cation-blocking methyl group (Fig. 2b,e).
Figure 2: Molecular mechanism of polyamine-mediated DNA sequence recognition.
(a–c) Representative configurations of Sm4+ molecules at the DNA–DNA distance of 28Å for the (AT)10–(AT)10 (a), (GC)10–(GC)10 (b) and (GmC)10–(GmC)10 (c) DNA pairs. The backbone and bases of DNA are shown as ribbon and molecular bond, respectively; Sm4+ molecules are shown as molecular bonds. Spheres indicate the location of the N7 atoms and the methyl groups. (d–f) The average distributions of cations for the three sequence pairs featured in a–c. Top: density of Sm4+ nitrogen atoms (d=28Å) averaged over the corresponding MD trajectory and the z axis. White circles (20Å in diameter) indicate the location of the DNA helices. Bottom: the average density of Sm4+ nitrogen (blue), DNA phosphate (black) and sodium (red) atoms projected onto the DNA–DNA distance axis (x axis). The plot was obtained by averaging the corresponding heat map data over y=[−10, 10] Å. See Supplementary Figs 4 and 5 for the cation distributions at d=30, 32, 34 and 36Å.
If indeed the extra methyl group in thymine, which is not found in cytosine, is responsible for stronger DNA–DNA interactions, we can predict that cytosine methylation, which occurs naturally in many eukaryotic organisms and is an essential epigenetic regulation mechanism26, would also increase the strength of DNA–DNA attraction. MD simulations showed that the GC-rich helices containing methylated cytosines (mC) lose the adsorbed Sm4+ (Fig. 2c,f) and that |ΔGmin| of (GC)10 increases on methylation of cytosines to become similar to |ΔGmin| of (AT)10 (Fig. 1b).
To experimentally assess the effect of cytosine methylation, we designed another GC-rich construct GC2 that had the same GC content as GC1 but a higher density of CpG sites (Supplementary Table 1). The CpG sites were then fully methylated using M. SssI methyltransferase (Supplementary Fig. 3; Methods). As predicted from the MD simulations, methylation of the GC-rich constructs increased the binding fraction to the level of the AT-rich constructs (Fig. 1f).
The sequence dependence of |ΔGmin| and its relation to the Sm4+ adsorption patterns can be rationalized by examining the number of Sm4+ molecules shared by the dsDNA molecules (Fig. 3a). An Sm4+ cation adsorbed to the major groove of one dsDNA is separated from the other dsDNA by at least 10Å, contributing much less to the effective DNA–DNA attractive force than a cation positioned between the helices, that is, the ‘bridging’ Sm4+ (ref. 23). An adsorbed Sm4+ also repels other Sm4+ molecules due to like-charge repulsion, lowering the concentration of bridging Sm4+. To demonstrate that the concentration of bridging Sm4+ controls the strength of DNA–DNA attraction, we computed the number of bridging Sm4+ molecules, Nspm (Fig. 3b). Indeed, the number of bridging Sm4+ molecules ranks in the same order as |ΔGmin|: Nspm of (A)20>Nspm of (AT)10≈Nspm of (GmC)10>Nspm of (GC)10>Nspm of (G)20. Thus, the number density of nucleotides carrying a methyl group (T and mC) is the primary determinant of the strength of attractive interaction between two dsDNA molecules. At the same time, the spatial arrangement of the methyl group carrying nucleotides can affect the interaction strength as well (Fig. 3c). The number of methyl groups and their distribution in the (AT)10 and (GmC)10 duplex DNA are identical, and so are their interaction free energies, |ΔGmin| of (AT)10≈|ΔGmin| of (GmC)10. For AT-rich DNA sequences, clustering of the methyl groups repels Sm4+ from the major groove more efficiently than when the same number of methyl groups is distributed along the DNA (Fig. 3b). Hence, |ΔGmin| of (A)20>|ΔGmin| of (AT)10. For GC-rich DNA sequences, clustering of the cation-binding sites (N7 nitrogen) attracts more Sm4+ than when such sites are distributed along the DNA (Fig. 3b), hence |ΔGmin| is larger for (GC)10 than for (G)20.
Figure 3: Methylation modulates the interaction free energy of two dsDNA molecules by altering the number of bridging Sm4+.
(a) Typical spatial arrangement of Sm4+ molecules around a pair of DNA helices. The phosphates groups of DNA and the amine groups of Sm4+ are shown as red and blue spheres, respectively. ‘Bridging’ Sm4+molecules reside between the DNA helices. Orange rectangles illustrate the volume used for counting the number of bridging Sm4+ molecules. (b) The number of bridging amine groups as a function of the inter-DNA distance. The total number of Sm4+ nitrogen atoms was computed by averaging over the corresponding MD trajectory and the 10Å (x axis) by 20Å (y axis) rectangle prism volume (a) centred between the DNA molecules. (c) Schematic representation of the dependence of the interaction free energy of two DNA molecules on their nucleotide sequence. The number and spatial arrangement of nucleotides carrying a methyl group (T or mC) determine the interaction free energy of two dsDNA molecules.
Genome-wide investigations of chromosome conformations using the Hi–C technique revealed that AT-rich loci form tight clusters in human nucleus27, 28. Gene or chromosome inactivation is often accompanied by increased methylation of DNA29 and compaction of facultative heterochromatin regions30. The consistency between those phenomena and our findings suggest the possibility that the polyamine-mediated sequence-dependent DNA–DNA interaction might play a role in chromosome folding and epigenetic regulation of gene expression.
Rau, D. C., Lee, B. & Parsegian, V. A.Measurement of the repulsive force between polyelectrolyte molecules in ionic solution: hydration forces between parallel DNA double helices. Proc. Natl Acad. Sci. USA81, 2621–2625 (1984).
Raspaud, E., Olvera de la Cruz, M., Sikorav, J. L. & Livolant, F.Precipitation of DNA by polyamines: a polyelectrolyte behavior. Biophys. J.74, 381–393 (1998).
Grosberg, A. Y., Nguyen, T. T. & Shklovskii, B. I.The physics of charge inversion in chemical and biological systems. Rev. Mod. Phys.74, 329–345 (2002).
Thomas, T. & Thomas, T. J.Polyamines in cell growth and cell death: molecular mechanisms and therapeutic applications. Cell. Mol. Life Sci.58, 244–258 (2001).
This is a lovely method and should find wide applicability in many settings, especially for microorganisms and cell lines. However, it is not clear that this approach will be, as implied by the discussion, an efficient mapping method for all multicellular organisms. I have performed similar experiments in Drosophila, focused on meiotic recombination, on a much smaller scale, and found that CRISPR-Cas9 can indeed generate targeted recombination at gRNA target sites. In every case I tested, I found that the recombination event was associated with a deletion at the gRNA site, which is probably unimportant for most mapping efforts, but may be a concern in some specific cases, for example for clinical applications. It would be interesting to know how often mutations occurred at the targeted gRNA site in this study.
The wider issue, however, is whether CRISPR-mediated recombination will be more efficient than other methods of mapping. After careful consideration of all the costs and the time involved in each of the steps for Drosophila, we have decided that targeted meiotic recombination using flanking visible markers will be, in most cases, considerably more efficient than CRISPR-mediated recombination. This is mainly due to the large expense of injecting embryos and the extensive effort and time required to screen injected animals for appropriate events. It is both cheaper and faster to generate markers (with CRISPR) and then perform a large meiotic recombination mapping experiment than it would be to generate the lines required for CRISPR-mediated recombination mapping. It is possible to dramatically reduce costs by, for example, mapping sequentially at finer resolution. But this approach would require much more time than marker-assisted mapping. If someone develops a rapid and cheap method of reliably introducing DNA into Drosophila embryos, then this calculus might change.
However, it is possible to imagine situations where CRISPR-mediated mapping would be preferable, even for Drosophila. For example, some genomic regions display extremely low or highly non-uniform recombination rates. It is possible that CRISPR-mediated mapping could provide a reasonable approach to fine mapping genes in these regions.
The authors also propose the exciting possibility that CRISPR-mediated loss of heterozygosity could be used to map traits in sterile species hybrids. It is not entirely obvious to me how this experiment would proceed and I hope the authors can illuminate me. If we imagine driving a recombination event in the early embryo (with maternal Cas9 from one parent and gRNA from a second parent), then at best we would end up with chimeric individuals carrying mitotic clones. I don’t think one could generate diploid animals where all cells carried the same loss of heterozygosity event. Even if we could, this experiment would require construction of a substantial number of stable transgenic lines expressing gRNAs. Mapping an ~20Mbp chromosome arm to ~10kb would require on the order of two-thousand transgenic lines. Not an undertaking to be taken lightly. It is already possible to perform similar tests (hemizygosity tests) using D. melanogaster deficiency lines in crosses with D. simulans, so perhaps CRISPR-mediated LOH could complement these deficiency screens for fine mapping efforts. But, at the moment, it is not clear to me how to do the experiment.