gene sequencing | Leaders in Pharmaceutical Business Intelligence Group, LLC, Doing Business As LPBI Group, Newton, MA

Posts Tagged ‘gene sequencing’

The Use of ChatGPT in the World of BioInformatics and Cancer Research and Development of BioGPT by MIT

Posted in Advanced Computing Platform, Artificial Intelligence - Breakthroughs in Theories and Technologies, Artificial Intelligence - General, Artificial Intelligence Applications in Health Care, Artificial Intelligence in CANCER, Artificial Intelligence in Health Care - Tools & Innovations, BioIT: BioInformatics, BioIT: BioInformatics, NGS, Clinical & Translational, Pharmaceutical R&D Informatics, Clinical Genomics, Cancer Informatics, Biological Networks, Biological Networks, Gene Regulation and Evolution, ChatGPT in Academic Education, ChatGPT, GPT-4, Deep Learning, Intelligent Information Systems, Machine Learning, Natural Language Processing (NLP), Next Generation Sequencing (NGS), tagged Artificial intelligence, Artificial Intelligence (AI), artificial intelligence in drug design, BioGPT, bioinformatic tools, BioPERL, ChatGPT, gene sequencing, microsoft, MIT, Natural Language Processing (NLP), systems biology on June 12, 2023| 3 Comments »

The Use of ChatGPT in the World of BioInformatics and Cancer Research and Development of BioGPT by MIT

Curator: Stephen J. Williams, Ph.D.

Chatbots are being used in multiple interdisciplanary areas in research and medicine and so it was a natural progression to incorporate artificial intelligence (AI), natural language processing (NLP), and chatbot technology like ChatGPT into bioinformatic analysis.

“In domains like informatics, management and marketing, media and communication science, languages and philosophy, psychology and sociology, engineering, design, and human-computer interaction, the fast expanding body of chatbot study is clearly interdisciplinary.

The field of bioinformatics is a natural fit for incorporting this technology. The curated information below shows some use examples of this technology in bioinformatics related to cancer research.

The text in this reference source was actually generated by ChatGPT (https://chat.openai.com/chat)

Source: https://omicstutorials.com/using-chatgpt-in-bioinformatics-and-biomedical-research/

How ChatGPT can be used in bioinformatics research?
ChatGPT is a variant of the GPT (Generative Pre-trained Transformer) language model that was designed to generate human-like text in a conversational setting. It is not directly related to bioinformatics, which is the field of study that deals with the storage, retrieval, and analysis of biological data, particularly in the context of genetics and genomics. However, ChatGPT could potentially be used in a bioinformatics context as a tool for generating natural language descriptions of genetic or genomic data, or for generating responses to queries about such data.

For example, a researcher could use ChatGPT to generate descriptions of genetic variants or gene expression patterns based on input data. This could be useful for generating summaries of results for reports or papers, or for generating explanations of complex concepts for non-expert audiences. ChatGPT could also be used to generate responses to queries about genetic or genomic data, such as “What is the function of gene X?” or “What are the potential consequences of variant Y?”

It’s worth noting that ChatGPT is just one of many tools and techniques that can be used in bioinformatics, and it is unlikely to be the most suitable or effective option in every situation. There are many specialized tools and resources available for working with biological data, and it is often necessary to use a combination of these tools to fully analyze and understand the data.

Generating descriptions of genetic or genomic data: ChatGPT could be used to generate natural language descriptions of genetic or genomic data based on input data. For example, suppose a researcher has a dataset containing information about gene expression levels in different tissues. The researcher could use ChatGPT to generate a description of the data, such as:
“Gene X is highly expressed in the liver and kidney, with moderate expression in the brain and heart. Gene Y, on the other hand, shows low expression in all tissues except for the lung, where it is highly expressed.”

Thereby ChatGPT, at its simplest level, could be used to ask general questions like “What is the function of gene product X?” and a ChatGPT could give a reasonable response without the scientist having to browse through even highly curated databases lie GeneCards or UniProt or GenBank. Or even “What are potential interactors of Gene X, validated by yeast two hybrid?” without even going to the curated InterActome databases or using expensive software like Genie.

Summarizing results: ChatGPT could be used to generate summaries of results from genetic or genomic studies. For example, a researcher might use ChatGPT to generate a summary of a study that found a association between a particular genetic variant and a particular disease. The summary might look something like this:
“Our study found that individuals with the variant form of gene X are more likely to develop disease Y. Further analysis revealed that this variant is associated with changes in gene expression that may contribute to the development of the disease.”

It’s worth noting that ChatGPT is just one tool that could potentially be used in these types of applications, and it is likely to be most effective when used in combination with other bioinformatics tools and resources. For example, a researcher might use ChatGPT to generate a summary of results, but would also need to use other tools to analyze the data and confirm the findings.

ChatGPT is a variant of the GPT (Generative Pre-training Transformer) language model that is designed for open-domain conversation. It is not specifically designed for generating descriptions of genetic variants or gene expression patterns, but it can potentially be used for this purpose if you provide it with a sufficient amount of relevant training data and fine-tune it appropriately.

To use ChatGPT to generate descriptions of genetic variants or gene expression patterns, you would first need to obtain a large dataset of examples of descriptions of genetic variants or gene expression patterns. You could use this dataset to fine-tune the ChatGPT model on the task of generating descriptions of genetic variants or gene expression patterns.

Here’s an example of how you might use ChatGPT to generate a description of a genetic variant:

First, you would need to pre-process your dataset of descriptions of genetic variants to prepare it for use with ChatGPT. This might involve splitting the descriptions into individual sentences or phrases, and encoding them using a suitable natural language processing (NLP) library or tool.

Next, you would need to fine-tune the ChatGPT model on the task of generating descriptions of genetic variants. This could involve using a tool like Hugging Face’s Transformers library to load the ChatGPT model and your pre-processed dataset, and then training the model on the task of generating descriptions of genetic variants using an appropriate optimization algorithm.

Once the model has been fine-tuned, you can use it to generate descriptions of genetic variants by providing it with a prompt or seed text and asking it to generate a response. For example, you might provide the model with the prompt “Generate a description of a genetic variant associated with increased risk of breast cancer,” and ask it to generate a response. The model should then generate a description of a genetic variant that is associated with increased risk of breast cancer.

It’s worth noting that generating high-quality descriptions of genetic variants or gene expression patterns is a challenging task, and it may be difficult to achieve good results using a language model like ChatGPT without a large amount of relevant training data and careful fine-tuning.

To train a language model like chatGPT to extract information about specific genes or diseases from research papers, you would need to follow these steps:

Gather a large dataset of research papers that contain information about the specific genes or diseases you are interested in. This dataset should be diverse and representative of the types of papers you want the model to be able to extract information from.

Preprocess the text data in the research papers by tokenizing the text and creating a vocabulary. You may also want to consider lemmatizing or stemming the text to reduce the dimensionality of the dataset.

Train the language model on the preprocessed text data. You may want to fine-tune a pre-trained model such as chatGPT on your specific dataset, or you can train a new model from scratch.

ChatGPT could also be useful for sequence analysis

A few examples of sequence analysis a ChatGPT could be useful include:

Protein structure
Identifying functional regions of a protein
Predicting protein-protein interactions
Identifying protein homologs
Generating Protein alignments

All this could be done without having access to UNIX servers or proprietary software or knowing GCG coding

ChatGPT in biomedical research
There are several potential ways that ChatGPT or other natural language processing (NLP) models could be applied in biomedical research:

Text summarization: ChatGPT or other NLP models could be used to summarize large amounts of text, such as research papers or clinical notes, in order to extract key information and insights more quickly.

Data extraction: ChatGPT or other NLP models could be used to extract structured data from unstructured text sources, such as research papers or clinical notes. For example, the model could be trained to extract information about specific genes or diseases from research papers, and then used to create a database of this information for further analysis.

Literature review: ChatGPT or other NLP models could be used to assist with literature review tasks, such as identifying relevant papers, extracting key information from papers, or summarizing the main findings of a group of papers.

Predictive modeling: ChatGPT or other NLP models could be used to build predictive models based on large amounts of text data, such as electronic health records or research papers. For example, the model could be trained to predict the likelihood of a patient developing a particular disease based on their medical history and other factors.

It’s worth noting that while NLP models like ChatGPT have the potential to be useful tools in biomedical research, they are only as good as the data they are trained on, and it is important to carefully evaluate the quality and reliability of any results generated by these models.

ChatGPT in text mining of biomedical data
ChatGPT could potentially be used for text mining in the biomedical field in a number of ways. Here are a few examples:

Extracting information from scientific papers: ChatGPT could be trained on a large dataset of scientific papers in the biomedical field, and then used to extract specific pieces of information from these papers, such as the names of compounds, their structures, and their potential uses.

Generating summaries of scientific papers: ChatGPT could be used to generate concise summaries of scientific papers in the biomedical field, highlighting the main findings and implications of the research.

Identifying trends and patterns in scientific literature: ChatGPT could be used to analyze large datasets of scientific papers in the biomedical field and identify trends and patterns in the data, such as emerging areas of research or common themes among different papers.

Generating questions for further research: ChatGPT could be used to suggest questions for further research in the biomedical field based on existing scientific literature, by identifying gaps in current knowledge or areas where further investigation is needed.

Generating hypotheses for scientific experiments: ChatGPT could be used to generate hypotheses for scientific experiments in the biomedical field based on existing scientific literature and data, by identifying potential relationships or associations that could be tested in future research.

PLEASE WATCH VIDEO

In this video, a bioinformatician describes the ways he uses ChatGPT to increase his productivity in writing bioinformatic code and conducting bioinformatic analyses.

He describes a series of uses of ChatGPT in his day to day work as a bioinformatian:

Using ChatGPT as a search engine: He finds more useful and relevant search results than a standard Google or Yahoo search. This saves time as one does not have to pour through multiple pages to find information. However, a caveat is ChatGPT does NOT return sources, as highlighted in previous postings on this page. This feature of ChatGPT is probably why Microsoft bought OpenAI in order to incorporate ChatGPT in their Bing search engine, as well as Office Suite programs

ChatGPT to help with coding projects: Bioinformaticians will spend multiple hours searching for and altering open access available code in order to run certain function like determining the G/C content of DNA (although there are many UNIX based code that has already been established for these purposes). One can use ChatGPT to find such a code and then assist in debugging that code for any flaws

ChatGPT to document and add coding comments: When writing code it is useful to add comments periodically to assist other users to determine how the code works and also how the program flow works as well, including returned variables.

One of the comments was interesting and directed one to use BIOGPT instead of ChatGPT

@tzvi7989

1 month ago (edited)

0:54 oh dear. You cannot use chatgpt like that in Bioinformatics as it is rn without double checking the info from it. You should be using biogpt instead for paper summarisation. ChatGPT goes for human-like responses over precise information recal. It is quite good for debugging though and automating boring awkward scripts

So what is BIOGPT?

BioGPT https://github.com/microsoft/BioGPT

The BioGPT model was proposed in BioGPT: generative pre-trained transformer for biomedical text generation and mining by Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon and Tie-Yan Liu. BioGPT is a domain-specific generative pre-trained Transformer language model for biomedical text generation and mining. BioGPT follows the Transformer language model backbone, and is pre-trained on 15M PubMed abstracts from scratch.

The abstract from the paper is the following:

Pre-trained language models have attracted increasing attention in the biomedical domain, inspired by their great success in the general natural language domain. Among the two main branches of pre-trained language models in the general language domain, i.e. BERT (and its variants) and GPT (and its variants), the first one has been extensively studied in the biomedical domain, such as BioBERT and PubMedBERT. While they have achieved great success on a variety of discriminative downstream biomedical tasks, the lack of generation ability constrains their application scope. In this paper, we propose BioGPT, a domain-specific generative Transformer language model pre-trained on large-scale biomedical literature. We evaluate BioGPT on six biomedical natural language processing tasks and demonstrate that our model outperforms previous models on most tasks. Especially, we get 44.98%, 38.42% and 40.76% F1 score on BC5CDR, KD-DTI and DDI end-to-end relation extraction tasks, respectively, and 78.2% accuracy on PubMedQA, creating a new record. Our case study on text generation further demonstrates the advantage of BioGPT on biomedical literature to generate fluent descriptions for biomedical terms.

Tips:

BioGPT is a model with absolute position embeddings so it’s usually advised to pad the inputs on the right rather than the left.
BioGPT was trained with a causal language modeling (CLM) objective and is therefore powerful at predicting the next token in a sequence. Leveraging this feature allows BioGPT to generate syntactically coherent text as it can be observed in the run_generation.py example script.
The model can take the past_key_values (for PyTorch) as input, which is the previously computed key/value attention pairs. Using this (past_key_values or past) value prevents the model from re-computing pre-computed values in the context of text generation. For PyTorch, see past_key_values argument of the BioGptForCausalLM.forward() method for more information on its usage.

This model was contributed by kamalkraj. The original code can be found here.

This repository contains the implementation of BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining, by Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon and Tie-Yan Liu. BioGPT is a github which is being developed by MIT in collaboration with Microsoft. It is based on Python.

License

BioGPT is MIT-licensed. The license applies to the pre-trained models as well.

Contributing

This project welcomes contributions and suggestions. Most contributions require you to agree to a Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us the rights to use your contribution. For details, visit https://cla.opensource.microsoft.com.

When you submit a pull request, a CLA bot will automatically determine whether you need to provide a CLA and decorate the PR appropriately (e.g., status check, comment). Simply follow the instructions provided by the bot. You will only need to do this once across all repos using our CLA.

This project has adopted the Microsoft Open Source Code of Conduct. For more information see the Code of Conduct FAQ or contact opencode@microsoft.com with any additional questions or comments.

As of right now this does not seem Open Access, however a sign up is required!

We provide our pre-trained BioGPT model checkpoints along with fine-tuned checkpoints for downstream tasks, available both through URL download as well as through the Hugging Face 🤗 Hub.

Model	Description	URL	🤗 Hub
BioGPT	Pre-trained BioGPT model checkpoint	link	link
BioGPT-Large	Pre-trained BioGPT-Large model checkpoint	link	link
BioGPT-QA-PubMedQA-BioGPT	Fine-tuned BioGPT for question answering task on PubMedQA	link
BioGPT-QA-PubMedQA-BioGPT-Large	Fine-tuned BioGPT-Large for question answering task on PubMedQA	link
BioGPT-RE-BC5CDR	Fine-tuned BioGPT for relation extraction task on BC5CDR	link
BioGPT-RE-DDI	Fine-tuned BioGPT for relation extraction task on DDI	link
BioGPT-RE-DTI	Fine-tuned BioGPT for relation extraction task on KD-DTI	link
BioGPT-DC-HoC	Fine-tuned BioGPT for document classification task on HoC	link

So far as to date (6/12/2023) there are three papers referencing BioGPT on PubMed (abstracts and links below)

BioGPT: generative pre-trained transformer for biomedical text generation and mining.

Luo R, Sun L, Xia Y, Qin T, Zhang S, Poon H, Liu TY.Brief Bioinform. 2022 Nov 19;23(6):bbac409. doi: 10.1093/bib/bbac409.PMID: 36156661

We evaluate BioGPT on six biomedical natural language processing tasks and demonstrate that our model outperforms previous models on most tasks. …Our case study on text generation further demonstrates the advantage of BioGPT on biomedical literature to generate fl …

Here they show the advantages of BioGPT over BioBERT and ChatGPT in an example of text extraction and question answer

Table 7. Example of generated text

Input Model Text

Bicalutamide GPT-2 Bicalutamide, a synthetic bicalutamide derivative, was found to increase the cellular proliferation

of C. elegans cells. This effect was not observed with the other synthetic bicalutamide derivatives. This is in line with our previous work which showed that bicalutamide increases the proliferation of C. elegans cells but not in the presence of synthetic bicalutamide.

BioGPT Bicalutamide (Casodex) is an androgen receptor (AR) antagonist approved for the treatment of metastatic castration-resistant prostate cancer (mCRPC) in patients who have progressed on or are ineligible for docetaxel chemotherapy, as well as for the treatment of early-stage prostate cancer in men who have not progressed on or are ineligible for docetaxel chemotherapy, as well as for the treatment of metastatic castration-sensitive prostate cancer (mCSPC) in men who have not progressed on or are ineligible for docetaxel chemotherapy.

Janus kinase 3 (JAK-3) GPT-2 Janus kinase 3 (JAK-3) mediates the conversion of glucose to glycogen in skeletal muscle. The
increase in glucose uptake by skeletal muscle is believed to be associated with a reduction in muscle protein breakdown and with the release of ketone bodies, which contribute to the metabolism of glucose.

BioGPT Janus kinase 3 (JAK-3) is a member of the Janus kinase (JAK) family of non-receptor tyrosine kinases and plays an important role in the regulation of cell proliferation, differentiation, survival, migration and angiogenesis.

Apricitabine GPT-2 Apricitabine, duloxetine and citalopram) were investigated in patients with a history of
treatment-resistant depression (n = 10). We found that treatment with the three drugs significantly decreased depressive symptoms, but not depressive symptoms over the next 12 weeks. We conclude that antidepressant drugs that reduce depressive symptoms and reduce depressive symptoms over time may be beneficial in patients with recurrent depression.

BioGPT Apricitabine is an oral prodrug of 5-aza-2 ’-deoxycytidine (5-aza-CdR), a DNA methyltransferase (DNMT) inhibitor, which has been approved by the US Food and Drug Administration (FDA) for the treatment of myelodysplastic syndrome (MDS) and acute myeloid leukemia (AML) in combination with low-dose cytarabine (Ara-C) and granulocyte colony-stimulating factor (G-CSF) for patients with intermediate-2 or high-risk MDS or AML.

Xylazine GPT-2 Xylazine, the ‘active ingredient’ in ‘bath salts’ (see above) has been linked to numerous deaths,
including those of former University of California, Berkeley, professor Andrew Anglin, who posted

a video of himself having a bath filled with it.

BioGPT Xylazine is an alpha 2-adrenoceptor agonist which has been used as a sedative and analgesic in veterinary medicine for many years, but its effects on the cardiovascular system have not been extensively studied in the dog, and its effects on the central nervous system (CNS) have not been well characterized in the dog, despite the fact that xylazine has been widely used as a sedative and analgesic in veterinary medicine for more than 30 years.

Psoralen GPT-2 Psoralen, the city of the spiritualist, and had been given to him by God, who wished to show him
the way of salvation, and to draw him back to the true life of Christ, and to bring him into communion with his fellow-man.

BioGPT Psoralen photochemotherapy (PUVA) is a well-established treatment for psoriasis, but its use is limited by the risk of skin cancer, particularly squamous cell carcinoma (SCC) of the head and neck (H & N), which is the most common site of skin cancer in the United States (US)

CP-673451 GPT-2 CP-673451) and 4-5-6-3-1-2-3-4-6-7-8-9-10-11-12-13-14-15-16-17-18-19-20-21-22-23-24-25-26-27-

28-29-30-31-32-33-34-35-36-37-38-39-40-41-42-43-44-45-46-47-48-49-50…

BioGPT CP-673451 is a potent, selective, and orally active inhibitor of human neutrophil elastase (HNE) and human cathepsin G (CatG) with in vitro and in vivo anti-inflammatory activity in a variety of animal models of inflammation and in a model of acute lung injury (ALI) in the rat induced by intratracheal instillation of lipopolysaccharide (LPS) and tumor necrosis factor-alpha (TNF-alpha), a model of acute lung injury (ALI) in which neutrophils play an important role.

BIIB-021 GPT-2 BIIB-021A-2-1.pdf. Wright, C. (2002). ‘The War on Terror: A Legal Theory of Terrorist Attacks,’ in
L.E.R. Wright, S.R. Warrick and J.H. Rothman (Eds.), The War on Terror: Essays in Honor of Michael T. Klare (New York: Oxford University Press), 143-57.

BioGPT BIIB-021 is a novel, orally active, non-peptide bradykinin B2 receptor antagonist with potent and long-lasting anti-inflammatory activity in animal models of acute and chronic inflammation and in a rat model of adjuvant-induced arthritis (AIA), an animal model of rheumatoid arthritis (RA) and in a rat model of collagen-induced arthritis (CIA), an animal model of collagen-induced arthritis (CIA), in which arthritis is induced by immunization with bovine type II collagen (CII).

Note how BioGPT is more descriptive and accurate!

EGFI: drug-drug interaction extraction and generation with fusion of enriched entity and sentence information.

Huang L, Lin J, Li X, Song L, Zheng Z, Wong KC.Brief Bioinform. 2022 Jan 17;23(1):bbab451. doi: 10.1093/bib/bbab451.PMID: 34791012

The rapid growth in literature accumulates diverse and yet comprehensive biomedical knowledge hidden to be mined such as drug interactions. However, it is difficult to extract the heterogeneous knowledge to retrieve or even discover the latest and novel knowledge in an efficient manner. To address such a problem, we propose EGFI for extracting and consolidating drug interactions from large-scale medical literature text data. Specifically, EGFI consists of two parts: classification and generation. In the classification part, EGFI encompasses the language model BioBERT which has been comprehensively pretrained on biomedical corpus. In particular, we propose the multihead self-attention mechanism and packed BiGRU to fuse multiple semantic information for rigorous context modeling. In the generation part, EGFI utilizes another pretrained language model BioGPT-2 where the generation sentences are selected based on filtering rules.

Results: We evaluated the classification part on ‘DDIs 2013’ dataset and ‘DTIs’ dataset, achieving the F1 scores of 0.842 and 0.720 respectively. Moreover, we applied the classification part to distinguish high-quality generated sentences and verified with the existing growth truth to confirm the filtered sentences. The generated sentences that are not recorded in DrugBank and DDIs 2013 dataset demonstrated the potential of EGFI to identify novel drug relationships.

Availability: Source code are publicly available at https://github.com/Layne-Huang/EGFI.

GeneGPT: Augmenting Large Language Models with Domain Tools for Improved Access to Biomedical Information.

Jin Q, Yang Y, Chen Q, Lu Z.ArXiv. 2023 May 16:arXiv:2304.09667v3. Preprint.PMID: 37131884 Free PMC article.

While large language models (LLMs) have been successfully applied to various tasks, they still face challenges with hallucinations. Augmenting LLMs with domain-specific tools such as database utilities can facilitate easier and more precise access to specialized knowledge. In this paper, we present GeneGPT, a novel method for teaching LLMs to use the Web APIs of the National Center for Biotechnology Information (NCBI) for answering genomics questions. Specifically, we prompt Codex to solve the GeneTuring tests with NCBI Web APIs by in-context learning and an augmented decoding algorithm that can detect and execute API calls. Experimental results show that GeneGPT achieves state-of-the-art performance on eight tasks in the GeneTuring benchmark with an average score of 0.83, largely surpassing retrieval-augmented LLMs such as the new Bing (0.44), biomedical LLMs such as BioMedLM (0.08) and BioGPT (0.04), as well as GPT-3 (0.16) and ChatGPT (0.12). Our further analyses suggest that: (1) API demonstrations have good cross-task generalizability and are more useful than documentations for in-context learning; (2) GeneGPT can generalize to longer chains of API calls and answer multi-hop questions in GeneHop, a novel dataset introduced in this work; (3) Different types of errors are enriched in different tasks, providing valuable insights for future improvements.

PLEASE WATCH THE FOLLOWING VIDEOS ON BIOGPT

This one entitled

Microsoft’s BioGPT Shows Promise as the Best Biomedical NLP

gives a good general description of this new MIT/Microsoft project and its usefullness in scanning 15 million articles on PubMed while returning ChatGPT like answers.

Please note one of the comments which is VERY IMPORTANT

@rufus9322

2 months ago

bioGPT is difficult for non-developers to use, and Microsoft researchers seem to default that all users are proficient in Python and ML.

Much like Microsoft Azure it seems this BioGPT is meant for developers who have advanced programming skill. Seems odd then to be paying programmers multiK salaries when one or two Key Opinion Leaders from the medical field might suffice but I would be sure Microsoft will figure this out.

ALSO VIEW VIDEO

This is a talk from Microsoft on BioGPT

The world’s most innovative intersection

Posted in BioIT: BioInformatics, NGS, Clinical & Translational, Pharmaceutical R&D Informatics, Clinical Genomics, Cancer Informatics, tagged gene biology, gene sequencing, genetics, genomics, Harvard, MIT on January 2, 2016| Leave a Comment »

The world’s most innovative intersection

Reported by: Irina Robu

The world’s most innovative crossroads in history is the intersection of

Vassar Street and Main Street, in the new world’s Cambridge, Massachusetts, would be a leading candidate.

According to the article published in Wired Magazine in November 2015 “when the Whitehead got too small for genomicist Eric Lander’s ambitions, he launched a flashier and brasher newcomer next door. The Broad Institute’s gargantuan gleaming glass lobby is filled with early gene-sequencing instruments. Its multimedia screens boast that this is one of the world’s largest gene-sequencing and research factories. The Broad’s strategy is different from that of the Whitehead; instead of concentrating a few in an ultra-exclusive bioclub, Broad bridges MIT, Harvard and most of the hospitals in Boston. Its 2,000 members extend outwards, partnering with tens of thousands of others globally. Those working at the Broad are not averse to commerce; its director alone helped to build Foundation Medicine, Verastem, Millennium, Fidelity Biosciences, Courtagen and Aclara among many other leading companies.

The sixth building on this extraordinary corner, Novartis, focuses on private research, and represents a huge migration from Basel in Switzerland towards the MIT campus, becoming Cambridge’s largest employer. Pfizer, Sanofi, Amgen, Biogen-Idec and hundreds of others cluster nearby. “

Attracting the best and the brightest, one can change not just a city but the world.

Source

http://www.wired.co.uk/magazine/archive/2015/11/ideas-bank/vassar-main-cambridge-massachusetts-innovaton

Read Full Post »

What comes after finishing the Euchromatic Sequence of the Human Genome?

Posted in Chemical Genetics, Genome Biology, tagged gene sequencing, Human Genome Project on April 21, 2014| Leave a Comment »

What comes after finishing the Euchromatic Sequence of the Human Genome?

Author and Curator: Larry H Bernstein, MD, FCAP

Finishing the euchromatic sequence of the human genome.

Oct 21 2004 ; 431(7011): 931-45. http://dx.doi.org/10.1038/nature03001
International Human Genome Sequencing Consortium.

The sequence of the human genome encodes the genetic instructions for human physiology, as well as rich information about human evolution. In 2001, the International Human Genome Sequencing Consortium reported a draft sequence of the euchromatic portion of the human genome. They then worked to complete a genome sequence with high accuracy and completeness. The result of this is reported in Nature (2004), here cited. The current genome sequence (Build 35) contains 2.85 billion nucleotides interrupted by only 341 gaps. It covers approximately 99% of the euchromatic genome and is accurate to an error rate of approximately 1 event per 100,000 bases. Many of the remaining euchromatic gaps are associated with segmental duplications and will require focused work with new methods. The near-complete sequence, the first for a vertebrate, greatly improves the precision of biological analyses of the human genome including studies of gene number, birth and death. Notably, the human genome seems to encode only 20,000-25,000 protein-coding genes
PMID: 15496913 [PubMed – indexed for MEDLINE]

Comment in Human genome: end of the beginning. [Nature. 2004]

Human genome: End of the beginning

Lincoln D. Stein
Nature 21 Oct 2004; 431, 915-916 | http://dx.doi.org/10.1038/431915a

Just over three years ago, a first draft of the human genome sequence had been completed. Gaps and errors remained, but the job of fixing those problems is now largely done.

The featured article in this issue of Nature, entitled “Finishing the euchromatic sequence of the human genome”, has been authored by members of the International Human Genome Sequencing Consortium (IHGSC). It is the latest, but by no means the last, milestone in this historic project.

Early in 2001, the duelling IHGSC (public) and Celera Corporation (private) groups published papers in Nature² and Science³describing the completion of so-called ‘draft’ sequences. These sequences have revolutionized molecular biology by largely eliminating the need to clone and sequence genes involved in human health and disease.

But the draft sequences were far from perfect. Some 10% of the so-called ‘euchromatin’ — the gene-rich portion of the genome — and some 30% of the genome as a whole (which includes the gene-poor regions of ‘heterochromatin’), were not disclosed. There were hundreds of thousands of gaps, and there were misassembled regions where portions of the genome were flipped or misplaced. As a result, large-scale analyses of the genome, had to contend with numerous uncertainties and artefacts. For example, studies of the dying remnants of genes that have accumulated mutations that render them non-functional, left the possibility that such a ‘pseudogene’ was a sequencing error.

Since the publication of the drafts, the IHGSC sequencing centers have quietly undertaken a laborious ‘finishing’ process, in which each gap in the draft was individually examined and subjected to a battery of steps involving cloning and resequencing stretches of DNA. The sequence announced today has just 341 gaps remaining, and consists of contiguous runs of sequence averaging 38 million base pairs. The authors estimate that the finished sequence covers 99% of the euchromatic portion of the genome and that the overall error rate is less than 1 error per 100,000 base pairs. This substantially exceeds the original goals for the project.

The finishing procedure roughly doubled the total time and cost of the project. Does it contribute anything new to our understanding of the genome? It does indeed, and to prove the point the authors of the current paper¹ describe several large-scale analyses of the genome that would have been difficult to perform on the draft sequence. One analysis studied the processes of gene birth and death. The authors find 1,183 human genes that show evidence of having been recently ‘born’ by a process of gene duplication and divergence. They also find 37 genes that seem to have recently ‘died’ by acquiring a mutation that rendered the gene non-functional. The resulting pseudogene then slowly degrades and disappears.

The authors then use the finished sequence to map out segmental duplications — large regions of the genome that have duplicated. They find 5% of the genome involved in segmental duplications, and the duplications are distributed widely across the chromosomes. The nature and extent of such duplications sheds light on the evolution of the human genome, and is needed for studying the many medically relevant disorders that are involved in segmental duplications.

Another paper in this issue, by She et al.⁴ (page 927), directly compares the outcomes of this second analysis with results obtained on an unfinished version of the human genome (an improved version of the Celera draft). She et al. find that the draft version artefactually ‘simplifies’ the genome by eliminating many duplicated regions. Their results bear on one of the highly publicized differences between the public and private genome projects. The public project used an older strategy in which the genome was first cloned into bacterial artificial chromosomes (BACs); the clones were then mapped, and each clone was sequenced and their sequences assembled individually. Celera championed an untested technique, ‘whole-genome shotgun’ (WGS), in which the entire genome was shattered into bite-size pieces, sequenced, and then assembled by software in one conceptually simple step.

Celera proved that the WGS technique is both technically feasible and provides a dramatic cost-saving over the clone-by-clone approach. The Celera draft has had a significant impact on the public project. Almost all genome-sequencing projects since then have used some form of WGS. The cautionary results contained in the new papers from the IHGSC¹ and She et al.⁴argue for a hybrid strategy in which WGS is supplemented by a modest amount of BAC cloning and mapping. This would protect draft WGS sequences from some of the ‘simplification’ reported by She et al. and provide the clones needed for finishing selected regions of special interest.

What is next for the human genome project?

1) Develop the definitive catalogue of protein-coding genes – estimated to be between 20,000 and 25,000.

a) Natural selection ensures that functional regions are more highly conserved than non-functional ones, so a comparative approach highlights candidate protein-coding regions.

b) The same approach shows promise for finding other functional elements such as gene promoters, which control the timing and level of expression of genes, and micro-RNAs, which have been implicated as regulatory agents of many developmental processes.

2) Sequencing the remaining 20% of the genome that lies within heterochromatin, the gene-poor, highly repetitive sequence that is implicated in the processes of chromosome replication and maintenance.

a) The repetitiveness ofheterochromatin means that it cannot be tackled using current sequencing methods, and new technologies will have to be developed to attack it.

We are only at the end of the beginning: ahead lies another mountain range that we will need to map out and explore as we seek to understand how all the parts revealed by the genome sequence work together to make life.

References

International Human Genome Sequencing Consortium Nature 431, 931−945 (2004). | Article |
International Human Genome Sequencing Consortium Nature 409, 860−921 (2001). | Article | PubMed | ISI | ChemPort |
Venter, J. C. et al. Science 291, 1304−1351 (2001). | Article | PubMed | ISI | ChemPort |
She, X. et al. Nature 431, 927−930 (2004). | Article |

Shotgun sequence assembly and recent segmental duplications within the human genome

Xinwei She1, Zhaoshi Jiang1, Royden A. Clark2, Ge Liu2, Ze Cheng1, et al.
Nature 431, 927-930 (21 Oct 2004) | http://dx.doi.org/10.1038/nature03062;

Complex eukaryotic genomes are now being sequenced at an accelerated pace primarily using whole-genome shotgun (WGS) sequence assembly approaches. WGS assembly was initially criticized because of its perceived inability to resolve repeat structures within genomes. Here, we quantify the effect of WGS sequence assembly on large, highly similar repeats by comparison of the segmental duplication content of two different human genome assemblies. Our analysis shows that large (> 15 kilobases) and highly identical (> 97%) duplications are not adequately resolved by WGS assembly. This leads to significant reduction in genome length and the loss of genes embedded within duplications. Comparable analyses of mouse genome assemblies confirm that strict WGS sequence assembly will oversimplify our understanding of mammalian genome structure and evolution; a hybrid strategy using a targeted clone-by-clone approach to resolve duplications is proposed.

Department of Genome Sciences, University of Washington School of Medicine, Seattle, Washington
Department of Genetics, Case Western Reserve University, Cleveland, Ohio
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland
Applied Biosystems, and
The Center for the Advancement of Genomics, Rockville, Maryland

Read Full Post »

Whole exome somatic mutations analysis of malignant melanoma contributes to the development of personalized cancer therapy for this disease

Posted in Biomarkers & Medical Diagnostics, CANCER BIOLOGY & Innovations in Cancer Therapy, Cell Biology, Signaling & Cell Circuits, tagged Cancer - General, Cancer research, Conditions and Diseases, Exome sequencing, gene sequencing, Ipilimumab, Melanoma, Personalized medicine, skin cancer, treatment, Vemurafenib on April 23, 2013| 3 Comments »

Whole exome somatic mutations analysis of malignant melanoma contributes to the development of personalized cancer therapy for this disease

Author: Ziv Raviv, PhD

Article-7.2.2. Whole exome somatic mutations analysis of malignant melanoma contributes to the development of personalized cancer therapy for this disease

Introduction

Cutaneous melanoma is a type of skin cancer that originates in melanocytes, the cells that are producing melanin. While being the least common type of skin cancer, melanoma is the most aggressive one with invasive characteristics and accounts for the majority of death incidences among skin cancers. Melanoma has an annual rate of 160,000 new cases and 48,000 deaths worldwide. Melanoma affects mainly Caucasians exposed to sun high UV irradiation. Among the genetic factors that characterize the disease, BRAF mutation (V600E) is found in most cases of melanoma (80%). Awareness toward risk factors of melanoma should lead to prevention and early detection*. There are several developmental stages (I-IV) of the disease, starting from local non-invasive melanoma, through invasive and high risk melanoma, up to metastatic melanoma. As with other cancers, the earlier stage melanoma is being detected, the better odds for full recovery are. Treatment is usually involving surgery to remove the local tumor and its margins, and when necessary also to remove the proximal lymph node(s) that drain the tumor. In high stages melanoma, adjuvant therapy is given in the form of chemotherapy (Dacarbazine and Temozolomide) and immunotherapy (IL-2 and IFN). Until recently no useful treatment was available for metastatic melanoma. However, research efforts had led to the development of two new drugs against metastatic melanoma: Vemurafenib (Zelboraf), a B-Raf inhibitor; and Ipilimumab (Yervoy), a monoclonal antibody that blocks the inhibitory signal of cytotoxic T lymphocyte-associated antigen 4 (CTLA-4). Both drugs are now available for clinical use presenting good results.

Personalized therapy for melanoma

In an attempt to develop personalized therapies for malignant melanoma, a unique strategy has been taken by the group of Prof. Yardena Samuels at the NIH (now situated at the WIS) to identify recurring genetic alterations of metastatic cutaneous melanoma. The researchers approach employed the collections of hundreds of tumors samples taken from metastasized melanoma patients together with matched normal blood tissues samples. The samples are undergoing exome sequencing for the analysis of somatic mutations (namely mutations that evolved during the progress of the disease to the stage of metastatic melanoma, unlike genomic mutations that may have contribute to the formation of the disease). The discrimination of such tumor related somatic mutations is done by comparison to the exome sequencing of the patient’s matched blood cells DNA. In addition, the malignant cells derived from the removed cancer tissue of each patient are extracted to form a cell line and are grown in culture. These cells are easily cultivate in culture with no special media supplements, nor further genetic manipulations such as hTERT are needed, and are extremely aggressive as determined by various cell culture and in vivo tests. The ability to grow these primary tumor-derived cell lines in culture has a great value as a tool for studying and characterizing the biochemical, functional, and clinical aspects of the mutated genes identified.

In one study [1] Samuels and her colleagues performed this sequencing process for mutation analysis for the protein tyrosine kinase (PTK) gene family, as PTKs are frequently mutated in cancer. Using high-throughput gene sequencing to analyze the entire PTK gene family, the researchers have identified 30 somatic mutations affecting the kinase domains of 19 PTKs and subsequently evaluated the entire coding regions of the genes encoding these 19 PTKs for somatic mutations in 79 melanoma samples. The most frequent mutations were found in ERBB4, a member of the EGFR/ErbB family of receptor tyrosine kinase (RTK), were 19% of melanoma patients had such mutations. Seven missense mutations in the ERBB4 gene were found to induce increased kinase activity and transformation capability. Melanoma derived cell lines that were expressing these mutant ERBB4 forms had reduced cell growth after silencing ERBB4 by RNAi or after treatment with the ERBB inhibitor Lapatinib. Lapatinib is already in use in the clinic for the treatment of HER2 (ErbB2) positive breast cancers patients. Following this study, a clinical trial is now conducted with this drug to evaluate its effect in cutaneous metastatic melanoma patients harboring mutations in ERBB4.

In another study of this group [2], the scientists employed the exome sequencing method to analyze the somatic mutations of 734 G protein coupled receptors (GPCRs) in melanoma. GPCRs are regulating various signaling pathways including those that affect cell growth and play also important role in human diseases. This screen revealed that GRM3 gene that encode the metabotropic glutamate receptor 3 (mGluR3), was frequently mutated and that one of its mutations clustered within one position. Mutant GRM3 was found to selectively regulate the phosphorylation of MEK1 leading to increased anchorage-independent cell growth and cellular migration. Tumor derived melanoma cells expressing mutant GRM3 exhibited reduced cell growth and migration upon knockdown of GRM3 by RNAi or by treatment with the selective MEK inhibitor, Selumetinib (AZD-6244), a drug that is being testing in clinical trials. Altogether, the results of this study point to the increased violent characteristics of melanomas bearing mutational GRM3.

In a third study, melanoma samples were examined for somatic mutations in 19 human genes that encode ADAMTS proteins [3]. Some of the ADAMTS genes have been suggested before to have implication in tumorigenesis. ADAMTS18, which was previously found to be a candidate cancer gene, was found in this study to be highly mutated in melanoma. ADAMTS18 mutations were biologically examined and were found to induce an increased proliferation of melanoma cells, as well as increased cell migration and metastasis. Moreover, melanoma cells expressing these mutated ADAMTS18 had reduced cell migration after RNAi-mediated knockdown of ADAMTS18. Thus, these results suggest that genetic alteration of ADAMTS18 plays a major role in melanoma tumorigenesis. Since ADAMTS genes encode extracellular proteins, their accessibility to systematically delivered drugs makes them excellent therapeutic targets.

Conclusive remarks

The above illustrated research approach intends to discover frequent melanoma-specific mutations by employing high-throughput whole exome and genome sequencing means. For the most highly mutated genes identified, the biochemical, functional, and clinical aspects are being characterized to examine their relevancy to the disease outcomes. This approach therefore introduces new opportunities for clinical intervention for the treatment of cutaneous melanoma. In addition to the discovery of novel highly mutated genes, this approach may also help determine which pathways are altered in melanoma and how these genes and pathways interact. Finding melanoma-associated highly mutated genes could lead to personalized therapeutics specifically targeting these altered genes in individual melanomas. Along with the opportunity to develop new agents to treat melanoma, the approach takes advantage of existing anti-cancer drugs, utilizing them to treat these mutated genes melanoma individuals. In addition to their potential for therapeutics, the discovery of highly mutated genes in melanoma patients may lead to the discovery of new markers that may assist the diagnosis of the disease. The implications of these screenings findings on other types of cancer bearing common pathways similar to melanoma should be examined as well. Finally, this elegant approach should be adopted in research efforts of other cancer types.

* Special review will be published further in the cancer prevention section of Pharmaceutical Intelligence

References

1. Prickett TD, Agrawal NS, Wei X, Yates KE, Lin JC, Wunderlich JR, Cronin JC, Cruz P, Rosenberg SA, Samuels Y (2009) Analysis of the tyrosine kinome in melanoma reveals recurrent mutations in ERBB4. Nat Genet 41 (10):1127-1132

2. Prickett TD, Wei X, Cardenas-Navia I, Teer JK, Lin JC, Walia V, Gartner J, Jiang J, Cherukuri PF, Molinolo A, Davies MA, Gershenwald JE, Stemke-Hale K, Rosenberg SA, Margulies EH, Samuels Y (2011) Exon capture analysis of G protein-coupled receptors identifies activating mutations in GRM3 in melanoma. Nat Genet 43 (11):1119-1126

3. Wei X, Prickett TD, Viloria CG, Molinolo A, Lin JC, Cardenas-Navia I, Cruz P, Rosenberg SA, Davies MA, Gershenwald JE, Lopez-Otin C, Samuels Y (2010) Mutational and functional analysis reveals ADAMTS18 metalloproteinase as a novel driver in melanoma. Mol Cancer Res 8 (11):1513-1525

Related articles on melanoma on this open access online scientific journal:

1. In focus: Melanoma Genetics. Curator: Ritu Saxena, Ph.D.

2. In focus: Melanoma therapeutics. Author and Curator: Ritu Saxena, Ph.D.

3. A New Therapy for Melanoma. Reporter- Larry H Bernstein, M.D.

4. Thymosin alpha1 and melanoma. Author, Editor: Tilda Barliya, Ph.D.

5. Exome sequencing of serous endometrial tumors shows recurrent somatic mutations in chromatin-remodeling and ubiquitin ligase complex genes. Reporter and Curator: Dr. Sudipta Saha, Ph.D.

6. How Genome Sequencing is Revolutionizing Clinical Diagnostics. Reporter: Aviva Lev-Ari, PhD, RN.

7. Issues in Personalized Medicine in Cancer: Intratumor Heterogeneity and Branched Evolution Revealed by Multiregion Sequencing. Curator and Reporter: Stephen J. Williams, Ph.D.

Read Full Post »

Leaders in Pharmaceutical Business Intelligence Group, LLC, Doing Business As LPBI Group, Newton, MA

Posts Tagged ‘gene sequencing’

The Use of ChatGPT in the World of BioInformatics and Cancer Research and Development of BioGPT by MIT

The Use of ChatGPT in the World of BioInformatics and Cancer Research and Development of BioGPT by MIT

@tzvi7989

Microsoft’s BioGPT Shows Promise as the Best Biomedical NLP

@rufus9322

Other Relevant Articles on Natural Language Processing in BioInformatics, Healthcare and ChatGPT for Medicine on this Open Access Scientific Journal Include

Medicine with GPT-4 & ChatGPT

Explanation on “Results of Medical Text Analysis with Natural Language Processing (NLP) presented in LPBI Group’s NEW GENRE Edition: NLP” on Genomics content, standalone volume in Series B and NLP on Cancer content as Part B New Genre Volume 1 in Series C

Like this:

The world’s most innovative intersection

The world’s most innovative intersection

Like this:

What comes after finishing the Euchromatic Sequence of the Human Genome?

What comes after finishing the Euchromatic Sequence of the Human Genome?

Finishing the euchromatic sequence of the human genome.

Human genome: End of the beginning

Shotgun sequence assembly and recent segmental duplications within the human genome

Like this:

Whole exome somatic mutations analysis of malignant melanoma contributes to the development of personalized cancer therapy for this disease

Whole exome somatic mutations analysis of malignant melanoma contributes to the development of personalized cancer therapy for this disease

Like this:

Follow Blog via Email

Recent Posts

Archives

Categories

Meta

Posts Tagged ‘gene sequencing’

The Use of ChatGPT in the World of BioInformatics and Cancer Research and Development of BioGPT by MIT

Microsoft’s BioGPT Shows Promise as the Best Biomedical NLP

Other Relevant Articles on Natural Language Processing in BioInformatics, Healthcare and ChatGPT for Medicine on this Open Access Scientific Journal Include

Share this:

Like this:

The world’s most innovative intersection

Share this:

Like this:

What comes after finishing the Euchromatic Sequence of the Human Genome?

Finishing the euchromatic sequence of the human genome.

Human genome: End of the beginning

Shotgun sequence assembly and recent segmental duplications within the human genome

Share this:

Like this:

Whole exome somatic mutations analysis of malignant melanoma contributes to the development of personalized cancer therapy for this disease

Share this:

Like this:

Follow Blog via Email

Recent Posts

Archives

Categories

Meta