Feeds:
Posts
Comments

Posts Tagged ‘Artificial intelligence’

Google, Verily’s Uses AI to Screen for Diabetic Retinopathy

Reporter : Irina Robu, PhD

Google and Verily, the life science research organization under Alphabet designed a machine learning algorithm to better screen for diabetes and associated eye diseases. Google and Verily believe the algorithm can be beneficial in areas lacking optometrists.

The algorithm is being integrated for the first time in a clinical setting at Aravind Eye Hospital in Madurai, India where it is designed to screen for diabetic retinopathy and diabetic macular edema. After a patient is imaged by trained staff using a fundus camera, the image is uploaded to the screening algorithm through management software. The algorithm then analyzes the images for the diabetic eye diseases before returning the results.

Numerous AI-driven approaches have lately been effective in detecting diabetic retinopathy with high accuracy. An AI-based grading system was able to effectively diagnose two patients with the disease. Furthermore, an AI-driven approach for detecting an early sign of diabetic retinopathy attained an accuracy rate of more than 98 percent.

According to the R. Usha Kim, Chief of retina services at the Aravind Eye Hospital the algorithm permits physicians to work closely with patients on treatment and management of their disease, whereas increasing the volume of screenings we can perform. Automated grading of diabetic retinopathy has possible benefits such as increasing efficiency, reproducible, and coverage of screening programs and improving patient outcomes by providing early detection and treatment.

Even if the technology sounds promising, current research show there are long way until it can directly transfer from the lab into clinic.

SOURCE
https://www.healthcareitnews.com/news/google-verily-using-ai-screen-diabetic-retinopathy-india

Read Full Post »

Artificial intelligence can be a useful tool to predict Alzheimer

Reporter: Irina Robu, PhD

3.3.10

3.3.10   Artificial intelligence can be a useful tool to predict Alzheimer, Volume 2 (Volume Two: Latest in Genomics Methodologies for Therapeutics: Gene Editing, NGS and BioInformatics, Simulations and the Genome Ontology), Part 2: CRISPR for Gene Editing and DNA Repair

The Alzheimer’s Association estimate that around 5.7 million people live with Alzheimer’s disease in the United States which will rise to almost 14 million by 2050. Earlier diagnosis would not only benefit those affected, but it could also jointly save about $7.9 trillion in medical care and related costs over time. As Alzheimer’s disease progresses, it changes how brain cells use glucose. This alteration in glucose metabolism shows up in a type of PET imaging that tracks the uptake of a radioactive form of glucose called 18F-fluorodeoxyglucose. By giving instructions about what to look for, the scientists were able to train the deep learning algorithm to assess the PET images for early signs of Alzheimer’s.
The researchers from University of California San Francisco used positron-emission tomography images of 1002 people’s brain to train the deep learning algorithm they developed. They used 90 percent of images to teach the algorithm to spot features of Alzheimer’s disease and the remaining 10 percent to verify its performance. The researchers tested the algorithm on PET images of brains from 40 people, from which they were able to predict which individuals would receive a final diagnosis of Alzheimer’s. On average, the people who were tested were diagnosed with the disease more than 6 years after the scans.
According to the Radiology journal in which the research was published, the team describes how the algorithm “achieved 82 percent specificity at 100 percent sensitivity, an average of 75.8 months prior to the final diagnosis.” The researchers taught the algorithm with the help of more than 2,109 PET images of 1,002 individuals’ brains. The algorithm uses deep learning, which allows the algorithm to “teach itself” what to look for by spotting subtle differences among the thousands of images. The algorithm was as good as, if not better than, human experts at analyzing the FDG PET images.
Future advances will involve using larger data sets and additional images taken over time from people at various clinics and institutions. In the future, the algorithm could be a beneficial addition to the radiologist’s toolbox and advance opportunities for the early treatment of Alzheimer’s disease.

Source

https://www.medicalnewstoday.com/articles/323608.php

Read Full Post »

2,000 human brains yield clues to how genes raise risk for mental illnesses

Reporter: Irina Robu, PhD

It’s one thing to detect sites in the genome associated with mental disorders; it’s quite another to discover the biological mechanisms by which these changes in DNA work in the human brain to boost risk. In their first concerted effort to tackle the problem, 15 collaborating research teams of the National Institutes of Health-funded PsychENCODE Consortium evaluated data of 2000 human brains which might yield clues to how genes raise risk for mental illnesses.

Applying newly uncovered secrets of the brain’s molecular architecture, they established an artificial intelligence model that is six times better than preceding ones at predicting risk for mental disorders. They also identified several hundred previously unknown risk genes for mental illnesses and linked many known risk variants to specific genes. In the brain tissue and single cells, the researchers identified patterns of gene expression, marks in gene regulation as well as genetic variants that can be linked to mental illnesses.

Dr. Nenad Sestan of Yale University explained that “ the consortium’s integrative genomic analyses elucidate the mechanisms by which cellular diversity and patterns of gene expression change throughout development and reveal how neuropsychiatric risk genes are concentrated into distinct co-expression modules and cell types”. The implicated variants are typically small-effect genetic variations that fall within regions of the genome that don’t code for proteins, but instead are thought to regulate gene expression and other aspects of gene function.

In addition to the 2000 postmortem human brains, researchers examined brain tissue from prenatal development as well as people with schizophrenia, bipolar disorder,  and typical development compared findings with parallel data from non-human primates. Their findings indicate that gene variants linked to mental illnesses exert more effects when they jointly form “modules”, communicating genes with related functions and at specific developmental time points that seem to coincide with the course of illness. Variability in risk gene expression and cell types increases during formative stages in early prenatal development and again during the teen years. However, in postmortem brains of people with a mental illness, thousands of RNAs were found to have anomalies.

According to NIMH, Geetha Senthil the multi-omic data resource caused by the PsychENCODE collaboration will pave a path for building molecular models of disease and developmental processes and may offer a platform for target identification for pharmaceutical research.

SOURCE
https://www.nih.gov/news-events/news-releases/2000-human-brains-yield-clues-how-genes-raise-risk-mental-illnesses

Read Full Post »

Can Blockchain Technology and Artificial Intelligence Cure What Ails Biomedical Research and Healthcare, Volume 2 (Volume Two: Latest in Genomics Methodologies for Therapeutics: Gene Editing, NGS and BioInformatics, Simulations and the Genome Ontology), Part 1: Next Generation Sequencing (NGS)

Can Blockchain Technology and Artificial Intelligence Cure What Ails Biomedical Research and Healthcare

Curator: Stephen J. Williams, Ph.D.

Updated 12/18/2018

In the efforts to reduce healthcare costs, provide increased accessibility of service for patients, and drive biomedical innovations, many healthcare and biotechnology professionals have looked to advances in digital technology to determine the utility of IT to drive and extract greater value from healthcare industry.  Two areas of recent interest have focused how best to use blockchain and artificial intelligence technologies to drive greater efficiencies in our healthcare and biotechnology industries.

More importantly, with the substantial increase in ‘omic data generated both in research as well as in the clinical setting, it has become imperative to develop ways to securely store and disseminate the massive amounts of ‘omic data to various relevant parties (researchers or clinicians), in an efficient manner yet to protect personal privacy and adhere to international regulations.  This is where blockchain technologies may play an important role.

A recent Oncotarget paper by Mamoshina et al. (1) discussed the possibility that next-generation artificial intelligence and blockchain technologies could synergize to accelerate biomedical research and enable patients new tools to control and profit from their personal healthcare data, and assist patients with their healthcare monitoring needs. According to the abstract:

The authors introduce new concepts to appraise and evaluate personal records, including the combination-, time- and relationship value of the data.  They also present a roadmap for a blockchain-enabled decentralized personal health data ecosystem to enable novel approaches for drug discovery, biomarker development, and preventative healthcare.  In this system, blockchain and deep learning technologies would provide the secure and transparent distribution of personal data in a healthcare marketplace, and would also be useful to resolve challenges faced by the regulators and return control over personal data including medical records to the individual.

The review discusses:

  1. Recent achievements in next-generation artificial intelligence
  2. Basic concepts of highly distributed storage systems (HDSS) as a preferred method for medical data storage
  3. Open source blockchain Exonium and its application for healthcare marketplace
  4. A blockchain-based platform allowing patients to have control of their data and manage access
  5. How advances in deep learning can improve data quality, especially in an era of big data

Advances in Artificial Intelligence

  • Integrative analysis of the vast amount of health-associated data from a multitude of large scale global projects has proven to be highly problematic (REF 27), as high quality biomedical data is highly complex and of a heterogeneous nature, which necessitates special preprocessing and analysis.
  • Increased computing processing power and algorithm advances have led to significant advances in machine learning, especially machine learning involving Deep Neural Networks (DNNs), which are able to capture high-level dependencies in healthcare data. Some examples of the uses of DNNs are:
  1. Prediction of drug properties(2, 3) and toxicities(4)
  2. Biomarker development (5)
  3. Cancer diagnosis (6)
  4. First FDA approved system based on deep learning Arterys Cardio DL
  • Other promising systems of deep learning include:
    • Generative Adversarial Networks (https://arxiv.org/abs/1406.2661): requires good datasets for extensive training but has been used to determine tumor growth inhibition capabilities of various molecules (7)
    • Recurrent neural Networks (RNN): Originally made for sequence analysis, RNN has proved useful in analyzing text and time-series data, and thus would be very useful for electronic record analysis. Has also been useful in predicting blood glucose levels of Type I diabetic patients using data obtained from continuous glucose monitoring devices (8)
    • Transfer Learning: focused on translating information learned on one domain or larger dataset to another, smaller domain. Meant to reduce the dependence on large training datasets that RNN, GAN, and DNN require.  Biomedical imaging datasets are an example of use of transfer learning.
    • One and Zero-Shot Learning: retains ability to work with restricted datasets like transfer learning. One shot learning aimed to recognize new data points based on a few examples from the training set while zero-shot learning aims to recognize new object without seeing the examples of those instances within the training set.

Highly Distributed Storage Systems (HDSS)

The explosion in data generation has necessitated the development of better systems for data storage and handling. HDSS systems need to be reliable, accessible, scalable, and affordable.  This involves storing data in different nodes and the data stored in these nodes are replicated which makes access rapid. However data consistency and affordability are big challenges.

Blockchain is a distributed database used to maintain a growing list of records, in which records are divided into blocks, locked together by a crytosecurity algorithm(s) to maintain consistency of data.  Each record in the block contains a timestamp and a link to the previous block in the chain.  Blockchain is a distributed ledger of blocks meaning it is owned and shared and accessible to everyone.  This allows a verifiable, secure, and consistent history of a record of events.

Data Privacy and Regulatory Issues

The establishment of the Health Insurance Portability and Accountability Act (HIPAA) in 1996 has provided much needed regulatory guidance and framework for clinicians and all concerned parties within the healthcare and health data chain.  The HIPAA act has already provided much needed guidance for the latest technologies impacting healthcare, most notably the use of social media and mobile communications (discussed in this article  Can Mobile Health Apps Improve Oral-Chemotherapy Adherence? The Benefit of Gamification.).  The advent of blockchain technology in healthcare offers its own unique challenges however HIPAA offers a basis for developing a regulatory framework in this regard.  The special standards regarding electronic data transfer are explained in HIPAA’s Privacy Rule, which regulates how certain entities (covered entities) use and disclose individual identifiable health information (Protected Health Information PHI), and protects the transfer of such information over any medium or electronic data format. However, some of the benefits of blockchain which may revolutionize the healthcare system may be in direct contradiction with HIPAA rules as outlined below:

Issues of Privacy Specific In Use of Blockchain to Distribute Health Data

  • Blockchain was designed as a distributed database, maintained by multiple independent parties, and decentralized
  • Linkage timestamping; although useful in time dependent data, proof that third parties have not been in the process would have to be established including accountability measures
  • Blockchain uses a consensus algorithm even though end users may have their own privacy key
  • Applied cryptography measures and routines are used to decentralize authentication (publicly available)
  • Blockchain users are divided into three main categories: 1) maintainers of blockchain infrastructure, 2) external auditors who store a replica of the blockchain 3) end users or clients and may have access to a relatively small portion of a blockchain but their software may use cryptographic proofs to verify authenticity of data.

 

YouTube video on How #Blockchain Will Transform Healthcare in 25 Years (please click below)

 

 

In Big Data for Better Outcomes, BigData@Heart, DO->IT, EHDN, the EU data Consortia, and yes, even concepts like pay for performance, Richard Bergström has had a hand in their creation. The former Director General of EFPIA, and now the head of health both at SICPA and their joint venture blockchain company Guardtime, Richard is always ahead of the curve. In fact, he’s usually the one who makes the curve in the first place.

 

 

 

Please click on the following link for a podcast on Big Data, Blockchain and Pharma/Healthcare by Richard Bergström:

https://soundcloud.com/vitalhealth/real-world-data-pay-for-performance-or-blockchain-richard-bergstrom-is-always-ahead-of-the-curve

References

  1. Mamoshina, P., Ojomoko, L., Yanovich, Y., Ostrovski, A., Botezatu, A., Prikhodko, P., Izumchenko, E., Aliper, A., Romantsov, K., Zhebrak, A., Ogu, I. O., and Zhavoronkov, A. (2018) Converging blockchain and next-generation artificial intelligence technologies to decentralize and accelerate biomedical research and healthcare, Oncotarget 9, 5665-5690.
  2. Aliper, A., Plis, S., Artemov, A., Ulloa, A., Mamoshina, P., and Zhavoronkov, A. (2016) Deep Learning Applications for Predicting Pharmacological Properties of Drugs and Drug Repurposing Using Transcriptomic Data, Molecular pharmaceutics 13, 2524-2530.
  3. Wen, M., Zhang, Z., Niu, S., Sha, H., Yang, R., Yun, Y., and Lu, H. (2017) Deep-Learning-Based Drug-Target Interaction Prediction, Journal of proteome research 16, 1401-1409.
  4. Gao, M., Igata, H., Takeuchi, A., Sato, K., and Ikegaya, Y. (2017) Machine learning-based prediction of adverse drug effects: An example of seizure-inducing compounds, Journal of pharmacological sciences 133, 70-78.
  5. Putin, E., Mamoshina, P., Aliper, A., Korzinkin, M., Moskalev, A., Kolosov, A., Ostrovskiy, A., Cantor, C., Vijg, J., and Zhavoronkov, A. (2016) Deep biomarkers of human aging: Application of deep neural networks to biomarker development, Aging 8, 1021-1033.
  6. Vandenberghe, M. E., Scott, M. L., Scorer, P. W., Soderberg, M., Balcerzak, D., and Barker, C. (2017) Relevance of deep learning to facilitate the diagnosis of HER2 status in breast cancer, Scientific reports 7, 45938.
  7. Kadurin, A., Nikolenko, S., Khrabrov, K., Aliper, A., and Zhavoronkov, A. (2017) druGAN: An Advanced Generative Adversarial Autoencoder Model for de Novo Generation of New Molecules with Desired Molecular Properties in Silico, Molecular pharmaceutics 14, 3098-3104.
  8. Ordonez, F. J., and Roggen, D. (2016) Deep Convolutional and LSTM Recurrent Neural Networks for Multimodal Wearable Activity Recognition, Sensors (Basel) 16.

Articles from clinicalinformaticsnews.com

Healthcare Organizations Form Synaptic Health Alliance, Explore Blockchain’s Impact On Data Quality

From http://www.clinicalinformaticsnews.com/2018/12/05/healthcare-organizations-form-synaptic-health-alliance-explore-blockchains-impact-on-data-quality.aspx

By Benjamin Ross

December 5, 2018 | The boom of blockchain and distributed ledger technologies have inspired healthcare organizations to test the capabilities of their data. Quest Diagnostics, in partnership with Humana, MultiPlan, and UnitedHealth Group’s Optum and UnitedHealthcare, have launched a pilot program that applies blockchain technology to improve data quality and reduce administrative costs associated with changes to healthcare provider demographic data.

The collective body, called Synaptic Health Alliance, explores how blockchain can keep only the most current healthcare provider information available in health plan provider directories. The alliance plans to share their progress in the first half of 2019.

Providing consumers looking for care with accurate information when they need it is essential to a high-functioning overall healthcare system, Jason O’Meara, Senior Director of Architecture at Quest Diagnostics, told Clinical Informatics News in an email interview.

“We were intentional about calling ourselves an alliance as it speaks to the shared interest in improving health care through better, collaborative use of an innovative technology,” O’Meara wrote. “Our large collective dataset and national footprints enable us to prove the value of data sharing across company lines, which has been limited in healthcare to date.”

O’Meara said Quest Diagnostics has been investing time and resources the past year or two in understanding blockchain, its ability to drive purpose within the healthcare industry, and how to leverage it for business value.

“Many health care and life science organizations have cast an eye toward blockchain’s potential to inform their digital strategies,” O’Meara said. “We recognize it takes time to learn how to leverage a new technology. We started exploring the technology in early 2017, but we quickly recognized the technology’s value is in its application to business to business use cases: to help transparently share information, automate mutually-beneficial processes and audit interactions.”

Quest began discussing the potential for an alliance with the four other companies a year ago, O’Meara said. Each company shared traits that would allow them to prove the value of data sharing across company lines.

“While we have different perspectives, each member has deep expertise in healthcare technology, a collaborative culture, and desire to continuously improve the patient/customer experience,” said O’Meara. “We also recognize the value of technology in driving efficiencies and quality.”

Following its initial launch in April, Synaptic Health Alliance is deploying a multi-company, multi-site, permissioned blockchain. According to a whitepaper published by Synaptic Health, the choice to use a permissioned blockchain rather than an anonymous one is crucial to the alliance’s success.

“This is a more effective approach, consistent with enterprise blockchains,” an alliance representative wrote. “Each Alliance member has the flexibility to deploy its nodes based on its enterprise requirements. Some members have elected to deploy their nodes within their own data centers, while others are using secured public cloud services such as AWS and Azure. This level of flexibility is key to growing the Alliance blockchain network.”

As the pilot moves forward, O’Meara says the Alliance plans to open ability to other organizations. Earlier this week Aetna and Ascension announced they joined the project.

“I am personally excited by the amount of cross-company collaboration facilitated by this project,” O’Meara says. “We have already learned so much from each other and are using that knowledge to really move the needle on improving healthcare.”

 

US Health And Human Services Looks To Blockchain To Manage Unstructured Data

http://www.clinicalinformaticsnews.com/2018/11/29/us-health-and-human-services-looks-to-blockchain-to-manage-unstructured-data.aspx

By Benjamin Ross

November 29, 2018 | The US Department of Health and Human Services (HHS) is making waves in the blockchain space. The agency’s Division of Acquisition (DA) has developed a new system, called Accelerate, which gives acquisition teams detailed information on pricing, terms, and conditions across HHS in real-time. The department’s Associate Deputy Assistant Secretary for Acquisition, Jose Arrieta, gave a presentation and live demo of the blockchain-enabled system at the Distributed: Health event earlier this month in Nashville, Tennessee.

Accelerate is still in the prototype phase, Arrieta said, with hopes that the new system will be deployed at the end of the fiscal year.

HHS spends around $25 billion a year in contracts, Arrieta said. That’s 100,000 contracts a year with over one million pages of unstructured data managed through 45 different systems. Arrieta and his team wanted to modernize the system.

“But if you’re going to change the way a workforce of 20,000 people do business, you have to think your way through how you’re going to do that,” said Arrieta. “We didn’t disrupt the existing systems: we cannibalized them.”

The cannibalization process resulted in Accelerate. According to Arrieta, the system functions by creating a record of data rather than storing it, leveraging machine learning, artificial intelligence (AI), and robotic process automation (RPA), all through blockchain data.

“We’re using that data record as a mechanism to redesign the way we deliver services through micro-services strategies,” Arrieta said. “Why is that important? Because if you have a single application or data use that interfaces with 55 other applications in your business network, it becomes very expensive to make changes to one of the 55 applications.”

Accelerate distributes the data to the workforce, making it available to them one business process at a time.

“We’re building those business processes without disrupting the existing systems,” said Arrieta, and that’s key. “We’re not shutting off those systems. We’re using human-centered design sessions to rebuild value exchange off of that data.”

The first application for the system, Arrieta said, can be compared to department stores price-matching their online competitors.

It takes the HHS close to a month to collect the amalgamation of data from existing system, whether that be terms and conditions that drive certain price points, or software licenses.

“The micro-service we built actually analyzes that data, and provides that information to you within one second,” said Arrieta. “This is distributed to the workforce, to the 5,000 people that do the contracting, to the 15,000 people that actually run the programs at [HHS].”

This simple micro-service is replicated on every node related to HHS’s internal workforce. If somebody wants to change the algorithm to fit their needs, they can do that in a distributed manner.

Arrieta hopes to use Accelerate to save researchers money at the point of purchase. The program uses blockchain to simplify the process of acquisition.

“How many of you work with the federal government?” Arrieta asked the audience. “Do you get sick of reentering the same information over and over again? Every single business opportunity you apply for, you have to resubmit your financial information. You constantly have to check for validation and verification, constantly have to resubmit capabilities.”

Wouldn’t it be better to have historical notes available for each transaction? said Arrieta. This would allow clinical researchers to be able to focus on “the things they’re really good at,” instead of red tape.

“If we had the top cancer researcher in the world, would you really want her spending her time learning about federal regulations as to how to spend money, or do you want her trying to solve cancer?” Arrieta said. “What we’re doing is providing that data to the individual in a distributed manner so they can read the information of historical purchases that support activity, and they can focus on the objectives and risks they see as it relates to their programming and their objectives.”

Blockchain also creates transparency among researchers, Arrieta said, which says creates an “uncomfortable reality” in the fact that they have to make a decision regarding data, fundamentally changing value exchange.

“The beauty of our business model is internal investment,” Arrieta said. For instance, the HHS could take all the sepsis data that exists in their system, put it into a distributed ledger, and share it with an external source.

“Maybe that could fuel partnership,” Arrieta said. “I can make data available to researchers in the field in real-time so they can actually test their hypothesis, test their intuition, and test their imagination as it relates to solving real-world problems.”

 

Shivom is creating a genomic data hub to elongate human life with AI

From VentureBeat.com
Blockchain-based genomic data hub platform Shivom recently reached its $35 million hard cap within 15 seconds of opening its main token sale. Shivom received funding from a number of crypto VC funds, including Collinstar, Lateral, and Ironside.

The goal is to create the world’s largest store of genomic data while offering an open web marketplace for patients, data donors, and providers — such as pharmaceutical companies, research organizations, governments, patient-support groups, and insurance companies.

“Disrupting the whole of the health care system as we know it has to be the most exciting use of such large DNA datasets,” Shivom CEO Henry Ines told me. “We’ll be able to stratify patients for better clinical trials, which will help to advance research in precision medicine. This means we will have the ability to make a specific drug for a specific patient based on their DNA markers. And what with the cost of DNA sequencing getting cheaper by the minute, we’ll also be able to sequence individuals sooner, so young children or even newborn babies could be sequenced from birth and treated right away.”

While there are many solutions examining DNA data to explain heritage, intellectual capabilities, health, and fitness, the potential of genomic data has largely yet to be unlocked. A few companies hold the monopoly on genomic data and make sizeable profits from selling it to third parties, usually without sharing the earnings with the data donor. Donors are also not informed if and when their information is shared, nor do they have any guarantee that their data is secure from hackers.

Shivom wants to change that by creating a decentralized platform that will break these monopolies, democratizing the processes of sharing and utilizing the data.

“Overall, large DNA datasets will have the potential to aid in the understanding, prevention, diagnosis, and treatment of every disease known to mankind, and could create a future where no diseases exist, or those that do can be cured very easily and quickly,” Ines said. “Imagine that, a world where people do not get sick or are already aware of what future diseases they could fall prey to and so can easily prevent them.”

Shivom’s use of blockchain technology and smart contracts ensures that all genomic data shared on the platform will remain anonymous and secure, while its OmiX token incentivizes users to share their data for monetary gain.

Rise in Population Genomics: Local Government in India Will Use Blockchain to Secure Genetic Data

Blockchain will secure the DNA database for 50 million citizens in the eighth-largest state in India. The government of Andhra Pradesh signed a Memorandum of Understanding with a German genomics and precision medicine start-up, Shivom, which announced to start the pilot project soon. The move falls in line with a trend for governments turning to population genomics, and at the same time securing the sensitive data through blockchain.

Andhra Pradesh, DNA, and blockchain

Storing sensitive genetic information safely and securely is a big challenge. Shivom builds a genomic data-hub powered by blockchain technology. It aims to connect researchers with DNA data donors thus facilitating medical research and the healthcare industry.

With regards to Andhra Pradesh, the start-up will first launch a trial to determine the viability of their technology for moving from a proactive to a preventive approach in medicine, and towards precision health. “Our partnership with Shivom explores the possibilities of providing an efficient way of diagnostic services to patients of Andhra Pradesh by maintaining the privacy of the individual data through blockchain technologies,” said J A Chowdary, IT Advisor to Chief Minister, Government of Andhra Pradesh.

Other Articles in this Open Access Journal on Digital Health include:

Can Mobile Health Apps Improve Oral-Chemotherapy Adherence? The Benefit of Gamification.

Medical Applications and FDA regulation of Sensor-enabled Mobile Devices: Apple and the Digital Health Devices Market

 

How Social Media, Mobile Are Playing a Bigger Part in Healthcare

 

E-Medical Records Get A Mobile, Open-Sourced Overhaul By White House Health Design Challenge Winners

 

Medcity Converge 2018 Philadelphia: Live Coverage @pharma_BI

 

Digital Health Breakthrough Business Models, June 5, 2018 @BIOConvention, Boston, BCEC

 

 

 

 

 

 

Read Full Post »

Live Coverage: MedCity Converge 2018 Philadelphia: AI in Cancer and Keynote Address

Reporter: Stephen J. Williams, PhD

3.3.4

3.3.4   Live Coverage: MedCity Converge 2018 Philadelphia: AI in Cancer and Keynote Address, Volume 2 (Volume Two: Latest in Genomics Methodologies for Therapeutics: Gene Editing, NGS and BioInformatics, Simulations and the Genome Ontology), Part 2: CRISPR for Gene Editing and DNA Repair

8:30 AM -9:15

Practical Applications of AI in Cancer

We are far from machine learning dictating clinical decision making, but AI has important niche applications in oncology. Hear from a panel of innovative startups and established life science players about how machine learning and AI can transform different aspects in healthcare, be it in patient recruitment, data analysis, drug discovery or care delivery.

Moderator: Ayan Bhattacharya, Advanced Analytics Specialist Leader, Deloitte Consulting LLP
Speakers:
Wout Brusselaers, CEO and Co-Founder, Deep 6 AI @woutbrusselaers ‏
Tufia Haddad, M.D., Chair of Breast Medical Oncology and Department of Oncology Chair of IT, Mayo Clinic
Carla Leibowitz, Head of Corporate Development, Arterys @carlaleibowitz
John Quackenbush, Ph.D., Professor and Director of the Center for Cancer Computational Biology, Dana-Farber Cancer Institute

Ayan: working at IBM and Thompon Rueters with structured datasets and having gone through his own cancer battle, he is now working in healthcare AI which has an unstructured dataset(s)

Carla: collecting medical images over the world, mainly tumor and calculating tumor volumetrics

Tufia: drug resistant breast cancer clinician but interested in AI and healthcareIT at Mayo

John: taking large scale datasets but a machine learning skeptic

moderator: how has imaging evolved?

Carla: ten times images but not ten times radiologists so stressed field needs help with image analysis; they have seen measuring lung tumor volumetrics as a therapeutic diagnostic has worked

moderator: how has AI affected patient recruitment?

Tufia: majority of patients are receiving great care but AI can offer profiles and determine which patients can benefit from tertiary care;

John: 1980 paper on no free lunch theorem; great enthusiasm about optimization algortihisms fell short in application; can extract great information from e.g. images

moderator: how is AI for healthcare delivery working at mayo?

Tufia: for every hour with patient two hours of data mining. for care delivery hope to use the systems to leverage the cognitive systems to do the data mining

John: problem with irreproducible research which makes a poor dataset:  also these care packages are based on population data not personalized datasets; challenges to AI is moving correlation to causation

Carla: algorithisms from on healthcare network is not good enough, Google tried and it failed

John: curation very important; good annotation is needed; needed to go in and develop, with curators, a systematic way to curate medial records; need standardization and reproducibility; applications in radiometrics can be different based on different data collection machines; developed a machine learning model site where investigators can compare models on a hub; also need to communicate with patients on healthcare information and quality information

Ayan: Australia and Canada has done the most concerning AI and lifescience, healthcare space; AI in most cases is cognitive learning: really two types of companies 1) the Microsofts, Googles, and 2) the startups that may be more pure AI

Final Notes: We are at a point where collecting massive amounts of healthcare related data is simple, rapid, and shareable.  However challenges exist in quality of datasets, proper curation and annotation, need for collaboration across all healthcare stakeholders including patients, and dissemination of useful and accurate information

9:15 AM–9:45 AM

Opening Keynote: Dr. Joshua Brody, Medical Oncologist, Mount Sinai Health System

The Promise and Hype of Immunotherapy

Immunotherapy is revolutionizing oncology care across various types of cancers, but it is also necessary to sort the hype from the reality. In his keynote, Dr. Brody will delve into the history of this new therapy mode and how it has transformed the treatment of lymphoma and other diseases. He will address the hype surrounding it, why so many still don’t respond to the treatment regimen and chart the way forward—one that can lead to more elegant immunotherapy combination paths and better outcomes for patients.

Speaker:
Joshua Brody, M.D., Assistant Professor, Mount Sinai School of Medicine @joshuabrodyMD

Director Lymphoma therapy at Mt. Sinai

  • lymphoma a cancer with high PD-L1 expression
  • hodgkin’s lymphoma best responder to PD1 therapy (nivolumab) but hepatic adverse effects
  • CAR-T (chimeric BCR and TCR); a long process which includes apheresis, selection CD3/CD28 cells; viral transfection of the chimeric; purification
  • complete remissions of B cell lymphomas (NCI trial) and long term remissions past 18 months
  • side effects like cytokine release (has been controlled); encephalopathy (he uses a hand writing test to see progression of adverse effect)

Vaccines

  •  teaching the immune cells as PD1 inhibition exhausting T cells so a vaccine boost could be an adjuvant to PD1 or checkpoint therapy
  • using Flt3L primed in-situ vaccine (using a Toll like receptor agonist can recruit the dendritic cells to the tumor and then activation of T cell response);  therefore vaccine does not need to be produced ex vivo; months after the vaccine the tumor still in remission
  • versus rituximab, which can target many healthy B cells this in-situ vaccine strategy is very specific for the tumorigenic B cells
  • HoWEVER they did see resistant tumor cells which did not overexpress PD-L1 but they did discover a novel checkpoint (cannot be disclosed at this point)

Please follow on Twitter using the following #hashtags and @pharma_BI

#MCConverge

#AI

#cancertreatment

#immunotherapy

#healthIT

#innovation

#precisionmedicine

#healthcaremodels

#personalizedmedicine

#healthcaredata

And at the following handles:

@pharma_BI

@medcitynews

Please see related articles on Live Coverage of Previous Meetings on this Open Access Journal

LIVE – Real Time – 16th Annual Cancer Research Symposium, Koch Institute, Friday, June 16, 9AM – 5PM, Kresge Auditorium, MIT

Real Time Coverage and eProceedings of Presentations on 11/16 – 11/17, 2016, The 12th Annual Personalized Medicine Conference, HARVARD MEDICAL SCHOOL, Joseph B. Martin Conference Center, 77 Avenue Louis Pasteur, Boston

Tweets Impression Analytics, Re-Tweets, Tweets and Likes by @AVIVA1950 and @pharma_BI for 2018 BioIT, Boston, 5/15 – 5/17, 2018

BIO 2018! June 4-7, 2018 at Boston Convention & Exhibition Center

https://pharmaceuticalintelligence.com/press-coverage/

Read Full Post »

Our Astrophysicist

Larry H. Bernstein, MD, FCAP, Curator

LPBI

 

Ray Kurzweil talks with host Neal deGrasse Tyson, PhD: on invention & immortality

part of the week long event series 7 Days of Genius at 92 Street Y

92 Street Y | 7 Days of Genius
Conversation on stage during the week long event series, held at the historic community center.

featured talk | Ray Kurzweil with host Neil DeGrasse Tyson, PhD — on Invention & Immortality

http://www.kurzweilai.net/images/video-A12.png

 

http://www.kurzweilai.net/images/Entrepreneur-A4.png

 

Inventor, author and futurist Ray Kurzweil is joined by astrophysicist and science communicator Neil deGrasse Tyson, PhD for a discussion of some of the biggest topics of our time. They explore the role of technology in the future, its impact on brain science — and coming innovations in artificial intelligence, energy, life extension and immortality.

Ray Kurzweil has been accurately predicting the future for decades. He explains to Star Talk show host Neil DeGrasse Tyson, PhD how he does it.

Kurzweil also says microscopic robots called nanobots will connect your neocortex to the cloud — the expansion of the human brain that he predicts will happen in the 2030s.

This featured talk is part of a week long series of events called 7 Days of Genius. Presented by the celebrated, historic 92 Street Y cultural arts and community center.

video | 1.
Highlights from the talk with Ray Kurzweil and host Neil deGrasse Tyson, PhD

https://youtu.be/1km56ka9Gnw

 

video | 2.
Highlights from the talk with Ray Kurzweil and host Neil deGrasse Tyson, PhD

https://youtu.be/6BsluRkxs78

 

Entrepreneur | The one tip for success shared by Ray Kurzweil and Neil deGrasse Tyson, PhD

March 9, 2016

Entrepreneur — March 8, 2016 | Catherine Clifford

This is a summary. Read original article in full here

Follow your passion deeply, Ray Kurzweil told an audience at an impressively humorous and entertaining talk hosted by astrophysicist Neil deGrasse Tyson, PhD at the 92 Street Y community center.

The talk with leading, innovative thinkers was part of 92 Street Y’s week long 7 Days of Genius festival. Kurzweil is an inventor, entrepreneur, author and futurist.

In the future, Kurzweil said, there will be a premium on specialized, comprehensive knowledge. If you have passion for art, music or literature — follow that, he says. Kurzweil learned when he was young he had a passion for inventing. “But for some people it’s not clear,” he says. “They should explore many different avenues.”

Money should not be the motivating factor, says Kurzweil, who is something of a romantic. “Don’t do what you think is practical, just because you think that’s a way to make a living. The best way to pursue the future is find an expression you have a passion for,” he says.

Tyson encourages people to seek out learning, visit museums and follow curiosity. Tyson says, “I’m here to make more people passionate, to transform the world for good.”

 

about | 7 Days of Genius at 92 Street Y
Background on the week long event series exploring science, innovation and culture.

92 Street Y |  7 Days of Genius is a multi-platform, week long festival with stage events featuring thought leaders in science, innovation and culture. It explores the concept of genius, and how it transforms lives and cultures.

Events are also hosted globally by partner organizations, and digital broadcast through partners MS • NBC and National Geographic.

Our yearly series of inspiring conversations with experts in politics, technology, knowledge, ethics is focused on the power of genius to change the world for the better.

 

92 Street Y | 7 Days of Genius
Some featured speakers from the series.

1.  Manjul Bhargava, PhD
2.  Esther Dyson
3.  Ray Kurzweil
4.  Martine Rothblatt, PhD
5.  Yancey Strickler
6.  Neil deGrasse Tyson, PhD

 

the festival celebrates Genius Revealed featuring:

1.  special installation at 92 Street Y on remarkable, historic female scientists and inventors throughout history
2.  series on female genius produced with Big Think
3.  20 world events with United Nations Women, exploring how genius can help gender equity
4.  global events celebrating innovative ideas of youth to improve communities with design, entrepreneurship
5.  look for Mental Floss campaign on women geniuses
6.  special programming on MS • NBC, and results of our Ultimate Genius Showdown
7.  see winners of our Global Challenges on design, entrepreneurship, religion

 

video | about 92 Street Y
Background on the historic cultural and community center

watch | video tour

about | 92 Street Y
Landmark community center for culture, arts and conversation.

The historic 92 Street Y is a famous cultural and community center where people from all over connect through culture, arts, entertainment and conversation. For 140 years, we have harnessed the power of arts and ideas to enrich, enlighten and change lives, and the power of community to repair the world. The 92 Street Y is a United States cultural institution in New York, New York at the corner of 92 Street and Lexington Avenue. It’s now a significant landmark center for music, arts, philosophy, celebrity talks and entertainment.

Its full name is the 92 Street Young Men’s and Young Women’s Hebrew Association. Founded in 1874 by German Jewish professionals, 92 Street Y has grown into an organization guided by Jewish principles but serves people of all races and faiths. We harness the power of arts and ideas to enrich, enlighten and change lives, and the power of community.

We enthusiastically reach out to all ages, backgrounds while embracing Jewish values like learning and self-improvement, importance of family, joy of life, and giving back to a wonderfully diverse and growing world.

We curate conversations with the world’s thought leaders — today’s most exceptional thinkers and influential partners for social good — to deepen understanding and engage.

Our performing arts center presents classical, jazz, popular and world music and dance performances. 92 Street Y is a legendary literary destination where the most celebrated writers and readers have gathered since 1939.

We’re a studio, school and workshop where dancers, musicians, jewelry makers, ceramicists, visual artists, poets, playwrights and novelists — professionals and eager amateurs — nourish the human spirit through the arts.

We provide an inspiring, safe and supportive home for families, decades of expertise in parenting, child development, after school sports and classes, special needs programs and summer camps. And offer seniors dozens of activities.

Our fitness center inspires health. 92 Street Y creates meaningful, relevant and joyous experiences for all those who want to connect, finding new ways to bring tradition into dialog with the modern world.

I would encourage anybody as well to watch the Intelligence Square debate video on 92nd Y Street. It is quiet interesting It’s called Don’t trust the promise of Artificial Intelligence. I think both sides of the debate bring interesting arguments.

 

PBS Newshour | Tech’s next feats? Maybe on-demand kidneys, robot sex, cheap solar, lab meat

PBS Newshour | Optimists at Silicon Valley thinktank Singularity University are pushing the frontiers of human progress through innovation and emerging technologies, looking to greater longevity and better health. As part of his series on “Making Sense” of financial news, Paul Solman explores a future of “exponential growth.”

Paul Solman: Admittedly, solar now provides less than 1 percent of U.S. energy needs. But Singularity University’s other cofounder, Ray Kurzweil, whom we interviewed by something called Teleportec, says the public is pointlessly pessimistic.

Ray Kurzweil, Chancellor, Singularity University: And I think the major reason that people are pessimistic is they don’t realize that these technologies are growing exponentially.

For example, solar energy is doubling every two years. It’s now only seven doublings from meeting 100 percent of the world’s energy needs, and we have 10,000 times more sunlight than we need to do that. […]

 

York University | “Google’s Ray Kurzweil receives honorary doctorate” — October 16, 2013

On October 16, 2013 York University conferred an honorary doctorate on Ray Kurzweil, Director of Engineering at Google, in a ceremony on campus. The Lassonde School of Engineering wishes to congratulate Ray Kurzweil on this tremendous honour.

An inventor, author, futurist and a thinker, Ray Kurzweil is most certainly a Renaissance Engineer

Read Full Post »

Unlocking the Microbiome

Larry H. Bernstein, MD, FCAP, Curator

LPBI

3.3.11

3.3.11   Unlocking the Microbiome, Volume 2 (Volume Two: Latest in Genomics Methodologies for Therapeutics: Gene Editing, NGS and BioInformatics, Simulations and the Genome Ontology), Part 2: CRISPR for Gene Editing and DNA Repair

Machine-learning technique uncovers unknown features of multi-drug-resistant pathogen

Relatively simple “unsupervised” learning system reveals important new information to microbiologists
January 29, 201   http://www.kurzweilai.net/machine-learning-technique-uncovers-unknown-features-of-pathogen

http://www.kurzweilai.net/images/Pseudomonas-aeruginosa.jpg

According to the CDC, Pseudomonas aeruginosa is a common cause of healthcare-associated infections, including pneumonia, bloodstream infections, urinary tract infections, and surgical site infections. Some strains of P. aeruginosa have been found to be resistant to nearly all or all antibiotics. (illustration credit: CDC)

A new machine-learning technique can uncover previously unknown features of organisms and their genes in large datasets, according to researchers from the Perelman School of Medicine at the University of Pennsylvania and the Geisel School of Medicine at Dartmouth University.

For example, the technique learned to identify the characteristic gene-expression patterns that appear when a bacterium is exposed in different conditions, such as low oxygen and the presence of antibiotics.

The technique, called “ADAGE” (Analysis using Denoising Autoencoders of Gene Expression), uses a “denoising autoencoder” algorithm, which learns to identify recurring features or patterns in large datasets — without being told what specific features to look for (that is, “unsupervised.”)*

Last year,  Casey Greene, PhD, an assistant professor of Systems Pharmacology and Translational Therapeutics at Penn, and his team published, in an open-access paper in the American Society for Microbiology’s mSystems, the first demonstration of ADAGE in a biological context: an analysis of two gene-expression datasets of breast cancers.

Tracking down gene patterns of a multi-drug-resistant bacterium

The new study, published Jan. 19 in an open-access paper in mSystems, was more ambitious. It applied ADAGE to a dataset of 950 gene-expression arrays publicly available at the time for the multi-drug-resistant bacteriumPseudomonas aeruginosa. This bacterium is a notorious pathogen in the hospital and in individuals with cystic fibrosis and other chronic lung conditions; it’s often difficult to treat due to its high resistance to standard antibiotic therapies.

The data included only the identities of the roughly 5,000 P. aeruginosa genes and their measured expression levels in each published experiment. The goal was to see if this “unsupervised” learning system could uncover important patterns in P. aeruginosa gene expression and clarify how those patterns change when the bacterium’s environment changes — for example, when in the presence of an antibiotic.

Even though the model built with ADAGE was relatively simple — roughly equivalent to a brain with only a few dozen neurons — it had no trouble learning which sets of P. aeruginosa genes tend to work together or in opposition. To the researchers’ surprise, the ADAGE system also detected differences between the main laboratory strain of P. aeruginosa and strains isolated from infected patients. “That turned out to be one of the strongest features of the data,” Greene said.

“We expect that this approach will be particularly useful to microbiologists researching bacterial species that lack a decades-long history of study in the lab,” said Greene. “Microbiologists can use these models to identify where the data agree with their own knowledge and where the data seem to be pointing in a different direction … and to find completely new things in biology that we didn’t even know to look for.”

Support for the research came from the Gordon and Betty Moore Foundation, the William H. Neukom Institute for Computational Science, the National Institutes of Health, and the Cystic Fibrosis Foundation.

* In 2012, Google-sponsored researchers applied a similar method to randomly selected YouTube images; their system learned to recognize major recurring features of those images — including cats of course.


Abstract of ADAGE-Based Integration of Publicly Available Pseudomonas aeruginosa Gene Expression Data with Denoising Autoencoders Illuminates Microbe-Host Interactions

The increasing number of genome-wide assays of gene expression available from public databases presents opportunities for computational methods that facilitate hypothesis generation and biological interpretation of these data. We present an unsupervised machine learning approach, ADAGE (analysis using denoising autoencoders of gene expression), and apply it to the publicly available gene expression data compendium for Pseudomonas aeruginosa. In this approach, the machine-learned ADAGE model contained 50 nodes which we predicted would correspond to gene expression patterns across the gene expression compendium. While no biological knowledge was used during model construction, cooperonic genes had similar weights across nodes, and genes with similar weights across nodes were significantly more likely to share KEGG pathways. By analyzing newly generated and previously published microarray and transcriptome sequencing data, the ADAGE model identified differences between strains, modeled the cellular response to low oxygen, and predicted the involvement of biological processes based on low-level gene expression differences. ADAGE compared favorably with traditional principal component analysis and independent component analysis approaches in its ability to extract validated patterns, and based on our analyses, we propose that these approaches differ in the types of patterns they preferentially identify. We provide the ADAGE model with analysis of all publicly available P. aeruginosa GeneChip experiments and open source code for use with other species and settings. Extraction of consistent patterns across large-scale collections of genomic data using methods like ADAGE provides the opportunity to identify general principles and biologically important patterns in microbial biology. This approach will be particularly useful in less-well-studied microbial species.


Abstract of Unsupervised feature construction and knowledge extraction from genome-wide assays of breast cancer with denoising autoencoders

Big data bring new opportunities for methods that efficiently summarize and automatically extract knowledge from such compendia. While both supervised learning algorithms and unsupervised clustering algorithms have been successfully applied to biological data, they are either dependent on known biology or limited to discerning the most significant signals in the data. Here we present denoising autoencoders (DAs), which employ a data-defined learning objective independent of known biology, as a method to identify and extract complex patterns from genomic data. We evaluate the performance of DAs by applying them to a large collection of breast cancer gene expression data. Results show that DAs successfully construct features that contain both clinical and molecular information. There are features that represent tumor or normal samples, estrogen receptor (ER) status, and molecular subtypes. Features constructed by the autoencoder generalize to an independent dataset collected using a distinct experimental platform. By integrating data from ENCODE for feature interpretation, we discover a feature representing ER status through association with key transcription factors in breast cancer. We also identify a feature highly predictive of patient survival and it is enriched by FOXM1 signaling pathway. The features constructed by DAs are often bimodally distributed with one peak near zero and another near one, which facilitates discretization. In summary, we demonstrate that DAs effectively extract key biological principles from gene expression data and summarize them into constructed features with convenient properties.

Read Full Post »

Future of Big Data for Societal Transformation, Volume 2 (Volume Two: Latest in Genomics Methodologies for Therapeutics: Gene Editing, NGS and BioInformatics, Simulations and the Genome Ontology), Part 1: Next Generation Sequencing (NGS)

Future of Big Data for Societal Transformation

Larry H. Bernstein, MD, FCAP, Curator

LPBI

 

Musk, others commit $1 billion to non-profit AI research company to ‘benefit humanity’

Open-sourcing AI development to prevent an AI superpower takeover
(credit: OpenAI)

Elon Musk and associates announced OpenAI, a non-profit AI research company, on Friday (Dec. 11), committing $1 billion toward their goal to “advance digital intelligence in the way that is most likely to benefit humanity as a whole, unconstrained by a need to generate financial return.”

The funding comes from a group of tech leaders including Musk, Reid Hoffman, Peter Thiel, and Amazon Web Services, but the venture expects to only spend “a tiny fraction of this in the next few years.”

The founders note that it’s hard to predict how much AI could “damage society if built or used incorrectly” or how soon. But the hope is to have a leading research institution that can “prioritize a good outcome for all over its own self-interest … as broadly and evenly distributed as possible.”

Brains trust

OpenAI’s co-chairs are Musk, who is also the principal funder of Future of Life Institute, and Sam Altman, president of  venture-capital seed-accelerator firm Y Combinator, who is also providing funding.

I think the best defense against the misuse of AI is to empower as many people as possible to have AI. If everyone has AI powers, then there’s not any one person or a small set of individuals who can have AI superpower.” — Elon Musk on Medium

The founders say the organization’s patents (if any) “will be shared with the world. We’ll freely collaborate with others across many institutions and expect to work with companies to research and deploy new technologies.”

OpenAI’s research director is machine learning expert Ilya Sutskever, formerly at Google, and its CTO is Greg Brockman, formerly the CTO of Stripe. The group’s other founding members are “world-class research engineers and scientists” Trevor Blackwell, Vicki Cheung, Andrej Karpathy, Durk Kingma, John Schulman, Pamela Vagata, and Wojciech Zaremba. Pieter Abbeel, Yoshua Bengio, Alan Kay, Sergey Levine, and Vishal Sikka are advisors to the group. The company will be based in San Francisco.


If I’m Dr. Evil and I use it, won’t you be empowering me?

“There are a few different thoughts about this. Just like humans protect against Dr. Evil by the fact that most humans are good, and the collective force of humanity can contain the bad elements, we think its far more likely that many, many AIs will work to stop the occasional bad actors than the idea that there is a single AI a billion times more powerful than anything else. If that one thing goes off the rails or if Dr. Evil gets that one thing and there is nothing to counteract it, then we’re really in a bad place.” — Sam Altman in an interview with Steven Levy on Medium.


The announcement follows recent announcements by Facebook to open-source the hardware design of its GPU-based “Big Sur” AI server (used for large-scale machine learning software to identify objects in photos and understand natural language, for example); by Google to open-source its TensorFlow machine-learning software; and by Toyota Corporation to invest $1 billion in a five-year private research effort in artificial intelligence and robotics technologies, jointly with Stanford University and MIT.

To follow OpenAI: @open_ai or info@openai.com    Topics: AI/Robotics | Survival/Defense

Spot on Elon! The threat is and currently the developments are unfortunately pointing exactly in that direction that AI will be controlled via a handful of big and powerful cooperation . None surprisingly none of those subjects are part of the OpenAI movement.

I like the sentiment, AI for all and for the common good, and at one level it seems doable but at another level it seems problematic on the scale of nation states and multinational entities.

If we all have AI systems then it will be those with control of the most energy to run their AI who will have the most influence, and that could be a “Dr. Evil”. It is the sum total of computing power on any given side of a conflict that will determine the outcome, if AI is a significant factor at all.

We could see bigger players looking at strategic questions such as, do they act now, or wait and put more resources into advancing the power of their AI so that they have better odds later, but at the risk of falling to a preemptive attack. Given this sort of thing I don’t see that AI will be a game changer, a leveller, rather it could just fit into the existing arms race type scenarios, at least until one group crosses a singularity threshold and then accelerates away from the pack while holding everyone else back so that they cannot catch up.

Not matter how I look at it I always see the scenarios running in the opposite direction to diversity, toward a singular dominant entity that “roots” all the other AI, sensor and actuator systems and then assimilates them.

How do they plan to stop this? How can one group of AIs have an ethical framework that allows them to “keep down” another group or single AI so that it does not get into a position to dominate them? How will this be any less messy than how the human super-powers have interacted in the last century?

 

I recommend the book “SuperIntelligence” by Nick Bostrom. Most thorough and penetrating. It covers many permutations of the intelligence explosion. The Allegory at the beginning is worth the price alone.

 

Elon, for goodness sake, focus! Get the big battery factory working, get space industry off the ground and America back in the ISS resupply and re-crew business, but enough with the non-profit expenditures already! Keep sinking your capital into non profits like the Hyperlink-a beautiful, high tech version of the old “I just know I can make trains profitable again outside of the northeast” dream and this non-profit AI and you’ll eventually go one financial step too far.

Both for you and for all of us who benefit from your efforts, consider this. At least change your attitude about profit; keep the option open that this AI will bring some profit, even with the open source aspect. This is a great effort, as I see you possibly becoming the “good AI” element that Ray writes about in his first essay, in the essay section on this site. There, Ray is confident that the good people with AI will out-think the bad people with AI and so good AI will prevail.

Read Full Post »

Artificial Intelligence Versus the Scientist: Who Will Win?

Will DARPA Replace the Human Scientist: Not So Fast, My Friend!

Writer, Curator: Stephen J. Williams, Ph.D.

Article ID #168: Artificial Intelligence Versus the Scientist: Who Will Win?. Published on 3/2/2015

WordCloud Image Produced by Adam Tubman

scientistboxingwithcomputer

Last month’s issue of Science article by Jia You “DARPA Sets Out to Automate Research”[1] gave a glimpse of how science could be conducted in the future: without scientists. The article focused on the U.S. Defense Advanced Research Projects Agency (DARPA) program called ‘Big Mechanism”, a $45 million effort to develop computer algorithms which read scientific journal papers with ultimate goal of extracting enough information to design hypotheses and the next set of experiments,

all without human input.

The head of the project, artificial intelligence expert Paul Cohen, says the overall goal is to help scientists cope with the complexity with massive amounts of information. As Paul Cohen stated for the article:

“‘

Just when we need to understand highly connected systems as systems,

our research methods force us to focus on little parts.

                                                                                                                                                                                                               ”

The Big Mechanisms project aims to design computer algorithms to critically read journal articles, much as scientists will, to determine what and how the information contributes to the knowledge base.

As a proof of concept DARPA is attempting to model Ras-mutation driven cancers using previously published literature in three main steps:

  1. Natural Language Processing: Machines read literature on cancer pathways and convert information to computational semantics and meaning

One team is focused on extracting details on experimental procedures, using the mining of certain phraseology to determine the paper’s worth (for example using phrases like ‘we suggest’ or ‘suggests a role in’ might be considered weak versus ‘we prove’ or ‘provide evidence’ might be identified by the program as worthwhile articles to curate). Another team led by a computational linguistics expert will design systems to map the meanings of sentences.

  1. Integrate each piece of knowledge into a computational model to represent the Ras pathway on oncogenesis.
  2. Produce hypotheses and propose experiments based on knowledge base which can be experimentally verified in the laboratory.

The Human no Longer Needed?: Not So Fast, my Friend!

The problems the DARPA research teams are encountering namely:

  • Need for data verification
  • Text mining and curation strategies
  • Incomplete knowledge base (past, current and future)
  • Molecular biology not necessarily “requires casual inference” as other fields do

Verification

Notice this verification step (step 3) requires physical lab work as does all other ‘omics strategies and other computational biology projects. As with high-throughput microarray screens, a verification is needed usually in the form of conducting qPCR or interesting genes are validated in a phenotypical (expression) system. In addition, there has been an ongoing issue surrounding the validity and reproducibility of some research studies and data.

See Importance of Funding Replication Studies: NIH on Credibility of Basic Biomedical Studies

Therefore as DARPA attempts to recreate the Ras pathway from published literature and suggest new pathways/interactions, it will be necessary to experimentally validate certain points (protein interactions or modification events, signaling events) in order to validate their computer model.

Text-Mining and Curation Strategies

The Big Mechanism Project is starting very small; this reflects some of the challenges in scale of this project. Researchers were only given six paragraph long passages and a rudimentary model of the Ras pathway in cancer and then asked to automate a text mining strategy to extract as much useful information. Unfortunately this strategy could be fraught with issues frequently occurred in the biocuration community namely:

Manual or automated curation of scientific literature?

Biocurators, the scientists who painstakingly sort through the voluminous scientific journal to extract and then organize relevant data into accessible databases, have debated whether manual, automated, or a combination of both curation methods [2] achieves the highest accuracy for extracting the information needed to enter in a database. Abigail Cabunoc, a lead developer for Ontario Institute for Cancer Research’s WormBase (a database of nematode genetics and biology) and Lead Developer at Mozilla Science Lab, noted, on her blog, on the lively debate on biocuration methodology at the Seventh International Biocuration Conference (#ISB2014) that the massive amounts of information will require a Herculaneum effort regardless of the methodology.

Although I will have a future post on the advantages/disadvantages and tools/methodologies of manual vs. automated curation, there is a great article on researchinformation.infoExtracting More Information from Scientific Literature” and also see “The Methodology of Curation for Scientific Research Findings” and “Power of Analogy: Curation in Music, Music Critique as a Curation and Curation of Medical Research Findings – A Comparison” for manual curation methodologies and A MOD(ern) perspective on literature curation for a nice workflow paper on the International Society for Biocuration site.

The Big Mechanism team decided on a full automated approach to text-mine their limited literature set for relevant information however was able to extract only 40% of relevant information from these six paragraphs to the given model. Although the investigators were happy with this percentage most biocurators, whether using a manual or automated method to extract information, would consider 40% a low success rate. Biocurators, regardless of method, have reported ability to extract 70-90% of relevant information from the whole literature (for example for Comparative Toxicogenomics Database)[3-5].

Incomplete Knowledge Base

In an earlier posting (actually was a press release for our first e-book) I had discussed the problem with the “data deluge” we are experiencing in scientific literature as well as the plethora of ‘omics experimental data which needs to be curated.

Tackling the problem of scientific and medical information overload

pubmedpapersoveryears

Figure. The number of papers listed in PubMed (disregarding reviews) during ten year periods have steadily increased from 1970.

Analyzing and sharing the vast amounts of scientific knowledge has never been so crucial to innovation in the medical field. The publication rate has steadily increased from the 70’s, with a 50% increase in the number of original research articles published from the 1990’s to the previous decade. This massive amount of biomedical and scientific information has presented the unique problem of an information overload, and the critical need for methodology and expertise to organize, curate, and disseminate this diverse information for scientists and clinicians. Dr. Larry Bernstein, President of Triplex Consulting and previously chief of pathology at New York’s Methodist Hospital, concurs that “the academic pressures to publish, and the breakdown of knowledge into “silos”, has contributed to this knowledge explosion and although the literature is now online and edited, much of this information is out of reach to the very brightest clinicians.”

Traditionally, organization of biomedical information has been the realm of the literature review, but most reviews are performed years after discoveries are made and, given the rapid pace of new discoveries, this is appearing to be an outdated model. In addition, most medical searches are dependent on keywords, hence adding more complexity to the investigator in finding the material they require. Third, medical researchers and professionals are recognizing the need to converse with each other, in real-time, on the impact new discoveries may have on their research and clinical practice.

These issues require a people-based strategy, having expertise in a diverse and cross-integrative number of medical topics to provide the in-depth understanding of the current research and challenges in each field as well as providing a more conceptual-based search platform. To address this need, human intermediaries, known as scientific curators, are needed to narrow down the information and provide critical context and analysis of medical and scientific information in an interactive manner powered by web 2.0 with curators referred to as the “researcher 2.0”. This curation offers better organization and visibility to the critical information useful for the next innovations in academic, clinical, and industrial research by providing these hybrid networks.

Yaneer Bar-Yam of the New England Complex Systems Institute was not confident that using details from past knowledge could produce adequate roadmaps for future experimentation and noted for the article, “ “The expectation that the accumulation of details will tell us what we want to know is not well justified.”

In a recent post I had curated findings from four lung cancer omics studies and presented some graphic on bioinformatic analysis of the novel genetic mutations resulting from these studies (see link below)

Multiple Lung Cancer Genomic Projects Suggest New Targets, Research Directions for

Non-Small Cell Lung Cancer

which showed, that while multiple genetic mutations and related pathway ontologies were well documented in the lung cancer literature there existed many significant genetic mutations and pathways identified in the genomic studies but little literature attributed to these lung cancer-relevant mutations.

KEGGinliteroanalysislungcancer

  This ‘literomics’ analysis reveals a large gap between our knowledge base and the data resulting from large translational ‘omic’ studies.

Different Literature Analyses Approach Yeilding

A ‘literomics’ approach focuses on what we don NOT know about genes, proteins, and their associated pathways while a text-mining machine learning algorithm focuses on building a knowledge base to determine the next line of research or what needs to be measured. Using each approach can give us different perspectives on ‘omics data.

Deriving Casual Inference

Ras is one of the best studied and characterized oncogenes and the mechanisms behind Ras-driven oncogenenis is highly understood.   This, according to computational biologist Larry Hunt of Smart Information Flow Technologies makes Ras a great starting point for the Big Mechanism project. As he states,” Molecular biology is a good place to try (developing a machine learning algorithm) because it’s an area in which common sense plays a minor role”.

Even though some may think the project wouldn’t be able to tackle on other mechanisms which involve epigenetic factors UCLA’s expert in causality Judea Pearl, Ph.D. (head of UCLA Cognitive Systems Lab) feels it is possible for machine learning to bridge this gap. As summarized from his lecture at Microsoft:

“The development of graphical models and the logic of counterfactuals have had a marked effect on the way scientists treat problems involving cause-effect relationships. Practical problems requiring causal information, which long were regarded as either metaphysical or unmanageable can now be solved using elementary mathematics. Moreover, problems that were thought to be purely statistical, are beginning to benefit from analyzing their causal roots.”

According to him first

1) articulate assumptions

2) define research question in counter-inference terms

Then it is possible to design an inference system using calculus that tells the investigator what they need to measure.

To watch a video of Dr. Judea Pearl’s April 2013 lecture at Microsoft Research Machine Learning Summit 2013 (“The Mathematics of Causal Inference: with Reflections on Machine Learning”), click here.

The key for the Big Mechansism Project may me be in correcting for the variables among studies, in essence building a models system which may not rely on fully controlled conditions. Dr. Peter Spirtes from Carnegie Mellon University in Pittsburgh, PA is developing a project called the TETRAD project with two goals: 1) to specify and prove under what conditions it is possible to reliably infer causal relationships from background knowledge and statistical data not obtained under fully controlled conditions 2) develop, analyze, implement, test and apply practical, provably correct computer programs for inferring causal structure under conditions where this is possible.

In summary such projects and algorithms will provide investigators the what, and possibly the how should be measured.

So for now it seems we are still needed.

References

  1. You J: Artificial intelligence. DARPA sets out to automate research. Science 2015, 347(6221):465.
  2. Biocuration 2014: Battle of the New Curation Methods [http://blog.abigailcabunoc.com/biocuration-2014-battle-of-the-new-curation-methods]
  3. Davis AP, Johnson RJ, Lennon-Hopkins K, Sciaky D, Rosenstein MC, Wiegers TC, Mattingly CJ: Targeted journal curation as a method to improve data currency at the Comparative Toxicogenomics Database. Database : the journal of biological databases and curation 2012, 2012:bas051.
  4. Wu CH, Arighi CN, Cohen KB, Hirschman L, Krallinger M, Lu Z, Mattingly C, Valencia A, Wiegers TC, John Wilbur W: BioCreative-2012 virtual issue. Database : the journal of biological databases and curation 2012, 2012:bas049.
  5. Wiegers TC, Davis AP, Mattingly CJ: Collaborative biocuration–text-mining development task for document prioritization for curation. Database : the journal of biological databases and curation 2012, 2012:bas037.

Other posts on this site on include: Artificial Intelligence, Curation Methodology, Philosophy of Science

Inevitability of Curation: Scientific Publishing moves to embrace Open Data, Libraries and Researchers are trying to keep up

A Brief Curation of Proteomics, Metabolomics, and Metabolism

The Methodology of Curation for Scientific Research Findings

Scientific Curation Fostering Expert Networks and Open Innovation: Lessons from Clive Thompson and others

The growing importance of content curation

Data Curation is for Big Data what Data Integration is for Small Data

Cardiovascular Original Research: Cases in Methodology Design for Content Co-Curation The Art of Scientific & Medical Curation

Exploring the Impact of Content Curation on Business Goals in 2013

Power of Analogy: Curation in Music, Music Critique as a Curation and Curation of Medical Research Findings – A Comparison

conceived: NEW Definition for Co-Curation in Medical Research

Reconstructed Science Communication for Open Access Online Scientific Curation

Search Results for ‘artificial intelligence’

 The Simple Pictures Artificial Intelligence Still Can’t Recognize

Data Scientist on a Quest to Turn Computers Into Doctors

Vinod Khosla: “20% doctor included”: speculations & musings of a technology optimist or “Technology will replace 80% of what doctors do”

Where has reason gone?

Read Full Post »

Multiple Lung Cancer Genomic Projects Suggest New Targets, Research Directions for Non-Small Cell Lung Cancer

Curator, Writer: Stephen J. Williams, Ph.D.

UPDATED 08/11/2025

Human Curation vs. AI tools: ChatGPT & Knowledge Graphs [KG] Output: A case study for the following original curation:

  • Multiple Lung Cancer Genomic Projects Suggest New Targets, Research Directions for Non-Small Cell Lung Cancer

https://pharmaceuticalintelligence.com/2014/09/05/multiple-lung-cancer-genomic-projects-suggest-new-targets-research-directions-for-non-small-cell-lung-cancer/

 

This update was performed by the following methods:
A. GPT 5 Text analysis and Reasoning
B. Insertion of Knowledge Graph on topic Curation of Genomic Analysis from Non Small Cell Lung Cancer Studies  from Nodus Labs using InfraNodus software
C. Domain Knowledge Expert evaluation of the Update outcomes
This article has the following Structure:
Part A: Introduction to LLM, Knowledge Graph software InfraNodus, ChatGPT5 and Background Information on curated material for Test Case
Part B: InfraNodus Analysis of manual curation and Knowledge Graph Creation
Part C: Chat GPT 5 Analysis of Manually Curated Material
Part D: Curation entitled Multiple Lung Cancer Genomic Projects Suggest New Targets, Research Directions for Non-Small Cell Lung Cancer originally published on 09/05/2014
Results of Article Update with GPT 5
1. GPT5 alone was not able to understand the goal of the article, namely to determine knowledge gaps in a particular research area involving 5 genomic studies on lung cancer patients
2. GPT5 alone was not able to group concepts or comonalities between biological pathways unless supplied with a manually curated list of KEGG pathways from a list of mutated genes.  However this precluded any effect that fusion proteins had on the analysis and so GPT5 would only concentrate on mutated genes commonly found in literature
3. GPT was not able to access some of the open Access databases like NCBI Gene Ontology database
Results of Article Update with KnowledgeGraph presentation to GPT 5
4. As the Knowledge Graph understood the importance of fusion proteins and transversions, the knowledgegraph augmented the GPT analysis and so enriched the known pathways as well as could correctly identify the less represented pathways in the knowledge graph
5.  This led to the identification of many novel signaling pathways not identified in the original analysis, and was able to perform this task with ease and speed

6. GPT with InfraNodus Analysis was able to propose pertinent questions for future research (the goal of the original curation) such as:

  • How does the interaction between [[EGFR]] mutations and sex-specific gene alterations, including [[RBM10]], influence treatment outcomes in lung adenocarcinoma?
  • How does the intersection of mutational patterns from smoking influence pathway activation in NSCLC, and can identifying these interactions improve targeted therapy development?
Novelty in comparison to Original article published on 09/05/2014
7. it appears that manual curation is necessary to assist in the building of relevant knowledge graphs in the biomedical fields to augment generative AI analysis
8. by itself, generative AI is not optimized for inference of higher concepts from biomedical text, and therefore, at this point, requires the input from human curators developing domain-specific knowledge graphs
9.  The combination of ChatGPT5 and Knowledge graphs of this manually curated biomedical text added a further layer of complexity of gaps of knowledge not seen in the original curations including the need to study noncanonical signaling pathways like WNT and Hedgehog in smoker versus nonsmoker cohorts of lung cancer patients

A Comparison of Manual Expert-Curative and an LLM-based analysis of Knowledge Gaps  in Non Small Lung Cancer Whole Exome Sequencing Studies and a Use Case Example of Chat GPT 5

Part A: Introduction to LLM, Knowledge Graph software InfraNodus, ChatGPT5 and Background Information on curated material for Test Case

The development of Large Language Models (LLMs), together with development of knowledge graphs, have facilitated the ability to analyze text and determine the relationships among the various concepts contained within series of texts.  These concepts and relationships can be visualized, and new insights inferred from these visualizations.  As a result, this type of analysis suggests new directions and lines of research.

 

Alternatively,  these types of visualizations can also reveal gaps in knowledge which should be addressed. A new type of LLM and visualization tools have been developed to understand the gaps in knowledge in biomedical text.

Nodus Labs InfrNodus AI Knowledge Graph Software Tools Allow Text Relationship Visualization and Integrated AI Functionality

 

Infranodus makes knowlegde graphs from text and then is able to visualize the relationships between concepts (or nodes).  In doing so, the tool also highlights the various knowledge gaps (or large differences between nodes) which can be used to investigate new hypotheses and research directions of previously univestigated relationships between concepts.  This generates new research questions, in which these gaps can be used as prompts in the software’s integrated AI tool.  The AI tool, much like a GPT, returns recommendations for research to be conducted in the area.

https://infranodus.com/

In addition, the InfraNodus software can detect if text is too biased on a particular concept or conclusion, and using a GPT3 or GPT4, can determine if the nodes are too dispersed and will recommend which gaps should be focused on.

The software can upload any biomedical text in various formats

A full demonstration is on their website but a good summary is found on their Youtube site at

https://www.youtube.com/watch?v=wCEhiIJsmrg

A couple of use cases include

 

 

Previously we had manually curated and analyzed the knowledge gaps from a series of publications on whole exome sequencing  of biopsied tumors from cohorts of non small lung cancer patients. This curation (from 2016) is seen in the lower half of this updated link below and I separated with a bar and highlighted in Yellow as Text for AI Analysis.

https://pharmaceuticalintelligence.com/2014/09/05/multiple-lung-cancer-genomic-projects-suggest-new-targets-research-directions-for-non-small-cell-lung-cancer/

A literature analysis of the driver mutations found in five NSLC exome sequencing projects:

  1. Comprehensive genomic characterization of squamous cell lung cancersNature 2012, 489(7417):519-525.
  2. A genomics-based classification of human lung tumorsScience translational medicine 2013, 5(209):209ra153.
  3. Govindan R, Ding L, Griffith M, Subramanian J, Dees ND, Kanchi KL, Maher CA, Fulton R, Fulton L, Wallis J et alGenomic landscape of non-small cell lung cancer in smokers and never-smokersCell 2012, 150(6):1121-1134.
  4. Imielinski M, Berger AH, Hammerman PS, Hernandez B, Pugh TJ, Hodis E, Cho J, Suh J, Capelletti M, Sivachenko A et alMapping the hallmarks of lung adenocarcinoma with massively parallel sequencingCell 2012, 150(6):1107-1120.
  5. Peifer M, Fernandez-Cuesta L, Sos ML, George J, Seidel D, Kasper LH, Plenker D, Leenders F, Sun R, Zander T et alIntegrative genome analyses identify key somatic driver mutations of small-cell lung cancerNature genetics 2012, 44(10):1104-1110.

 

were performed.

The purpose of this analysis was to uncover biological functions related to the sets of mutated genes with limited research publications in the area of  non small cell lung cancer.  The identification of such biological functions would represent a gap in knowledge in this disease.  In addition, this analysis attempted to find new lines of research or potential new biotargets to investigate for lung cancer therapy.

 

 

 

However this manual method is time consuming and may miss relationships not defined in a GO ontology or gene knowledgebases.

Therefore we turned to an AI-driven approach:

  1. Using InfraNodus ability to develop a knowledge graph based on our curation and determine if the AI platform could infer knowledge gaps
  2. Utilize Chat GPT5 to analyze the same curated set to determine if OpenAI analysis would lead to the similar analysis from curated material
  3. Determine if combining a knowledge graph within GPT would lead to a higher level of analysis

See below (Part D) of this update for the curated studies which were included in this analysis and the text which was entered into both InfraNodus and Chat GPT5. 

As a summary, it seems that manual curation is necessary to assist in the building of relevant knowledge graphs in the biomedical fields to augment generative AI analysis.  In addition, it appears that , by itself, generative AI is not optimized for inference of higher concepts from biomedical text, and therefore, at this point, requires the input from human curators developing domain-specific knowledge graphs.

 

Part B. InfraNodus Analysis of manual curation and Knowledge Graph Creation

Methods: 

Text of the curation was copied and directly pasted into the text analysis module of InfraNodus.  There was no editing of words however genes in the curation were linked to their GeneCard entry. GeneCards is a database run by the Weizmann Institute.  InfraNodus utilizes a combination of LLMs and its own GraphRAG system to provide insights from text analysis. While it leverages various models, including those from OpenAI and Anthropic, it’s not limited to a single LLM. Instead, InfraNodus integrates these models within its GraphRAG framework, which enhances their capabilities by adding a relational understanding of the context through a knowledge graph.

InfraNodus then autogenerates a knowledge graph and returns entities and relationships between entities.  InfraNodus offers the opportunity to modify the knowledge graph however for this analysis we used the first graph InfraNodus generated.  Inspection of this graph (as shown below) was deemed reasonable.

 

Results

The knowledge graph of the input text is shown below:

InfraNodus generated Knowledge Graph of 5 WES Non Smal Cell Lung Cancer studies involving smokers and non smokers

 

Four main concepts were returned: tumors, genes, literature, and mutations.

A snapshot of the Analysis window is given below.  It should be noted that InfraNodus felt there needed to be more connections between Pathway and Mutational Patterns.

An InfraNodus reposrt with Knowlege Graph on Whole Exome Sequencing studies in NSCLC to determine mutational spectrum in smokers versus non smokers

Auto generated summary report

Context name: text_250808T0144

Created on: aug 7, 2025 9:47 pm

Last updated on: aug 7, 2025 10:10 pm

Main concepts:

[[tumors]], analysis, [[mutations]], identify, [[lung]], [[genes]]

Main topics:

  1. Tumor Genomics: [[tumors]] [[lung]] reveal
  2. Genetic Alterations: identify [[genes]] study
  3. Pathway Analysis: analysis pathway literature
  4. Mutation Patterns: [[mutations]] [[egfr]] [[rbm10]]

Structural gap (topics to connect):

  1. Pathway Analysis: analysis pathway
  2. Smoking Influence: mutational [[smoking]]

Topical connectors:

alk clinical [[egfr]] mutational pathway [[paper]] found key literature study [[genomic]] reveal [[transversion]]

 

Top relations / ngrams:

1) [[lung]] [[tumors]]

2) alk fusion

3) link function

4) eml alk

5) function [[gene_ontology]]

Modulary: 0.47

Relations:

InfraNodus identified 744 relations between entities (nodes)

A list of some of the more frequent are given here:

source target occurrences weight betweenness
[[lung]] [[tumors]] 8 24 0.4676
analysis pathway 5 12 0.2291
significantly [[genes]] 5 9 0.1074
significantly [[mutated]] 4 12 0.0281
[[mutated]] [[genes]] 4 12 0.0847
[[transversion]] high 3 12 0.0329
[[smoking]] history 3 10 0.0352
study identify 3 9 0.2051
mutational pattern 3 9 0.0921
[[rbm10]] [[mutations]] 3 8 0.1776
literature analysis 3 7 0.2218
[[egfr]] [[mutations]] 3 7 0.2139
[[transversion]] group 3 7 0.0259
enriched cohort 3 6 0.0219
[[whole_exome_sequencing]] [[tumors]] 3 6 0.3485
identify [[genes]] 3 6 0.2268
including analysis 3 5 0.1985
alteration [[genes]] 3 4 0.1298
[[tumors]] analysis 3 4 0.5192
alk fusion 2 15 0.0671
link function 2 14 0.0269
function [[gene_ontology]] 2 13 0.0054

 

Notice how the betweenness or importance of connection of disparate concepts vary but are high between concepts like tumors and analysis, or lung and tumor, however many important linked concepts like alk and fusion may have low betweenness but are mentioned frequently and have a much higher weight or closeness to each other. Gene-mutations-transversions-smoking seem to have a high correspondence to each other

For a full list please see the file

InfraNodus relations five WES studies on NSCLC and smoking

sjwilliams_text_250808T0144_nodedataon NSCLCWESstudies

Generated the following responses using AI to bridge the gap between

  1. Mutation Patterns: [[mutations]][[egfr]][[rbm10]]and
  2. Genetic Alterations: identify[[genes]]study:The recent comprehensive studies on lung adenocarcinoma have significantly advanced our understanding of the genetic landscape by identifying key mutations and their intricate interactions. Notably, EGFR and RBM10 exhibit distinct mutational patterns, with RBM10 inactivations being notably enriched in male cohorts. This gender-linked enrichment underscores a potential differential oncogenic pathway involving ERBB2 and RB1 alterations.Moreover, these projects emphasize the quest to map significant gene alterations within lung adenocarcinoma. The identification of such genes not only corroborates prior reports but also expands upon them by highlighting new connections between mutation signatures and clinical factors like smoking history. These findings are crucial as they can inform future therapeutic targeting strategies, ensuring that personalized treatment approaches consider both gender-specific genomic enrichments and mutation-driven tumorigenesis pathways elucidated through rigorous analyses.elaborate

 

<MainTopics>

  1. Tumor Genomics: [[tumors]], [[lung]], reveal, base, [[genomic]], [[data]], classification, subtype,
  2. Genetic Alterations: identify, [[genes]], study, alteration, reported, significantly, [[lung_adenocarcinoma]], key,
  3. Pathway Analysis: analysis, pathway, literature, including, activation, target, [[nsclc]], [[patients]],
  4. Mutation Patterns: [[mutations]], [[egfr]], [[rbm10]], male, enriched, cohort, [[erbb2]], noted,
  5. Smoking Influence: mutational, [[smoking]], [[transversion]], pattern, high, notably, [[smokers]], show,
  6. Research Strategies: found, [[paper]], [[omics]], unique, profile, result, multus, characterization,
  7. Targeted Therapy: alk, [[lung_cancer]], fusion, [[therapy]], ro, eml, include, targeted,
  8. Mutation Signatures: clinical, [[tumor]], approach, mutation, sample, signature, gender, specific,
  9. Gene Function: link, [[gene]], note, function, author, [[gene_ontology]],

</MainTopics>

 

<TopicalGap>:

  1. Genetic Alterations: identify [[genes]]
  2. Mutation Patterns: [[mutations]] [[egfr]]

</TopicalGap>

 

 

questions generated using AI to help you explore “alk, clinical, [[egfr]], mutational, pathway, [[paper]], found, key, literature, study, [[genomic]], reveal, [[transversion]]…”:How do mutational patterns, specifically EGFR mutations and transversions related to smoking history, influence the effectiveness of targeted therapies in NSCLC patients?elaborate

 

ideas generated using AI to help you explore “alk, clinical, [[egfr]], mutational, pathway, [[paper]], found, key, literature, study, [[genomic]], reveal, [[transversion]]…”:Develop a predictive model that utilizes genomic data and smoking history to forecast patient response to targeted therapies. This model would identify key mutational signatures linked to EGFR and other genes, highlighting the impact of smoking-induced transversions on drug efficacy.elaborate

 

Project Notes

 

The recent comprehensive studies on lung adenocarcinoma have significantly advanced our understanding of the genetic landscape by identifying key mutations and their intricate interactions. Notably, EGFR and RBM10 exhibit distinct mutational patterns, with RBM10 inactivations being notably enriched in male cohorts. This gender-linked enrichment underscores a potential differential oncogenic pathway involving ERBB2 and RB1 alterations.

 

Moreover, these projects emphasize the quest to map significant gene alterations within lung adenocarcinoma. The identification of such genes not only corroborates prior reports but also expands upon them by highlighting new connections between mutation signatures and clinical factors like smoking history. These findings are crucial as they can inform future therapeutic targeting strategies, ensuring that personalized treatment approaches consider both gender-specific genomic enrichments and mutation-driven tumorigenesis pathways elucidated through rigorous analyses.”

 

<ConceptualGateways>

alk

clinical

[[egfr]]

mutational

pathway

[[paper]]

found

key

literature

study

[[genomic]]

reveal

[[transversion]]

</ConceptualGateways>

 

 

How do mutational patterns, specifically EGFR mutations and transversions related to smoking history, influence the effectiveness of targeted therapies in NSCLC patients?

 

The report from the NCI Bulletin outlines significant advancements in understanding lung cancer through genome sequencing projects. These studies have revealed a plethora of genetic and epigenetic alterations across various forms of lung tumors, including adenocarcinomas, squamous cell carcinomas, and small cell lung cancers. Notably, some identified alterations could be targeted by existing therapies, providing potential new avenues for treatment.Dr. Meyerson emphasizes the complexity of these genetic changes, highlighting that distinct mechanisms inactivating genes can vary between tumors. The report also notes gaps in knowledge regarding non-coding DNA alterations, which comprise a major part of the human genome.Key findings include:1. Comprehensive genomic analyses revealing unique driver mutations in lung adenocarcinoma, such as those affecting MET and ERBB2, alongside significant mutations in known cancer drivers like TP53 and KRAS.2. A classification system based on genomic data enabling more accurate patient stratification—achieving a 75% classification rate of lung cancer subtypes.3. Smoking history is shown to influence mutational patterns significantly, with smokers exhibiting a higher incidence of point mutations compared to never-smokers.Moreover, the integration of genomic data and pathway analysis highlighted recurrent mutations across various pathways related to tumorigenesis, suggesting new therapeutic targets and underscoring the importance of personalized medicine approaches that factor in gender-specific mutation distributions.This synthesis of findings not only corroborates earlier studies but also extends our understanding of the interplay between genomic alterations, smoking habits, and clinical outcomes in lung cancer. Future research is needed to explore the implications of these findings further and to develop targeted therapies that leverage the identified genetic alterations.

 

  1. Tumor Genomics
  2. Genetic Alterations
  3. Pathway Analysis
  4. Mutation Patterns

 

Main Ideas Summary from Infranodus:

 

The report from the NCI Bulletin highlights significant developments in lung cancer genomics through five genome sequencing projects. These studies have uncovered various genetic and epigenetic alterations across multiple forms of lung cancer, including adenocarcinomas, squamous cell cancers, and small cell lung cancers. Notably, these alterations may correspond to existing drugs, offering new avenues for targeted therapy.

Key findings include the identification of unique mutations in genes like NF1, MET, and ERBB2, with particular attention to smoking’s role in mutational patterns. The analysis shows that smokers exhibit a higher incidence of specific point mutations, marking a distinct divergence in mutation profiles compared to never-smokers.

The research also emphasizes the need for continued exploration of the data to link mutation signatures with clinical outcomes, particularly regarding gender-specific differences in mutation prevalence. Emerging insights suggest that pathways involving MYC and alterations in splicing factors could be vital targets for future therapeutic strategies.

Overall, the studies collectively underscore an urgent need for integrating genomic data into clinical practice, which can lead to more personalized treatment approaches for lung cancer patients by considering both the genetic landscape of tumors and their histories, such as smoking exposure. The ongoing efforts to map gene alterations within lung adenocarcinoma not only reinforce existing knowledge but also reveal new potential therapeutic targets that can significantly impact patient outcomes.

 

Questions to Explore from InfraNodus:

 

How does the interaction between [[EGFR]] mutations and sex-specific gene alterations, including [[RBM10]], influence treatment outcomes in lung adenocarcinoma?

 

 

How does the intersection of mutational patterns from smoking influence pathway activation in NSCLC, and can identifying these interactions improve targeted therapy development?

 

Part C: Chat GPT 5 Analysis of Manually Curated Material

Now let us see how Chat GPT5 analyzed the same material.

Methods

 

Nice — I’ll work from the KEGG-pathway list shown in your text (the table of pathways with PubMed counts) and call out major KEGG pathways that are not represented there.

Important caveat: “Not represented in the table” below means not listed in the KEGG-pathway hits shown in your pasted literature analysis, not that these pathways never play a role in lung cancer. Absence from that table can reflect lack of reporting in the specific WES/WGS studies or the PubMed keyword search used — it’s not proof the pathway is uninvolved biologically.

The above Important caveat was very important because it showed that ChatGPT5 could understand what my experimental goals were, or at least understood the problem I was presenting to ChatGPT5!

Major KEGG pathways missing from the provided KEGG list

 

(However Note that GPT did NOT take into account the mutational spectrum of the studies or the gene lists supplied in the curation)

 

These are commonly cited cancer-related KEGG pathways that do not appear in the table you gave:

  • Wnt signaling pathway
  • Notch signaling pathway
  • Hedgehog signaling pathway
  • TGF-beta signaling pathway
  • Hippo signaling pathway
  • p53 signaling pathway(note: TP53 is mutated in many tumors, so p53 biology is implicated in the text, but “p53 signaling” does not appear in that KEGG hits table)
  • Apoptosis (intrinsic/extrinsic apoptotic pathways)
  • Autophagy
  • NF-kappa B signaling pathway
  • Ubiquitin-mediated proteolysis / proteasome
  • DNA repair pathways (several separate KEGG entries):
    • Homologous recombination
    • Mismatch repair
    • Nucleotide excision repair
    • Base excision repair
    • Fanconi anemia pathway
  • Cellular metabolism pathways commonly in cancer:
    • Glycolysis / Gluconeogenesis
    • Citrate cycle (TCA cycle) / Oxidative phosphorylation
    • Fatty acid metabolism / beta-oxidation
    • Pentose phosphate pathway
  • Immune-signaling / antigen processing:
    • T cell receptor signaling pathway
    • B cell receptor signaling pathway
    • Antigen processing and presentation
  • Autocrine/paracrine inflammation/chemokine signaling beyond the one low-count entry (general cytokine networks other than the single “cytokine–cytokine receptor interaction” hit)
  • Signal transduction modules sometimes separate in KEGG:g. Calcium signaling was listed in your table, but other modules like cGMP–PKG signaling or PKA/PKC broader signaling networks are not.

 

Quick summary

The KEGG table in your text highlights many cancer and receptor-tyrosine/MAPK/PI3K related pathways (cell cycle, MAPK, ErbB, JAK-STAT, focal adhesion, etc.). What’s not listed are numerous developmental morphogen pathways (Wnt/Notch/Hedgehog/TGF-β/Hippo), many DNA-repair and core metabolic pathways, canonical apoptosis/autophagy/ubiquitin pathways, and broad adaptive-immune signaling pathways.

If you want, I can:

  • (A) cross-check this against the full KEGG pathway list and produce a definitive “KEGG universe → present vs absent in your table” spreadsheet, or
  • (B) prioritize which of the missing pathways you’d most want investigated in lung tumor sequencing data and suggest genes to look for.

Which follow-up would you prefer?

 

Now taking out the KEGG list I gave GPT in order to force GPT to do a similar but unbiased analysis

 

From the studies you’ve shared, we can identify which KEGG pathways are under-represented (i.e., appear less frequently altered than expected) in the altered genes in lung cancer.

To do this, we note that the researchers used KEGG and other pathway analysis tools (PathScan, literature curation, proteomics) to see which biological pathways the mutated genes mapped to.
While many well-known cancer pathways were found to be activated (MAPK, PI3K-AKT, mTOR, cell cycle, oxidative stress, etc.), a few key pathways were reported as unexpectedly low in mutation frequency despite their known role in cancer biology.

KEGG Pathways Under-Represented in Altered Genes in Lung Cancer

From the descriptions in Papers 1–5, these are the main pathways that appear under-represented:

  1. MAPK signaling pathway
    • Normally, MAPK is a major cancer driver pathway. In this dataset, MAPK activity was high in proteomics, but many tumors lacked mutations in MAPK pathway genes.
    • This means MAPK activation is often non-mutational (possibly via upstream signaling, epigenetic changes, or ligand activation), so in mutation-based KEGG enrichment it is under-represented.
  2. PI3K-Akt signaling pathway
    • Like MAPK, the PI3K-Akt pathway was highly activated in proteomics, but often without mutations in PI3K pathway genes.
    • Under-represented in gene mutation lists, despite strong functional activation.
  3. mTOR signaling pathway
    • Functional activation was observed (especially when STK11 was mutated), but most tumors had no mTOR-pathway gene mutations.
    • Under-represented when looking only at altered genes.
  4. JAK-STAT signaling pathway
    • Mentioned as novel pathway alterations in some tumors, but in KEGG enrichment analysis, JAK-STAT genes did not appear as a dominant mutated group—implying low mutation frequency compared to other cancer types.
  5. Oxidative phosphorylation & metabolic pathways
    • While oxidative stress response genes (e.g., KEAP1) were sometimes mutated, the oxidative phosphorylation KEGG pathway itself did not emerge as mutation-enriched, suggesting functional changes occur via other mechanisms (e.g., hypoxia signaling).
  6. Histone modification / chromatin remodeling pathways
    • While some epigenetic regulators (ARID1A, CREBBP, EP300, MLL) were mutated, KEGG’s “Chromatin remodeling” and “Histone modification” groupings were less represented than expected given the role of epigenetics in lung cancer, indicating epigenetic dysregulation is often non-mutational.

Why They Are Under-Represented

  • Proteomics revealed activation without DNA mutation (e.g., ligand-driven activation, phosphorylation cascades).
  • Epigenetic or post-translational modifications can activate pathways without altering the coding sequence.
  • Tumor heterogeneity: some pathways are only altered in subsets of tumors and thus don’t reach statistical enrichment.
  • Limitations of mutation-based KEGG analysis: pathways that are activated downstream or by non-coding alterations are missed.

Summary Table

Pathway (KEGG ID) Representation in Mutated Genes Reason for Under-Representation
hsa04010 MAPK signaling Low Activated via non-mutational mechanisms
hsa04151 PI3K-Akt signaling Low Functional activation w/o gene mutations
hsa04150 mTOR signaling Low Activation often via upstream STK11 loss or growth factor signaling
hsa04630 JAK-STAT signaling Low Rarely mutated, altered mostly via cytokine/growth factor changes
hsa00190 Oxidative phosphorylation Low Functional metabolic shift via hypoxia & ROS stress, not mutations
Chromatin remodeling / histone modification Low Mostly altered via epigenetic changes, not coding mutations

If you want, I can cross-reference the exact KEGG IDs and run a quick gene-to-pathway mapping using the genes listed in the reports so we can get a precise statistical under-representation list. That way it’s backed by an enrichment p-value rather than just qualitative observation.

 

I tried to re-run the KEGG-backed analysis you requested, but the KEGG REST endpoints/web pages aren’t reachable from my browsing tool right now (the web.run calls to KEGG returned an internal/safety error). I can’t produce a properly-sourced, KEGG-ID–backed enrichment report unless I can successfully query KEGG (or a comparable pathway service) for the exact pathway membership.

Two clear, immediate options — tell me which you prefer and I’ll proceed right away:

Option 1 — Proceed with a KEGG-backed run using an alternate route (recommended)
I’ll attempt a small number of web.run search queries (to public pathway tools that mirror KEGG membership, e.g., g:Profiler / Reactome / Enrichr / DAVID) and then follow links to KEGG pathway pages. This often works even when direct KEGG REST calls fail. It will take a few web.run calls and I’ll cite the pages used. (If you want me to prefer a particular service — e.g., g:Profiler, Enrichr, DAVID — say so; otherwise I’ll pick g:Profiler as it’s robust.)

Option 2 — Run an immediate local (best-effort) analysis without web lookups
I’ll map the genes you listed to likely KEGG pathways from internal knowledge, run hypergeometric enrichment using a standard human background (20,000 genes), and give a ranked list of under-represented KEGG pathways with p-values — but I must be clear this is approximate (I cannot provide verified KEGG IDs or exact pathway sizes).

If you pick Option 1, I’ll start the web lookups now and produce the KEGG-ID mapped, statistically precise under-representation list with citations.
If you pick Option 2, I’ll run the approximation immediately and give results plus a confidence caveat.

Which do you want?

Sources

Ask ChatGPT

You said:

Use Option 1 to Proceed with a KEGG-backed run using an alternate route

ChatGPT said:

Thought for 37s

Thanks — I followed Option 1 trying to do a KEGG-backed run, but hit the same problem: my web tool (web.run) can reach g:Profiler/KEGG reliably in some contexts but it failed for the programmatic queries I attempted (internal / non-retryable errors). Because you explicitly requested a KEGG-backed run, I tried an alternate route (g:Profiler web query) and the site blocked the automated call.

Rather than leave you waiting, I made a transparent, reproducible best-effort analysis right now so you have immediate, concrete results — but please read the very important limitations below before using these numbers: (GPT having problems accessing material behind firewalls and paywalls)

What I did (short)

  1. I extracted a conservative set of 48 unique genes from the papers you supplied (recurrent drivers, fusion partners, and the druggable gene list).
    Gene list used (48 genes):
    TP53, KRAS, EGFR, STK11, KEAP1, SMARCA4, RBM10, NF1, MET, ERBB2, RIT1, BRAF, PIK3CA, PTEN, ARID1A, U2AF1, CREBBP, EP300, MLL, ALK, ROS1, DACH1, CFTR, RELN, ABCB5, HGF, JAK2, JAK3, HDAC2, HDAC6, HDAC9, BIRC6, ITGB1, ITGB3, MMP2, PRKCB, PIK3CG, TERT, MMP14, AURKC, ERBB4, FGFR1, EML4, SIK2, ZFHX4, EPHA3, EPHA7, SLIT2.
  2. Because direct KEGG lookups failed, I mapped those genes to a short list of common KEGG pathways by manual curation (based on canonical pathway membership knowledge) and used conservative, literature-style approximate pathway sizes (number of genes in that KEGG pathway) to compute exact hypergeometric lower-tail p-values (i.e., probability of observing as few or fewer genes in the pathway than we did — an under-representation test).
  3. I computed exact hypergeometric probabilities (sum of combinations) with a standard human gene background of N = 20,000. All calculations were executed programmatically (so no hand-approximate combinatorics).

Important limitations (please read)

  • This is NOT a true KEGG-backed run. I could not fetch KEGG pathway membership lists or KEGG IDs because the web tool could not access the KEGG/g:Profiler APIs in this session. (more GPT issues with access to data)
  • I used manually mapped genes → pathways (conservative, but not exhaustive) and approximate pathway sizes (estimates). Because both gene membership and exact pathway sizes matter a lot for enrichment statistics, the p-values below are illustrative only — they are useful to prioritize where to retry a proper KEGG query, but should not be taken as authoritative.
  • The correct, defensible way to produce what you asked is to re-run this on g:Profiler/Enrichr/DAVID (or download KEGG GMTs) so pathway membership and KEGG IDs are exact. I can do that as soon as the web tool can access those services (or you can paste exact KEGG gene sets if you prefer).

Results — immediate best-effort under-representation test

Analysis parameters

  • Number of input genes (n): 48
  • Background gene universe (N): 20,000 (standard approximation)
  • For each pathway I show: pathway size (K, approximate), number of genes from your list found in that pathway (k), and the hypergeometric lower-tail p-value P(X ≤ k) (probability of seeing k or fewer pathway genes by chance). Lower p-values (e.g., < 0.05) would suggest significant under-representation.
Pathway (approx KEGG) Pathway size K (approx) Genes found (k) Lower-tail p-value P(X ≤ k)
MAPK signaling (hsa04010) 270 10 0.9999999999678
PI3K-Akt signaling (hsa04151) 350 7 0.9999983354944
mTOR signaling (hsa04150) 120 3 0.9998044249093
JAK-STAT signaling (hsa04630) 150 3 0.9995423622408
Oxidative phosphorylation (hsa00190) 130 0 0.7309662417384
Spliceosome (hsa03040) 120 2 0.9970072261351
Chromatin / epigenetic regulators (KEGG-like grouping) 200 5 0.9999919751398
ECM / focal adhesion (hsa04510 / related) 150 4 0.9999706923476
Cell cycle (hsa04110) 200 3 0.9986613990288

Interpretation: none of the pathways above reach statistical significance for under-representation under the assumptions I used. The closest to a (moderate) signal is oxidative phosphorylation where we observed 0 genes from the list (p ≈ 0.73) — but that p-value is still far from being significant (p < 0.05 would be needed). In short: no pathway is robustly under-represented in this small gene set using the conservative approximations above.

What I recommend (next steps to get the exact KEGG-ID–backed answer you originally asked for)

  1. Allow me to retry the web-backed run (g:Profiler / Enrichr / DAVID) — I’ll re-run the enrichment and return:
    • gene → exact KEGG IDs mapping,
    • KEGG pathway sizes (exact K),
    • over- and under-representation p-values (Fisher/hypergeometric), and
    • multiple-testing corrected q-values (Benjamini–Hochberg).
      (This requires the web tool to be able to reach g:Profiler/KEGG; I already tried and the call failed. If you want, I’ll attempt again immediately.)
  2. If you prefer not to rely on the web tool, you can paste your final canonical gene list (e.g., a CSV of the exact genes you want analyzed) and I will run a conservative local analysis like the above but I will be explicit that KEGG IDs and pathway sizes are approximations.
  3. If you’d like, I can also send you the exact code/notebook I used for the above best-effort run so you (or your bioinformatics person) can re-run it locally against an up-to-date KEGG GMT file.

If you want me to retry the web-backed run now, say “Retry web run now” and I’ll immediately attempt g:Profiler/Enrichr again and produce the precise KEGG-ID–backed enrichment with citations. If you prefer I proceed without web access (i.e., run a local approximation on a larger set of genes or with different pathway groupings), tell me what gene list to use and I’ll compute immediately.

Top of Form

 

Bottom of Form

 

Done — I created a bar chart showing estimated PubMed article counts (prior to 2018) for “lung cancer” mentions of the pathways we discussed: MAPK, PI3K-Akt, mTOR, JAK-STAT, and Oxidative phosphorylation.

Important caveats (please read)

  • I could not access PubMed programmatically from this session (the web tool failed on PubMed/NCBI requests), so the numbers used are illustrative estimates, not exact counts.
  • If you need exact PubMed counts (filtered to publications before 2018), I can re-run the web-backed PubMed queries and regenerate the graph as soon as web access to NCBI is available — just say “Retry PubMed queries” and I’ll attempt that immediately.
  • The image file is available to download: Download the bar chart image.

If you want, I can:

  • Re-run the exact PubMed queries and update the chart with real counts, or
  • Expand the chart to include more pathways (e.g., chromatin modifiers, spliceosome, cell cycle), or
  • Provide the code/notebook used so you can run it locally against PubMed.

 

Here’s the expanded bar chart with additional pathways — Spliceosome, Chromatin remodeling / histone modification, ECM / focal adhesion, and Cell cycle — alongside the original set.

You can download the image here: Download expanded chart.

Would you like me to now replace these illustrative numbers with exact PubMed counts (filtered to before 2018) by running targeted PubMed searches?

 

 

Yes — even though the dataset and pathway lists we discussed focused mostly on signaling, metabolic, and genomic alterations, there is a clear immune component to lung cancer supported by both your provided text and the broader literature. (Now note that GPT noticed a clear literature coverage of cell cycle in lung cancer however it is claiming there was evidence of an immune component in the literature before 2018).

Evidence from your provided studies (Note GPT is using recent evidence not just restricting to periods before 2018)

  1. Pathways with immune relevance
    • JAK-STAT signaling (detected in the analysis)
      • Central to cytokine signaling, immune cell activation, and inflammatory responses in the tumor microenvironment.
      • Altered JAK-STAT can drive immune evasion or chronic inflammation that promotes tumor growth.
    • RTK/RAS/RAF pathway indirectly affects immune signaling by altering cytokine/chemokine secretion from tumor cells.
    • ROS1-ALK fusions
      • These can alter the tumor’s immunogenic profile; your text notes they may create epitopes that make tumors amenable to immunotherapy.
    • Chromatin remodeling / epigenetic regulators
      • Affect antigen presentation (MHC gene expression) and immune checkpoint ligand expression.
  1. Fusion proteins and immune sensitivity
    • The EML4–ALK fusion cohort:
      • Shown in cited literature to be less chemosensitive but potentially sensitive to immune-based approaches when certain epitopes are exposed.
      • Also exhibit HSP90 chaperone dependence, which links to immune stress response pathways.
  1. Tumor heterogeneity and immune escape
    • The genomic heterogeneity reported (monoclonal vs biclonal tumors) suggests different subclones may evade immune surveillance differently — a hallmark of tumor–immune system interaction.

From broader lung cancer research

Even when not highlighted in mutation frequency lists, immune-related processes in lung cancer are driven by:

  • Tumor-infiltrating lymphocytes (TILs) and their exhaustion profiles.
  • Immune checkpoints (PD-1/PD-L1, CTLA-4) — a major therapeutic target in NSCLC.
  • Inflammatory cytokine networks (many driven via JAK-STAT, NF-κB, MAPK cross-talk).
  • Neoantigen load (especially in smokers, due to high mutation rates) that can trigger immune recognition — unless suppressed by tumor-mediated immunosuppression.

Conclusion

Yes — there is an immune component to lung cancer.
It may not appear as strongly in mutation-based KEGG under-representation analysis (because immune evasion is often mediated by gene expression, epigenetics, and tumor–immune interactions rather than recurrent coding mutations), but immune signaling pathways like JAK-STAT and cytokine-related signaling are still altered and clinically relevant.

 

(Remember our initial test was to see if GPT could do our manual curative analysis without relying on current information!)

That’s why immune checkpoint inhibitors (anti–PD-1, anti–PD-L1, anti–CTLA-4) are now standard-of-care in many lung cancer treatment regimens (The evidence that led to immune checkpoint inhibitors was not based just on JAK-STAT or cytokine signaling).

 

Now Let Us See What Happens When a Knowledge Graph from InfraNodus is Imported into Chat GPT5

 

The import of a knowledge graph as a CSV from InfraNodus into ChatGPT5 is rather straightforward as can directly upload into the GPT interface.

 

ChatGPT was prompted with: “based on the imported knowledge graph tell me what Gene ontology biologic functions are missing?

ChatGPT said:

Your imported graph isn’t labeled with standard Gene Ontology (GO) IDs or names, so I can’t directly map it to “biological process” terms yet.

The next step is to align your node name column with the official GO Biological Process ontology and then report which expected functions are missing.

I can:

  1. Pull the latest GO Biological Process list.
  2. Match your nodes against it.
  3. Output the GO BP terms absent from your network.

Do you want me to fetch the current GO BP ontology from the Gene Ontology database and run that comparison?

You said:

based on the imported knowledge graph tell me what KEGG processes are missing?

ChatGPT said:

To identify which KEGG processes are missing from your imported knowledge graph, the first step is to retrieve an up-to-date list of KEGG human pathways and then compare that against the nodes present in your graph. Here’s how I can help:

  1. Fetch the current list of KEGG human pathways

I can pull the latest catalog of KEGG pathways for Homo sapiens, which includes major metabolic routes (e.g., purine metabolism, glycolysis) as well as signaling pathways (e.g., JAK-STAT, Wnt, PI3K-Akt) KEGG+1.

  1. Compare with your knowledge graph nodes

Once I have that KEGG list, I’ll match it against your graph’s node name entries. Any pathway names or identifiers (like hsa04110 for “Cell cycle”) not present in your graph would be flagged as missing.

  1. Deliver the findings

I’ll provide you with:

  • A list of KEGG processes present in your graph
  • A list of those missing, enabling you to fill in gaps or trigger further investigation

Does that sound good? If you’d like, please share a sample or the full list of your graph’s node name entries, especially if they include KEGG-style labels or pathway names—it’ll help speed up the matching process.

Sources

You said:

Fetch the current list of KEGG human pathways and compare with my knowledge graph nodes and provide me with a list of KEGG processes present and a list of KEGG processes missing.

ChatGPT said:

I’ve fetched the current list of KEGG human pathways (hsa) and compared it to the nodes in your imported knowledge graph. Here’s what we can do next—and what I need from you to complete the analysis:

KEGG Human Pathway List

Below are just a few highlights from the complete KEGG human pathway list (retrieved via KEGG REST /list/pathway/hsa) rest.kegg.jp:

  • hsa00010 – Glycolysis / Gluconeogenesis
  • hsa00020 – Citrate cycle (TCA cycle)
  • hsa01100 – Metabolic pathways
  • hsa04010 – MAPK signaling pathway
  • hsa04151 – PI3K-Akt signaling pathway
  • hsa04910 – Insulin signaling pathway
  • hsa05200 – Pathways in cancer
  • …plus hundreds more spanning metabolism, genetic information processing, organismal systems, disease, and more.

You can view the full list in the KEGG API or in your local copy if you’ve grabbed it via /list/pathway/hsa.

What I Need From You

To identify which KEGG pathways are present or missing in your graph, please provide the list of node names (or IDs) from your knowledge graph that correspond to KEGG pathways. For example, this might include:

  • KEGG identifiers like hsa04110
  • Pathway names like “Cell cycle” or “Glycolysis / Gluconeogenesis”—ideally matching exact KEGG naming

With those node names in hand, I can:

  1. Match them against the official KEGG list.
  2. Generate two clear lists:
    • Present in your graph
    • Missing from your graph (i.e., not represented)

If you can share your node set (even a sample if too large), we can proceed immediately. Let me know when you’re ready!

Sources

You said:

based on my knowledge graph I uploaded, is there any relation between lung cancer and the immune system?

ChatGPT said:

In the knowledge graph you uploaded, there aren’t any nodes explicitly labeled with immune-related terms (like “immune,” “T cell,” “macrophage,” “cytokine,” etc.) that directly link to lung cancer concepts.

That means your graph doesn’t appear to contain an obvious connection between lung cancer and the immune system—at least not through explicitly named immune processes or cell types.

If you want, I can run a broader search for indirect immune-related concepts (e.g., inflammation, checkpoint pathways like PD-1/PD-L1, MHC) to see if they’re present but labeled differently.

You said:

Based on my knowledge graph I uploaded is there any relation between lung cancer and inflammation or genes related to inflammation?

ChatGPT said:

Your uploaded knowledge graph doesn’t contain any nodes that explicitly mention inflammation or common inflammation-related genes (e.g., TNF, interleukins, NF-κB, COX-2, JAK/STAT, chemokines).

So based on the current data, there’s no recorded link between lung cancer and inflammation in your network—either directly or via inflammation-related genes.

If you want, I can try mapping your existing nodes to KEGG’s inflammation and immune-related pathways to catch indirect associations that aren’t obvious from node names alone.

You said:

Are there any nodes in my graph that maps to KEGG inflammation and immune-related pathways?

ChatGPT said:

Here’s what I discovered regarding KEGG immune- and inflammation-related pathways, and how they relate (or don’t) to the nodes in your uploaded knowledge graph:

  1. Key KEGG Immune-Related Pathways

KEGG organizes numerous human pathways under the “Immune System” category, including:

  • Toll-like receptor signaling (hsa:04620)
  • NOD-like receptor signaling (hsa:04621)
  • RIG-I-like receptor signaling (hsa:04622)
  • Chemokine signaling (hsa:04062)
  • T cell receptor signaling
  • B cell receptor signaling
  • Antigen processing & presentation
  • Th1/Th2/Th17 cell differentiation
  • IL-17 signaling
  • Natural killer cell mediated cytotoxicity
  • Complement & coagulation cascades
  • Hematopoietic cell lineage
  • Leukocyte transendothelial migration
  • Intestinal immune network for IgA production KEGG+2KEGG+2KEGG
  1. Does Your Graph Map to These KEGG Pathways?

From your graph’s node name entries, I didn’t see any nodes that correspond to the KEGG immune or inflammation pathways listed above. There are no direct matches like:

  • “Toll-like receptor signaling”
  • “Chemokine signaling”
  • “Antigen processing and presentation”
  • “Th1 cell differentiation,” etc.

So, based solely on node names, none of the canonical KEGG immune or inflammation pathways appear to be represented in your graph.

  1. Next Steps to Explore Potential Indirect Links

Even if those pathways aren’t explicitly named, your graph might still include immune- or inflammation-related genes under alternate labels. If you’d like, I can help:

  • Search for common immune-related genes in your graph (e.g., TLR4, NLRP3, interleukins, NF-κB) to spot indirect associations;
  • Perform pathway mapping, using tools like KEGG Mapper or Reactome, to see if the gene set from your graph corresponds to any immune-related pathways—even if they’re not explicitly named.

Would you like me to proceed with one of those approaches to dig deeper?

Sources

You said:

Search for common immune-related genes in my graph and pathways using KEGG Mapper

ChatGPT said:

Your uploaded knowledge graph doesn’t contain any nodes that match common immune-related genes from the KEGG immune and inflammation pathways list—at least not under standard gene symbols like IL6, TNF, TLR4, STAT3, etc.

This means KEGG Mapper wouldn’t find direct immune pathway hits from your current node set unless those genes are present under alternate labels or embedded inside other entity names.

If you want, I can:

  • Expand the search to look for partial matches (e.g., “IL-6” or “interleukin 6” instead of strictly “IL6”)
  • Try mapping all your nodes to KEGG to see if indirect immune pathway associations appear

You said:

Based on the above analysis and my knowledge graph would you say investigating the role of the immune system in lung cancer would uncover interesting therapeutic targets?

ChatGPT said:

Thought for 9s

Short answer: Yes — very likely.
Even though your uploaded graph does not currently show immune or inflammation nodes, independent evidence strongly supports that the immune system and inflammatory programs are central to lung-cancer biology and have already yielded major therapeutic targets (and will likely yield more). PMC+1

Why I say that (big points, with sources):

  • Immune checkpoint therapy (anti-PD-1/PD-L1, anti-CTLA-4) is an established, practice-changing class of treatments in non-small cell lung cancer (NSCLC) — i.e., the immune axis is clinically actionable in lung cancer. PMCAACR Journals
  • The tumor immune microenvironment (TIME) — infiltrating T cells, myeloid cells, cytokines and stromal factors — strongly influences lung cancer progression and patient outcome, and is a major focus for biomarker/target discovery. PMC+1
  • Chronic inflammation (IL-1, IL-6, TNF, NF-κB, STAT3 and related signaling) promotes tumor initiation/progression and shapes the microenvironment in lung cancer; those pathways are plausible therapeutic targets or combinatorial partners with immunotherapy. FrontiersPMC
  • Beyond PD-1/PD-L1, many immune/TME components are under active translational study (co-inhibitory/activating receptors, cytokines, myeloid suppressor cells, TME remodeling) — so there’s real opportunity for new targets or combination strategies. NatureBioMed Central

What this means for your graph / next steps I recommend (practical, doable):

  1. Expand node matching to include gene name variants and immune terms (e.g., IL6, IL-6, interleukin 6, PDCD1, PD-1, CD274, PD-L1, TLR4, NFKB1, STAT3) — you may have immune hits under alternate labels. I can do this mapping for you and return matches.
  2. Map your gene set (all nodes that are genes/keywords) to pathway databases (KEGG, Reactome) and to immune gene lists — that will reveal indirect associations even if pathway names aren’t nodes. I can run KEGG Mapper / Reactome mapping and give a present/missing report.
  3. If you have expression / mutation / clinical metadata, do: (a) differential expression / mutation enrichment for immune genes, (b) deconvolution to estimate immune cell infiltration, (c) network-neighborhood expansion to find immune neighbors — each can point to actionable targets.
  4. Prioritize candidate targets by combining evidence: presence in your graph, pathway membership, druggability, and clinical relevance (e.g., known checkpoint ligands, cytokines, myeloid regulators).

Would you like me to (A) run a relaxed search of your graph for immune/inflammation gene name variants now, and then (B) map any hits to KEGG/Reactome so we can produce a ranked list of immune-related targets and pathways? If yes I’ll start immediately and return the matched nodes + pathway hits.

Sources

 

 

Short answer: Yes — very likely.
Even though your uploaded graph does not currently show immune or inflammation nodes, independent evidence strongly supports that the immune system and inflammatory programs are central to lung-cancer biology and have already yielded major therapeutic targets (and will likely yield more). (PMC)

Why I say that (big points, with sources):

  • Immune checkpoint therapy (anti-PD-1/PD-L1, anti-CTLA-4) is an established, practice-changing class of treatments in non-small cell lung cancer (NSCLC) — i.e., the immune axis is clinically actionable in lung cancer. (PMC, AACR Journals)
  • The tumor immune microenvironment (TIME) — infiltrating T cells, myeloid cells, cytokines and stromal factors — strongly influences lung cancer progression and patient outcome, and is a major focus for biomarker/target discovery. (PMC)
  • Chronic inflammation (IL-1, IL-6, TNF, NF-κB, STAT3 and related signaling) promotes tumor initiation/progression and shapes the microenvironment in lung cancer; those pathways are plausible therapeutic targets or combinatorial partners with immunotherapy. (Frontiers, PMC)
  • Beyond PD-1/PD-L1, many immune/TME components are under active translational study (co-inhibitory/activating receptors, cytokines, myeloid suppressor cells, TME remodeling) — so there’s real opportunity for new targets or combination strategies. (Nature, BioMed Central)

What this means for your graph / next steps I recommend (practical, doable):

  1. Expand node matching to include gene name variants and immune terms (e.g., IL6, IL-6, interleukin 6, PDCD1, PD-1, CD274, PD-L1, TLR4, NFKB1, STAT3) — you may have immune hits under alternate labels. I can do this mapping for you and return matches.
  2. Map your gene set (all nodes that are genes/keywords) to pathway databases (KEGG, Reactome) and to immune gene lists — that will reveal indirect associations even if pathway names aren’t nodes. I can run KEGG Mapper / Reactome mapping and give a present/missing report.
  3. If you have expression / mutation / clinical metadata, do: (a) differential expression / mutation enrichment for immune genes, (b) deconvolution to estimate immune cell infiltration, (c) network-neighborhood expansion to find immune neighbors — each can point to actionable targets.
  4. Prioritize candidate targets by combining evidence: presence in your graph, pathway membership, druggability, and clinical relevance (e.g., known checkpoint ligands, cytokines, myeloid regulators).

Would you like me to (A) run a relaxed search of your graph for immune/inflammation gene name variants now, and then (B) map any hits to KEGG/Reactome so we can produce a ranked list of immune-related targets and pathways? If yes I’ll start immediately and return the matched nodes + pathway hits.

Based on the uploaded knowledge graph would there be any plausible reason to develop PD-L1 inhibitors for treatment of lung cancer?

ChatGPT said:

Thought for a few seconds

Short answer: Not from the uploaded graph alone.
Your current knowledge graph contains no node matching PD-L1 (CD274) or close synonyms (PD-L1, PDL1, PDCD1LG1/2) and—more broadly—no explicit immune/inflammation gene nodes were found. So, based solely on the graph you uploaded there’s no direct evidence in that dataset that would, by itself, justify developing PD-L1 inhibitors.

That said — important context and why the answer is still “probably yes” in real biology:

  • Clinically, PD-1/PD-L1 blockade is already an established, effective therapy in many lung cancers (especially NSCLC). The absence of PD-L1 in your graph likely reflects a gap in the data you uploaded (labels, scope, or data sources), not a biological counterargument.
  • Rationale for PD-L1 targeting typically depends on features that your graph currently lacks: PD-L1 (CD274) expression on tumor or immune cells, presence of PD-1 on T cells, tumor mutational burden/neoantigens, and evidence of T cell infiltration or an immunosuppressive microenvironment. None of those features appear in your nodes.

If you want to pursue PD-L1 as a target for this dataset, here are concrete next steps I recommend (I can do these for you):

  1. Search the graph for PD-L1 synonyms and related immune genes (CD274, PDCD1, PDCD1LG2, PD1, PDL1, PD-L2, CD8A, CTLA4, etc.) — I’ll run a relaxed search (IL-6 vs IL6, hyphens, full names) to catch alternate labels.
  2. Map all gene nodes to standard gene symbols (HGNC) and then run pathway mapping (KEGG/Reactome) to reveal immune pathway hits even if pathway names aren’t nodes.
  3. If you have expression or mutation metadata (linked to nodes), check: PD-L1 expression levels, TMB/neoantigen load, and immune-deconvolution signatures — these are strong predictors of benefit from PD-1/PD-L1 therapies.
  4. Network neighborhood expansion: find genes directly connected to known immune regulators — that can reveal indirect immune associations missed by exact-name matching.
  5. If no immune content is present, consider integrating orthogonal data (TCGA expression for lung cancer, published PD-L1 expression datasets) to augment the graph and inform target prioritization.

Summary and Discussions:  The goal of our original curation was to determine if gaps of knowledge exists between five whole exome studies on mutational spectrum in non small cell lung cancer (NSCLC) and if differences in mutational spectrum in smokers versus non-smokers.  Our manual curation involved a process in which individual genes and mutations were analayzed for literature coverage in Pubmed and if the potential affected pathways were also being investigated in lung cancer research (what we called a literomics approach).  Our manual analysis (as of 2016) revealed while many mutated genes were involved in the well researched fields of Cell Cycle, there were substantial gaps in knowledge of the role of the immune system in lung cancer, especially given the mutational spectrum seen in these studies.  We had also noticed a number of fusion proteins which may be interesting for further (post 2016) investigation.  This involved some inference into the use of ALK inhibitors and a suggestion of noncanonical pathways of EGFR to smoker versus nonsmoker patients, based on differences in mutational spectrum and KEGG analysis.

Using both an AI tool to generate knowledge graphs and gain insights into knowledge gaps (InfraNodus) and a generative AI new tool (Chat GPT5) we attempted to determine if our inital analysis in 2016 using more labor intensive manual curation methods could be similar to results that both AI tools could infer.  It is interesting to note that InfraNodus generated knowledge graphs could generate concepts and relationships pertinent to lung cancer, mutational spectrum and gave some interesting insights into the importance of transversions, especially relating to fusion proteins.  InfraNodus did not see much relations to immune functions however to further probe this we asked the same question to GPT5 in two different formats: with text alone and text with uploaded knowledge graph.   Surprisingly Chat GPT had some issues retrieving data from certain online open access databases such as NCBI GO but better luck with the KEGG database.  However GPT, being trained on the most recent data inferred there must be an immune component of lung cancer, although it admitted this was from recent studies; not the studies we supplied to it.  When we narrowed down GPT to look at studies before 2018 there was similarities in the relations and lack of relations we had found in our previous manual method.  We then supplied GPT with our knowledge graph and forced GPT to focus on our knowledge graph from older studies.  Under these constraints GPT correctly admitted there were no links between the immune system and lung cancer mutational specrum although it did give some interesting insights into the role of fusion proteins and reactive oxygen signaling.  After our intial curation, one of our experts Dr. Larry Bernstein had noticed that KEAP1 and 2 showed genetic alterations in the studies, as he suggested there were differences in redox signaling between smokers and nonsmokers.  KEAP1 and 2 are intracellular redox sensors.

 

Therefore it is possible that GPT alone, including the new 5 version, may not be as effective in complex inference into biomedical literature analysis, and a human expert curated knowledge graph incorporated into GPT analysis returns better inference and more novel insights than either modality alone.

For further reading on Artificial Intelligence, Machine Learning and Immunotherapy on this Open Access Scientific Journal please read these articles:

https://pharmaceuticalintelligence.com/2021/07/06/yet-another-success-story-machine-learning-to-predict-immunotherapy-response/

https://pharmaceuticalintelligence.com/2021/05/04/machine-learning-ml-in-cancer-prognosis-prediction-helps-the-researcher-to-identify-multiple-known-as-well-as-candidate-cancer-diver-genes/

Part D: Curation entitled Multiple Lung Cancer Genomic Projects Suggest New Targets, Research Directions for Non-Small Cell Lung Cancer originally published on 09/05/2014

  • Note the text below this point was used for all AI-based text analsysis

UPDATED 10/10/2021

lung cancer

(photo credit: cancer.gov)

A report Lung Cancer Genome Surveys Find Many Potential Drug Targets, in the NCI Bulletin,

http://www.cancer.gov/ncicancerbulletin/091812/page2

summarizes the clinical importance of five new lung cancer genome sequencing projects. These studies have identified genetic and epigenetic alterations in hundreds of lung tumors, of which some alterations could be taken advantage of using currently approved medications.

The reports, all published this month, included genomic information on more than 400 lung tumors. In addition to confirming genetic alterations previously tied to lung cancer, the studies identified other changes that may play a role in the disease.

Collectively, the studies covered the main forms of the disease—lung adenocarcinomas, squamous cell cancers of the lung, and small cell lung cancers.

“All of these studies say that lung cancers are genomically complex and genomically diverse,” said Dr. Matthew Meyerson of Harvard Medical School and the Dana-Farber Cancer Institute, who co-led several of the studies, including a large-scale analysis of squamous cell lung cancer by The Cancer Genome Atlas (TCGA) Research Network.

Some genes, Dr. Meyerson noted, were inactivated through different mechanisms in different tumors. He cautioned that little is known about alterations in DNA sequences that do not encode genes, which is most of the human genome.

Four of the papers are summarized below, with the first described in detail, as the Nature paper used a multi-‘omics strategy to evaluate expression, mutation, and signaling pathway activation in a large cohort of lung tumors. A literature informatics analysis is given for one of the papers.  Please note that links on GENE names usually refer to the GeneCard entry.

Paper 1. Comprehensive genomic characterization of squamous cell lung cancers[1]

The Cancer Genome Atlas Research Network Project just reported, in the journal Nature, the results of their comprehensive profiling of 230 resected lung adenocarcinomas. The multi-center teams employed analyses of

  • microRNA
  • Whole Exome Sequencing including
    • Exome mutation analysis
    • Gene copy number
    • Splicing alteration
  • Methylation
  • Proteomic analysis

Summary:

Some very interesting overall findings came out of this analysis including:

  • High rates of somatic mutations including activating mutations in common oncogenes
  • Newly described loss of function MGA mutations
  • Sex differences in EGFR and RBM10 mutations
  • driver roles for NF1, MET, ERBB2 and RITI identified in certain tumors
  • differential mutational pattern based on smoking history
  • splicing alterations driven by somatic genomic changes
  • MAPK and PI3K pathway activation identified by proteomics not explained by mutational analysis = UNEXPLAINED MECHANISM of PATHWAY ACTIVATION

however, given the plethora of data, and in light of a similar study results recently released, there appears to be a great need for additional mining of this CGAP dataset. Therefore I attempted to curate some of the findings along with some other recent news relevant to the surprising findings with relation to biomarker analysis.

Makeup of tumor samples

230 lung adenocarcinomas specimens were categorized by:

Subtype

33% acinar

25% solid

14% micro-papillary

9% papillary

8% unclassified

5% lepidic

4% invasive mucinous
Gender

Smoking status

81% of patients reported past of present smoking

The authors note that TCGA samples were combined with previous data for analysis purpose.

A detailed description of Methodology and the location of deposited data are given at the following addresses:

Publication TCGA Web Page: https://tcga-data.nci.nih.gov/docs/publications/luad_2014/

Sequence files: https://cghub.ucsc.edu

Results:

Gender and Smoking Habits Show different mutational patterns

 

WES mutational analysis

  1. a) smoking status

– there was a strong correlations of cytosine to adenine nucleotide transversions with past or present smoking. In fact smoking history separated into transversion high (past and previous smokers) and transversion low (never smokers) groups, corroborating previous results.

mutations in groups              Transversion High                   Transversion Low

TP53, KRAS, STK11,                 EGFR, RB1, PI3CA

     KEAP1, SMARCA4 RBM10

 

  1. b) Gender

Although gender differences in mutational profiles have been reported, the study found minimal number of significantly mutated genes correlated with gender. Notably:

  • EGFR mutations enriched in female cohort
  • RBM10 loss of function mutations enriched in male cohort

Although the study did not analyze the gender differences with smoking patterns, it was noted that RBM10 mutations among males were more prevalent in the transversion high group.

Whole exome Sequencing and copy number analysis reveal Unique, Candidate Driver Genes

Whole exome sequencing revealed that 62% of tumors contained mutations (either point or indel) in known cancer driver genes such as:

KRAS, EGFR, BRMF, ERBB2

However, authors looked at the WES data from the oncogene-negative tumors and found unique mutations not seen in the tumors containing canonical oncogenic mutations.

Unique potential driver mutations were found in

TP53, KEAP1, NF1, and RIT1

The genomics and expression data were backed up by a proteomics analysis of three pathways:

  1. MAPK pathway
  2. mTOR
  3. PI3K pathway

…. showing significant activation of all three pathways HOWEVER the analysis suggested that activation of signaling pathways COULD NOT be deduced from DNA sequencing alone. Phospho-proteomic analysis was required to determine the full extent of pathway modification.

For example, many tumors lacked an obvious mutation which could explain mTOR or MAPK activation.

 

Altered cell signaling pathways included:

  • Increased MAPK signaling due to activating KRAS
  • Higher mTOR due to inactivating STK11 leading to increased proliferation, translation

Pathway analysis of mutations revealed alterations in multiple cellular pathways including:

  • Reduced oxidative stress response
  • Nucleosome remodeling
  • RNA splicing
  • Cell cycle progression
  • Histone methylation

Summary:

Authors noted some interesting conclusions including:

  1. MET and ERBB2 amplification and mutations in NF1 and RIT1 may be unique driver events in lung adenocarcinoma
  2. Possible new drug development could be targeted to the RTK/RAS/RAF pathway
  3. MYC pathway as another important target
  4. Cluster analysis using multimodal omics approach identifies tumors based on single-gene driver events while other tumor have multiple driver mutational events (TUMOR HETEROGENEITY)

Paper 2. A Genomics-Based Classification of Human Lung Tumors[2]

The paper can be found at

http://stm.sciencemag.org/content/5/209/209ra153

by The Clinical Lung Cancer Genome Project (CLCGP) and Network Genomic Medicine (NGM),*,

Paper Summary

This sequencing project revealed discrepancies between histologic and genomic classification of lung tumors.

Methodology

– mutational analysis by whole exome sequencing of 1255 lung tumors of histologically

defined subtypes

– immunohistochemistry performed to verify reclassification of subtypes based on sequencing data

Results

  • 55% of all cases had at least one oncogenic alteration amenable to current personalized treatment approaches
  • Marked differences existed between cluster analysis within and between preclassified histo-subtypes
  • Reassignment based on genomic data eliminated large cell carcinomas
  • Prospective classification of 5145 lung cancers allowed for genomic classification in 75% of patients
  • Identification of EGFR and ALK mutations led to improved outcomes

Conclusions:

It is feasible to successfully classify and diagnose lung tumors based on whole exome sequencing data.

Paper 3. Genomic Landscape of Non-Small Cell Lung Cancer in Smokers and Never-Smokers[3]

A link to the paper can be found here with Graphic Summary: http://www.cell.com/cell/abstract/S0092-8674%2812%2901022-7?cc=y?cc=y

Methodology

  • Whole genome sequencing and transcriptome sequencing of cancerous and adjacent normal tissues from 17 patients with NSCLC
  • Integrated RNASeq with WES for analysis of
    • Variant analysis
    • Clonality by variant allele frequency anlaysis
    • Fusion genes
  • Bioinformatic analysis

Results

  • 3,726 point mutations and more than 90 indels in the coding sequence
  • Smokers with lung cancer show 10× the number of point mutations than never-smokers
  • Novel lung cancer genes, including DACH1, CFTR, RELN, ABCB5, and HGF were identified
  • Tumor samples from males showed high frequency of MYCBP2 MYCBP2 involved in transcriptional regulation of MYC.
  • Variant allele frequency analysis revealed 10/17 tumors were at least biclonal while 7/17 tumors were monoclonal revealing majority of tumors displayed tumor heterogeneity
  • Novel pathway alterations in lung cancer include cell-cycle and JAK-STAT pathways
  • 14 fusion proteins found, including ROS1-ALK fusion. ROS1-ALK fusions have been frequently found in lung cancer and is indicative of poor prognosis[4].
  • Novel metabolic enzyme fusions
  • Alterations were identified in 54 genes for which targeted drugs are available.           Drug-gable mutant targets include: AURKC, BRAF, HGF, EGFR, ERBB4, FGFR1, MET, JAK2, JAK3, HDAC2, HDAC6, HDAC9, BIRC6, ITGB1, ITGB3, MMP2, PRKCB, PIK3CG, TERT, KRAS, MMP14

Table. Validated Gene-Fusions Obtained from Ref-Seq Data

Note: Gene columns contain links for GeneCard while Gene function links are to the    gene’s GO (Gene Ontology) function.

GeneA (5′) GeneB (3′) GeneA function (link to Gene Ontology) GeneB function (link to Gene Ontology) known function (refs)
GRIP1 TNIP1 glutamate receptor IP transcriptional repressor
SGMS1 STK10 sphingolipid synthesis ser/thr kinase
RASSF3 TTYH2 GTP-binding protein chloride anion channel
KDELR2 ROS1, GOPC ER retention seq. binding proto-oncogenic tyr kinase
ACSL4 DCAF6 fatty acid synthesis ?
MARCH8 PRKG1 ubiquitin ligase cGMP dependent protein kinase
APAF1 UNC13B, TLN1 caspase activation cytoskeletal
EML4 ALK microtubule protein tyrosine kinase
EDR3,PHC3 LOC441601 polycomb pr/DNA binding ?
DKFZp761L1918,RHPN2 ANKRD27 Rhophilin (GTP binding pr ankyrin like
VANGL1 HAO2 tetraspanin family oxidase
CACNA2D3 FLNB VOC Ca++ channel filamin (actin binding)

Author’s Note:

There has been a recent literature on the importance of the EML4-ALK fusion protein in lung cancer. EML4-ALK positive lung tumors were found to be les chemo sensitive to cytotoxic therapy[5] and these tumor cells may exhibit an epitope rendering these tumors amenable to immunotherapy[6]. In addition, inhibition of the PI3K pathway has sensitized EMl4-ALK fusion positive tumors to ALK-targeted therapy[7]. EML4-ALK fusion positive tumors show dependence on the HSP90 chaperone, suggesting this cohort of patients might benefit from the new HSP90 inhibitors recently being developed[8].

Table. Significantly mutated genes (point mutations, insertions/deletions) with associated function.

Gene Function
TP53 tumor suppressor
KRAS oncogene
ZFHX4 zinc finger DNA binding
DACH1 transcription factor
EGFR epidermal growth factor receptor
EPHA3 receptor tyrosine kinase
ENSG00000205044
RELN cell matrix protein
ABCB5 ABC Drug Transporter

Table. Literature Analysis of pathways containing significantly altered genes in NSCLC reveal putative targets and risk factors, linkage between other tumor types, and research areas for further investigation.

Note: Significantly mutated genes, obtained from WES, were subjected to pathway analysis (KEGG Pathway Analysis) in order to see which pathways contained signicantly altered gene networks. This pathway term was then used for PubMed literature search together with terms “lung cancer”, “gene”, and “NOT review” to determine frequency of literature coverage for each pathway in lung cancer. Links are to the PubMEd search results.

KEGG pathway Name # of PUBMed entries containing Pathway Name, Gene ANDLung Cancer
Cell cycle 1237
Cell adhesion molecules (CAMs) 372
Glioma 294
Melanoma 219
Colorectal cancer 207
Calcium signaling pathway 175
Prostate cancer 166
MAPK signaling pathway 162
Pancreatic cancer 88
Bladder cancer 74
Renal cell carcinoma 68
Focal adhesion 63
Regulation of actin cytoskeleton 34
Thyroid cancer 32
Salivary secretion 19
Jak-STAT signaling pathway 16
Natural killer cell mediated cytotoxicity 11
Gap junction 11
Endometrial cancer 11
Long-term depression 9
Axon guidance 8
Cytokine-cytokine receptor interaction 8
Chronic myeloid leukemia 7
ErbB signaling pathway 7
Arginine and proline metabolism 6
Maturity onset diabetes of the young 6
Neuroactive ligand-receptor interaction 4
Aldosterone-regulated sodium reabsorption 2
Systemic lupus erythematosus 2
Olfactory transduction 1
Huntington’s disease 1
Chemokine signaling pathway 1
Cardiac muscle contraction 1
Amyotrophic lateral sclerosis (ALS) 1

A few interesting genetic risk factors and possible additional targets for NSCLC were deduced from analysis of the above table of literature including HIF1-α, mIR-31, UBQLN1, ACE, mIR-193a, SRSF1. In addition, glioma, melanoma, colorectal, and prostate and lung cancer share many validated mutations, and possibly similar tumor driver mutations.

KEGGinliteroanalysislungcancer

 please click on graph for larger view

Paper 4. Mapping the Hallmarks of Lung Adenocarcinoma with Massively Parallel Sequencing[9]

For full paper and graphical summary please follow the link: http://www.cell.com/cell/abstract/S0092-8674%2812%2901061-6

Highlights

  • Exome and genome characterization of somatic alterations in 183 lung adenocarcinomas
  • 12 somatic mutations/megabase
  • U2AF1, RBM10, and ARID1A are among newly identified recurrently mutated genes
  • Structural variants include activating in-frame fusion of EGFR
  • Epigenetic and RNA deregulation proposed as a potential lung adenocarcinoma hallmark

Summary

Lung adenocarcinoma, the most common subtype of non-small cell lung cancer, is responsible for more than 500,000 deaths per year worldwide. Here, we report exome and genome sequences of 183 lung adenocarcinoma tumor/normal DNA pairs. These analyses revealed a mean exonic somatic mutation rate of 12.0 events/megabase and identified the majority of genes previously reported as significantly mutated in lung adenocarcinoma. In addition, we identified statistically recurrent somatic mutations in the splicing factor gene U2AF1 and truncating mutations affecting RBM10 and ARID1A. Analysis of nucleotide context-specific mutation signatures grouped the sample set into distinct clusters that correlated with smoking history and alterations of reported lung adenocarcinoma genes. Whole-genome sequence analysis revealed frequent structural rearrangements, including in-frame exonic alterations within EGFR and SIK2 kinases. The candidate genes identified in this study are attractive targets for biological characterization and therapeutic targeting of lung adenocarcinoma.

Paper 5. Integrative genome analyses identify key somatic driver mutations of small-cell lung cancer[10]

Highlights

  • Whole exome and transcriptome (RNASeq) sequencing 29 small-cell lung carcinomas
  • High mutation rate 7.4 protein-changing mutations/million base pairs
  • Inactivating mutations in TP53 and RB1
  • Functional mutations in CREBBP, EP300, MLL, PTEN, SLIT2, EPHA7, FGFR1 (determined by literature and database mining)
  • The mutational spectrum seen in human data also present in a Tp53-/- Rb1-/- mouse lung tumor model

 

Curator Graphical Summary of Interesting Findings From the Above Studies

DGRAPHICSUMMARYNSLCSEQPOST

The above figure (please click on figure) represents themes and findings resulting from the aforementioned studies including

questions which will be addressed in Future Posts on this site.

UPDATED 10/10/2021

The following article uses RNASeq to screen lung adenocarcinomas for fusion proteins in patients with either low or high tumor mutational burden. Findings included presence of MET fusion proteins in addition to other fusion proteins irrespective if tumors were driver negative by DNASeq screening.

High Yield of RNA Sequencing for Targetable Kinase Fusions in Lung Adenocarcinomas with No Mitogenic Driver Alteration Detected by DNA Sequencing and Low Tumor Mutation Burden

Source:

High Yield of RNA Sequencing for Targetable Kinase Fusions in Lung Adenocarcinomas with No Mitogenic Driver Alteration Detected by DNA Sequencing and Low Tumor Mutation Burden
Ryma BenayedMichael OffinKerry MullaneyPurvil SukhadiaKelly RiosPatrice DesmeulesRyan PtashkinHelen WonJason ChangDarragh HalpennyAlison M. SchramCharles M. RudinDavid M. HymanMaria E. ArcilaMichael F. BergerAhmet ZehirMark G. KrisAlexander Drilon and Marc Ladanyi

Abstract

Purpose: Targeted next-generation sequencing of DNA has become more widely used in the management of patients with lung adenocarcinoma; however, no clear mitogenic driver alteration is found in some cases. We evaluated the incremental benefit of targeted RNA sequencing (RNAseq) in the identification of gene fusions and MET exon 14 (METex14) alterations in DNA sequencing (DNAseq) driver–negative lung cancers.

Experimental Design: Lung cancers driver negative by MSK-IMPACT underwent further analysis using a custom RNAseq panel (MSK-Fusion). Tumor mutation burden (TMB) was assessed as a potential prioritization criterion for targeted RNAseq.

Results: As part of prospective clinical genomic testing, we profiled 2,522 lung adenocarcinomas using MSK-IMPACT, which identified 195 (7.7%) fusions and 119 (4.7%) METex14 alterations. Among 275 driver-negative cases with available tissue, 254 (92%) had sufficient material for RNAseq. A previously undetected alteration was identified in 14% (36/254) of cases, 33 of which were actionable (27 in-frame fusions, 6 METex14). Of these 33 patients, 10 then received matched targeted therapy, which achieved clinical benefit in 8 (80%). In the 32% (81/254) of DNAseq driver–negative cases with low TMB [0–5 mutations/Megabase (mut/Mb)], 25 (31%) were positive for previously undetected gene fusions on RNAseq, whereas, in 151 cases with TMB >5 mut/Mb, only 7% were positive for fusions (P < 0.0001).

Conclusions: Targeted RNAseq assays should be used in all cases that appear driver negative by DNAseq assays to ensure comprehensive detection of actionable gene rearrangements. Furthermore, we observed a significant enrichment for fusions in DNAseq driver–negative samples with low TMB, supporting the prioritization of such cases for additional RNAseq.

Translational Relevance

Inhibitors targeting kinase fusions have shown dramatic and durable responses in lung cancer patients, making their comprehensive detection critical. Here, we evaluated the incremental benefit of targeted RNA sequencing (RNAseq) in the identification of gene fusions in patients where no clear mitogenic driver alteration is found by DNA sequencing (DNAseq)–based panel testing. We found actionable alterations (kinase fusions or MET exon 14 skipping) in 13% of cases apparently driver negative by previous DNAseq testing. Among the driver-negative samples tested by RNAseq, those with low tumor mutation burden (TMB) were significantly enriched for gene fusions when compared with the ones with higher TMB. In a clinical setting, such patients should be prioritized for RNAseq. Thus, a rational, algorithmic approach to the use of targeted RNA-based next-generation sequencing (NGS) to complement large panel DNA-based NGS testing can be highly effective in comprehensively uncovering targetable gene fusions or oncogenic isoforms not just in lung cancer but also more generally across different tumor types.

A Commentary is in the same issue at https://clincancerres.aacrjournals.org/content/25/15/4586?iss=15

Wake Up and Smell the Fusions: Single-Modality Molecular Testing Misses Drivers

by Kurtis D. Davies and Dara L. Aisner

Abstract

Multitarget assays have become common in clinical molecular diagnostic laboratories. However, all assays, no matter how well designed, have inherent gaps due to technical and biological limitations. In some clinical cases, testing by multiple methodologies is needed to address these gaps and ensure the most accurate molecular diagnoses.

See related article by Benayed et al., p. 4712

In this issue of Clinical Cancer Research, Benayed and colleagues illustrate the growing need to consider multiple molecular testing methodologies for certain clinical specimens (1). The rapidly expanding list of actionable molecular alterations across cancer types has resulted in the wide adoption of multitarget testing approaches, particularly those based on next-generation sequencing (NGS). NGS-based assays are commonly viewed as “one-stop shops” to detect a vast array of molecular variants. However, as Benayed and colleagues discuss, even well-designed and highly vetted NGS assays have inherent gaps that, under certain circumstances, are ideally addressed by analyzing the sample using an alternative approach.

In the article, the authors examined a cohort of lung adenocarcinoma patient samples that had been deemed “driver- negative” via MSK-IMPACT, an FDA-cleared test that is widely considered by experts in the field to be one of the best examples of a DNA-based large gene panel NGS assay (2). Of 589 driver-negative cases, 254 had additional material amenable for a different approach: RNA-based NGS designed specifically for gene fusion and oncogenic gene isoform detection. After accounting for quality control failures, 232 samples were successfully sequenced, and, among these, 36 samples (representing an astonishing 15.5% of tested cases) were found to be positive for a driver gene fusion or oncogenic isoform that had not been detected by DNA-based NGS. The real-world value derived from this orthogonal testing schema was more than theoretical, with 8 of 10 (80%) patients demonstrating clinical benefit when treated according to the alteration identified via the RNA-based approach.

To detect gene rearrangements that lead to oncogenic gene fusions (and to detect mutations and insertions/deletions that lead to MET exon 14 skipping), MSK-IMPACT employs hybrid capture-based enrichment of selected intronic regions from genomic DNA. While this approach has proven to be successful in a variety of settings, there are associated limitations that were determined in this study to underlie the discrepancies between MSK-IMPACT and the RNA-based assay. First, some introns that are involved in clinically actionable rearrangement events are very large, thus requiring substantial sequencing capital that can represent a disproportionate fraction of the assay. Despite the ability via NGS to perform sequencing at a large scale, this sequencing capacity is still finite, and thus decisions must be made to sacrifice coverage of certain large genomic regions to ensure sufficient sequencing depth for other desired genomic targets. In the case of MSK-IMPACT (and most other DNA-based NGS assays), certain important introns in NTRK3 and NRG1 are not included in covered content, simply because they are too large (>90 Kb each). The second primary problem with DNA-based analysis of introns is that they often contain highly repetitive elements that are extremely difficult to assess via NGS due to their recurring presence across the genome. Attempts to sequence these regions are largely unfruitful because any sequencing data obtained cannot be specifically aligned/mapped to the desired targeted region of the genome (3). This is particularly true for intron 31 of ROS1, because it contains two repetitive long interspersed nuclear elements, and many DNA-based assays, including MSK-IMPACT, poorly cover this intron (4). In this study by Benayed and colleagues, the most common discrepant alteration was fusion involving ROS1, which accounted for 10 of 36 (28%) cases. At least six of these, those that demonstrated fusion to ROS1 exon 32, were likely directly explained by incomplete intron 31 sequencing. RNA-based analysis is able to overcome the above described limitations owing to the simple fact that sequencing is focused on exons post-splicing and the need to sequence introns is entirely avoided (Fig. 1).

Figure 1.

Schematic representation of underlying genomic complexities that can lead to false-negative gene fusion results in DNA-based NGS analysis. In some cases, RNA-based approaches may overcome the limitations of DNA-based testing.

Lack of sufficient intronic coverage could not account for all of the discrepancies between DNA-based and RNA-based analysis however. Six samples in the cohort were found to be positive for MET exon 14 skipping based on RNA. In five of these, genomic alterations in MET introns 13 or 14 were observed, however they did not conform to canonical splice site alterations and thus were not initially called (although this was addressed by bioinformatics updates). In RNA-based testing, however, determination of exon skipping is simplified such that, regardless of the specific genomic alteration that interferes with splicing, absence of the exon in the transcript is directly observed (5). In another two of the discrepant cases, tumor purity was observed to be low in the sample, meaning that the expected variant allele frequency (VAF) for a genomic event would also likely be low, potentially below detectable levels. However, overexpression of the fusions at the transcript level was theorized to compensate for low VAF (Fig. 1). Additional explanations for discordant findings between the assays included sample-specific poor sequencing in selected introns and complex rearrangements that hindered proper capture (Fig. 1).

The take home message from Benayed and colleagues is simply this: there is no perfect assay that will detect 100% of the potential actionable alterations in patient samples. Even an extremely well designed, thoroughly vetted, and FDA-cleared assay such as MSK-IMPACT will have inherent and unavoidable “holes” due to intrinsic limitations. The solution to this dilemma, as adeptly described by Benayed and colleagues, is additional testing using a different approach. While in an ideal world every clinical tumor sample would be tested by multiple modalities to ensure the most comprehensive clinical assessment, the reality is that these samples are often scant and testing is fiscally burdensome (and often not reimbursed). Therefore, algorithms to determine which samples should be reflexed to secondary assays after testing with a primary assay are critical for maximizing benefit. In this study, the first algorithmic step was lack of an identified driver (because activated oncogenic drivers tend to exist exclusively of each other), which amounted to 23% of samples tested with the primary assay. In addition, the authors found a significantly higher rate of actionable gene fusions in samples with a low (<5 mut/Mb) tumor mutational burden, meaning that this metric, which was derived from the primary assay, could also be used to help inform decision making regarding additional testing. While this scenario is somewhat specific to lung cancer, similar approaches could be prescribed on a cancer type–specific basis.

These findings should be considered a “wake-up call” for oncologists in regard to the ordering and interpretation of molecular testing. It is clear from these and other published findings that advanced molecular analysis has limitations that require nuanced technical understanding. As this arena evolves, it is critical for oncologists (and trainees) to gain an increased comprehension of how to identify when the “gaps” in a test might be most clinically relevant. This requires a level of technical cognizance that has been previously unexpected of clinical practitioners, yet is underscored by the reality that opportunities for effective targeted therapy can and will be missed if the treating oncologist is unaware of how to best identify patients for whom additional testing is warranted. This study also highlights the mantra of “no test is perfect” regardless of prestige of the testing institution, number of past tests performed, or regulatory status. NGS, despite its benefits, does not mean all-encompassing. It is only through the adaptability of laboratories to utilize knowledge such as is provided by Benayed and colleagues that advances in laboratory medicine can be quickly deployed to maximize benefits for oncology patients.

References:

  1. Comprehensive genomic characterization of squamous cell lung cancers. Nature 2012, 489(7417):519-525.
  2. A genomics-based classification of human lung tumors. Science translational medicine 2013, 5(209):209ra153.
  3. Govindan R, Ding L, Griffith M, Subramanian J, Dees ND, Kanchi KL, Maher CA, Fulton R, Fulton L, Wallis J et al: Genomic landscape of non-small cell lung cancer in smokers and never-smokers. Cell 2012, 150(6):1121-1134.
  4. Takeuchi K, Soda M, Togashi Y, Suzuki R, Sakata S, Hatano S, Asaka R, Hamanaka W, Ninomiya H, Uehara H et al: RET, ROS1 and ALK fusions in lung cancer. Nature medicine 2012, 18(3):378-381.
  5. Morodomi Y, Takenoyama M, Inamasu E, Toyozawa R, Kojo M, Toyokawa G, Shiraishi Y, Takenaka T, Hirai F, Yamaguchi M et al: Non-small cell lung cancer patients with EML4-ALK fusion gene are insensitive to cytotoxic chemotherapy. Anticancer research 2014, 34(7):3825-3830.
  6. Yoshimura M, Tada Y, Ofuzi K, Yamamoto M, Nakatsura T: Identification of a novel HLA-A 02:01-restricted cytotoxic T lymphocyte epitope derived from the EML4-ALK fusion gene. Oncology reports 2014, 32(1):33-39.
  7. Yang L, Li G, Zhao L, Pan F, Qiang J, Han S: Blocking the PI3K pathway enhances the efficacy of ALK-targeted therapy in EML4-ALK-positive nonsmall-cell lung cancer. Tumour biology : the journal of the International Society for Oncodevelopmental Biology and Medicine 2014.
  8. Workman P, van Montfort R: EML4-ALK fusions: propelling cancer but creating exploitable chaperone dependence. Cancer discovery 2014, 4(6):642-645.
  9. Imielinski M, Berger AH, Hammerman PS, Hernandez B, Pugh TJ, Hodis E, Cho J, Suh J, Capelletti M, Sivachenko A et al: Mapping the hallmarks of lung adenocarcinoma with massively parallel sequencing. Cell 2012, 150(6):1107-1120.
  10. Peifer M, Fernandez-Cuesta L, Sos ML, George J, Seidel D, Kasper LH, Plenker D, Leenders F, Sun R, Zander T et al: Integrative genome analyses identify key somatic driver mutations of small-cell lung cancer. Nature genetics 2012, 44(10):1104-1110.

Other posts on this site which refer to Lung Cancer and Cancer Genome Sequencing include:

Multi-drug, Multi-arm, Biomarker-driven Clinical Trial for patients with Squamous Cell Carcinoma called the Lung Cancer Master Protocol, or Lung-MAP launched by NCI, Foundation Medicine, and Five Pharma Firms

US Personalized Cancer Genome Sequencing Market Outlook 2018 –

Comprehensive Genomic Characterization of Squamous Cell Lung Cancers

International Cancer Genome Consortium Website has 71 Committed Cancer Genome Projects Ongoing

Non-small Cell Lung Cancer drugs – where does the Future lie?

Lung cancer breathalyzer trialed in the UK

Diagnosing Lung Cancer in Exhaled Breath using Gold Nanoparticles

Multi-drug, Multi-arm, Biomarker-driven Clinical Trial for patients with Squamous Cell Carcinoma called the Lung Cancer Master Protocol, or Lung-MAP launched by NCI, Foundation Medicine, and Five Pharma Firms

Read Full Post »

« Newer Posts - Older Posts »