WO2026089611A1 - Circulating cancer catalog - Google Patents
Circulating cancer catalogInfo
- Publication number
- WO2026089611A1 WO2026089611A1 PCT/NL2025/050538 NL2025050538W WO2026089611A1 WO 2026089611 A1 WO2026089611 A1 WO 2026089611A1 NL 2025050538 W NL2025050538 W NL 2025050538W WO 2026089611 A1 WO2026089611 A1 WO 2026089611A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- ctcs
- subject
- tumor
- cancer
- ccc
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6876—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
- C12Q1/6883—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
- C12Q1/6886—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material for cancer
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61K—PREPARATIONS FOR MEDICAL, DENTAL OR TOILETRY PURPOSES
- A61K39/00—Medicinal preparations containing antigens or antibodies
- A61K39/0005—Vertebrate antigens
- A61K39/0011—Cancer antigens
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q2535/00—Reactions characterised by the assay type for determining the identity of a nucleotide base or a sequence of oligonucleotides
- C12Q2535/122—Massive parallel sequencing
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q2600/00—Oligonucleotides characterized by their use
- C12Q2600/156—Polymorphic or mutational markers
Landscapes
- Health & Medical Sciences (AREA)
- Chemical & Material Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Organic Chemistry (AREA)
- Immunology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Zoology (AREA)
- Microbiology (AREA)
- Oncology (AREA)
- Wood Science & Technology (AREA)
- Pathology (AREA)
- Engineering & Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- Analytical Chemistry (AREA)
- Genetics & Genomics (AREA)
- Mycology (AREA)
- Biotechnology (AREA)
- Veterinary Medicine (AREA)
- Public Health (AREA)
- Physics & Mathematics (AREA)
- Biophysics (AREA)
- Animal Behavior & Ethology (AREA)
- Hospice & Palliative Care (AREA)
- Molecular Biology (AREA)
- Epidemiology (AREA)
- Pharmacology & Pharmacy (AREA)
- Medicinal Chemistry (AREA)
- Biochemistry (AREA)
- Bioinformatics & Cheminformatics (AREA)
- General Engineering & Computer Science (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
The invention relates to the field of cancer. In particular, it relates to preparing a circulating cancer catalog (CCC) of a subject afflicted with or previously afflicted with cancer. The circulating cancer catalog includes a representative set of circulating tumor cells (CTCs) and for each CTC in the set the somatic mutations plus tumor-associated antigens (TAAs) and/or overexpressed tumor antigens expressed in it. The circulating cancer catalog is useful for exploring novel therapeutic targets, determining appropriate therapy based on available cancer treatments, and generating personalized cancer therapies. The invention further relates to the methods of preparing cancer vaccines and the like based on this circulating cancer catalog.
Description
Title: Circulating Cancer Catalog
FIELD OF THE INVENTION
The invention relates to the field of cancer. In particular, it relates to preparing a circulating cancer catalog (CCC) of a subject afflicted with or previously afflicted with cancer. The circulating cancer catalog includes a representative set of circulating tumor cells (CTCs) and for each CTC in the set the somatic mutations plus overexpressed tumor antigens and tumor -associated antigens (TAAs) expressed in it. The circulating cancer catalog is useful for exploring novel therapeutic targets, determining appropriate therapy based on available cancer treatments, and generating and designing personalized cancer therapies. The invention further relates to the methods of preparing cancer vaccines and cancer immunotherapies and the like based on this circulating cancer catalog.
BACKGROUND OF THE INVENTION
Cancer remains one of the leading causes of mortality worldwide, with various forms of the disease manifesting in different tissues and organs. Despite advances in cancer treatments, including surgery, radiation, chemotherapy, and targeted therapies, the ability to early detect and manage cancer and its metastatic growth continues to be a significant challenge. This difficulty is largely due to tumor heterogeneity, evolution, and the complexity of metastasis.
A tumor is typically composed of billions of cells with diverse genotypes and phenotypes, making it difficult to obtain a complete picture of its characteristics and to target it effectively from all angles. Tumors also exhibit dynamic behavior, continuously evolving over time. However, most conventional therapies are based on tissue samples taken at a single time point, which can contribute to treatment resistance as the tumor adapts. Furthermore, cancer cells can spread throughout the body, enter a dormant state, and cause recurrence years after the initial treatment. In fact, metastasis, the process by which cancer cells spread from the primary tumor to distant sites in the body, is responsible for the vast majority of cancer-related deaths; in many cases the primary tumor itself is not causing death. Therefore, there is a critical need for improved methods for early detection, monitoring during and post-treatment, and treating cancer.
In addition, biopsies from the primary tumor, currently the method of choice for designing even personalized cancer treatments, are almost by definition not ideal: the primary tumor is heterogeneous, but also only a small subset of the cells in the primary tumor have the ability to leave the solid tumor, pass through cell layers to enter the bloodstream (or other fluids such as saliva or urine) and survive there to spread to distant loci. Therefore, this invention focuses on the actual culprits in tumor lethality: the ciculating tumor cells (CTCs).
SUMMARY OF THE INVENTION
In one aspect the disclosure provides a method of preparing a circulating cancer catalog (CCC) of a subject afflicted with or previously afflicted with cancer or is at elevated risk of developing cancer (e.g. a solid tumor), the method comprising: performing whole genome sequencing from a healthy sample from said subject, preferably wherein the healthy sample is white blood cells or buccal cells;
providing a bodily fluid sample of the subject, preferably blood, saliva or urine;
isolating a plurality of circulating tumor cells (CTCs) from the bodily fluid sample, preferably at least 48 CTCs or 24 CTCs
performing whole genome sequencing of said CTCs, preferably long-read whole genome sequencing;
performing RNA sequencing on RNA of said CTCs, preferably long-read RNA sequencing or long-read sequencing of the corresponding cDNA, and wherein RNA is preferably poly-(A) selected mRNA;
identifying somatic mutations in the genome sequence of said CTCs, said identification comprising for each CTC:
a) comparing the genome sequences from the healthy sample and the CTC to identify potential somatic mutations; and
b) identifying the presence of somatic mutations in said CTC if said potential somatic mutations also occur in the RNA of said CTC, and/or if said potential somatic mutations are identified in the genome sequence of at least one other CTC; and
identifying overexpressed tumor antigens and/or tumor- associated antigens (TAAs) expressed by the CTCs,
wherein said somatic mutations and overexpressed tumor antigens and/or TAAs form the CCC.
In one aspect the disclosure provides a method of preparing a circulating cancer catalog of a subject afflicted with or previously afflicted with cancer or is at elevated risk of developing cancer (e.g. a solid tumor), the method comprising the steps of:
performing whole genome sequencing from a healthy sample from said subject, preferably wherein the healthy sample is white blood cells or saliva; providing a bodily fluid sample of the subject, preferably blood, saliva or urine;
isolating a plurality of circulating tumor cells (CTCs) from said sample, preferably at least 48 CTCs
performing whole genome sequencing of a part of the plurality of CTCs, preferably at least 20 of said CTCs, preferably long-read whole genome sequencing,
performing RNA sequencing of a part of the plurality of CTCs, preferably at least 20 of said CTCs, preferably long-read RNA sequencing or long-read sequencing of the corresponding cDNA,
identifying somatic mutations in the genome sequence of said CTCs, said step comprising:
a) comparing the genome sequences from the healthy sample and the CTC to identify potential somatic mutations; and
b) identifying the presence of somatic mutations in a CTC if said potential somatic mutations occur in the RNA of at least one other CTC, and/or if said potential somatic mutations are identified in the genome sequence of at least one other CTC; and
identifying overexpressed tumor antigens and/or tumor- associated antigens (TAAs) expressed by the CTCs,
wherein said somatic mutations and overexpressed tumor antigens and/or TAAs form the CCC.
In some embodiments, whole genome sequencing is performed to achieve between 5x to 30x sequencing depth and/or whole genome sequencing is performed to obtain between 4 x108 and 8 x108 sequencing reads.
In some embodiments, the identification of TAAs expressed by CTCs comprises comparing the RNA levels of a CTC against the RNA levels of TAAs from a database.
In some embodiments, the method further comprises identifying neoantigens encoded by the somatic mutations.
In some embodiments, at least 8 or 48 CTCs are isolated from a bodily fluid sample.
In some embodiments, the method comprises determining the HLA-type of at least one MHC molecule in said subject.
In some embodiments, said method further comprises identifying a treatment for said subject based on the CCC of said subject. In some embodiments, the method comprises identifying one or more vaccines for targeting
- one or more TAAs expressed by at least one CTC of the CCC,
- one or more overexpressed tumor antigens expressed by at least one CTC and/or - one or more neoantigens encoded by the somatic mutations in at least one CTCs of the CCC.
In some embodiments, the method comprises identifying at least one vaccine that targets a TAA, an overexpressed tumor antigen or a neoantigen expressed by at least two CTCs from said subject.
In some embodiments, the method further comprises preparing one or more vaccines for targeting a CTC in the subject.
In some embodiments a method is provided for preparing a treatment for a subject having cancer, said method comprising preparing a circulating cancer catalog (CCC) of the subject from a plurality of circulating tumor cells (CTCs) as disclosed herein and preparing one or more vaccines for targeting
- one or more TAAs expressed by at least one CTC of the CCC
- one or more overexpressed tumor antigens expressed by at least one CTC and/or - one or more neoantigens encoded by the somatic mutations in at least one CTCs of the CCC.
In some embodiments, the method further comprises preparing at least one vaccine (e.g., a peptide vaccine) that targets a TAA, overexpressed tumor antigen, or a neoantigen expressed by at least two CTCs from said subject.
In some embodiments, wherein the method comprises selecting one or more vaccines for targeting at least two or more TAAs, two or more overexpressed tumor antigens, and/or at least two or more neoantigens expressed by at least one CTC. In some embodiments, the targeting of a TAA, overexpressed tumor antigen or neoantigen comprises identifying and/or preparing at least three different vaccines targeting said TAA, overexpressed tumor antigen or neoantigen, respectively. In some embodiments, the method further comprises identifying a treatment for said subject, wherein the treatment comprises one or more of the following:
- preparing a chimeric antigen receptor (CAR)-T cell specific for a TAA, overexpressed tumor antigen or neoantigen identified in the CCC;
- preparing natural killer (NK) cells engineered to target one or more TAAs, overexpressed tumor antigens, or neoantigens identified in the CCC;
- preparing an antibody or antigen binding fragment thereof specific for one or more TAAs or neoantigens identified in the CCC;
- preparing a nucleic acid vaccine encoding one or more TAAs, overexpressed tumor antigen, or neoantigens identified in the CCC.
In some embodiments, said tumor- associated antigens (TAAs) are not expressed in healthy tissue of said subject and/or said TAAs are cancer testis antigens.
In some embodiments, wherein cancer vaccines are identified and/or prepared which target at least 2 CTCs from said subject.
Further provided are uses of the cancer vaccines disclosed herein for treatment as well as methods of treating a subject comprising administering to a subject in need thereof an effective amount of the cancer vaccines disclosed herein. The subject may be afflicted with cancer. The treatment may result in the treatment of cancer, including the treatment/re duction of symptoms and the delay/reduction of metastases.
In one aspect the disclosure provides a method for identifying a treatment for an individual expressing a TAA, overexpressed tumor antigen, and/or neoantigen, said method comprising
-preparing a CCC according to any one of the preceding embodiments for a plurality of subjects afflicted with or previously afflicted with cancer, -optionally, preparing one or more cancer vaccines according to any one of the preceding embodiments for each subject,
-optionally, administering said vaccines to each subject,
-collecting data on the subjects related to the outcome of treatment, in particular the ability of said vaccines to kill CTCs, preferably wherein such data is used as input for artificial intelligence and/or machine learning, and
-identifying one or more cancer vaccines that are likely to treat said individual. Preferably, wherein the HLA-type of the individual is also determined as described herein. Preferably, wherein the one or more vaccines administered to said subjects target the TAA or neoantigen expressed by said individual.
In one aspect the disclosure provides a method implemented by one or more processors executing computer program instructions that, when executed, perform the method, the method comprising:
i) providing data related to a plurality of patients to a prediction model (e.g., a neural network or other machine learning model),
wherein each patient is or was afflicted with cancer and has been administered one or more vaccines targeting a TAA, overexpressed tumor antigen, or neoantigen expressed by one or more CTCs of said patient,
wherein said data comprises for each patient
- the HLA-type of the patient,
- the amino acid sequences of the one or more peptide vaccines administered to said patient,
- the immunogenicity and/or clearance of CTCs expressing the target of said one or more peptide vaccines, following administration of said peptide vaccines;
ii) training the prediction model with said data in order to predict which amino acid sequences are most likely to increase immunogenicity and/or clearance of CTCs.
In one aspect the disclosure provides a method implemented by one or more processors executing computer program instructions that, when executed, perform the method, the method comprising:
i) preparing a CCC according to any one of the preceding embodiments for a plurality of subjects afflicted with or previously afflicted with cancer,
ii) providing data related to the plurality of subjects to a prediction model (e.g., a neural network or other machine learning model),
wherein each subject has been administered one or more peptide vaccines targeting a TAA, overexpressed tumor antigen, or neoantigen expressed by one or more CTCs of said subject,
wherein said data comprises for each subject
- the HLA-type of the subject,
- the amino acid sequences of the one or more peptide vaccines administered to said subject,
- the immunogenicity and/or clearance of CTCs expressing the target of said one or more peptide vaccines, following administration of said peptide vaccines;
iii) training the prediction model with said data.
In some embodiments, the method further comprises:
iv) providing to the trained prediction model, the HLA-type of an individual and the sequence of at least one TAA, overexpressed tumor antigen, or neoantigen expressed in one or more CTC’s of said individual and
v) obtaining from the trained prediction model, one or more peptide sequences predicted to be effective in increasing the immunogenicity and/or clearing CTCs expressing said TAA, overexpressed tumor antigen, or neoantigen.
BRIEF DESCRIPTION OF THE DRAWINGS
Figure 1. Figure depicts the distinction between tumor cells obtained through a traditional thin needle biopsy and CTCs obtained from the blood stream. Tumor heterogeneity is a common hallmark of cancer. A thin needle biopsy is only sampling part of a large solid tumor and therefore misses tumor heterogeneity. By contrast, CTCs isolated from the blood represent different populations of heterogeneous cells from a solid tumor and therefore provide a more realistic representation of the tumor in a patient.
Figure 2. Possibilities of capturing and analysing CTCs during different stages of disease (breast cancer). Since liquid biopsies (e.g. blood) is readily available from a patient with cancer, CTCs can be isolated from a patient at different time points before, during and after cancer treatment. Hence, ongoing evolution of cancer cells will be robustly detected and provide an optimal and up-to-date basis for therapy choice.
Figure 3. Schematic drawing depicting workflow for diagnostic leukapheresis (DLA) and subsequent capture of CTCs from DLA products.
Figure 4. Venn diagram showing overlap of non-synonymous somatic mutations (single-nucleotide variants and short indels) for a primary tumor, a metastatic tumor and CTCs. The data demonstrate that the somatic genetic mutations present in CTCs more resemble metastatic cancer vs primary cancer. Data and figure was obtained from Ni, Xiaohui, et al. "Reproducible copy number variation patterns among single circulating tumor cells of lung cancer patients." Proceedings of the National Academy of Sciences 110.52 (2013): 21083-21088.
Figure 5. Covering the full genome for single-cell whole genome sequencing data is challenging, because of the risk of loss of DNA during genomic library preparation. Increasing sequencing throughput improves genome coverage. The graph shows percentages of mappable genome covered (Y-axis) as a function of the number of sequence reads generated for genomic DNA derived from a single cell (X-axis)
Figure 6. Plot depicting mRNA expression values (i.e. number of RNA sequencing reads) for genes expressed in a single LNCaP cell. Genes are ranked based on their mRNA expression values. Red dots represent TAAs, in particular cancer-testis antigens (CTAs). X-axis: each gene sequenced and ordered based on its expression. Y-axis: number of RNA sequencing reads.
Figure 7. Example of a peptide vaccine composition generated on the bases of TAA expression in LNCaP cells. Each peptide is depicted as a horizontal bar with amino acids in different shades of grey. For each of the top three most expressed TAAs a peptide vaccine is designed. A mixture of peptides targeting said selected TAAs (and/or possible neoantigens) can subsequently be manufactured and administered to a patient afflicted with cancer.
Figure 8. Capture efficiency from spiked healthy donor blood samples with cells from different cancer cell lines (LNCaP, 22Rv1, PC3-9, PC3) using the FETCH enrichment system showing ability to capture tumour cells with varying Ep CAM -expression from blood.
Figure 9. Whole Genome Sequencing data of CTCs from patients with prostate cancer. For 4 CTCs isolated from a patient with prostate cancer, Whole Genome Sequencing data were generated. For each of the cells, somatic genomic structural variants and DNA copy number changes were determined by comparison to reference white blood cells from the same patient.
Figure 10. Barplot depicting short- and long-read RNA sequencing gene counts for small batches of MCF7 cancer cells and single MCF7 cancer cells. Two versions of RNA- sequencing library prep were used and the RNA libraries were sequenced on Illumina (short-read) and Oxford Nanopore (long-read) instruments. Per experiment the number of genes with quantifiable expression were counted and plotted.
Figure 11. Full-length RNA sequencing coverage of the GAPDH gene. For MCF7 cells, long-read RNA sequencing was performed on the Oxford Nanopore sequencer. Many RNA sequencing reads cover the full-length GAPDH gene from the 5’ end to the 3’ polyA tail (Full). Also partial GAPDH covering RNA reads are shown (Partial). The lower panel shows known GAPDH transcript isoforms.
DETAILED DESCRIPTION OF THE DISCLOSED EMBODIMENTS
One promising area of research in cancer diagnostics and treatment is the study of circulating tumor cells (CTCs). CTCs are cells that have broken away from a solid tumor and enter bodily fluids (primarily the bloodstream). While a primary cancer (e.g., tumor) may be removed or destroyed during treatment such that a patient may enter remission, relapse of cancer as well as eventual death is often the result of metastases. Metastases is the spread of cancer cells from one part of the body to another. Such cancer cells (i.e., circulating tumor cells) generally travel through the blood or lymph system and can form a new tumor in other parts of the body. These CTCs can be isolated while they are in circulation, e.g., from the blood or
other bodily fluids (such as saliva for some Head & Neck cancers and urine for bladder cancer). Disseminated Tumor Cells (DTCs) are cancer cells that have exited the blood stream or lymph system and have settled is distant organs such as bone marrow or lymph nodes.
The detection and characterization of CTCs offer a minimally invasive way to monitor the progression of cancer, assess prognosis, and evaluate the efficacy of treatments. CTCs also provide insights into the molecular characteristics of tumors, including the identification of specific genetic mutations, which can guide personalized treatment strategies. Despite their clinical significance, the detection and analysis of CTCs is complicated by their rarity in the bloodstream, often appearing at very low concentrations amidst a vast number of normal blood cells.
There is an unmet need to not only enhance the characterization of CTCs but also to generate new therapeutic approaches to target these cells and prevent metastasis. One object of the present disclosure to provide methods for identifying somatic mutations in CTCs and overexpressed tumor antigens and/or tumor-associated antigens (TAAs) expressed in CTCs. A further object of the present disclosure is to provide subject- specific cancer vaccines and treatments based on said identified somatic mutations, overexpressed tumor antigens, and TAAs.
The disclosure provides methods for preparing a circulating cancer catalog (CCC). The CCC provides information regarding the expression of TAAs and/or overexpressed tumor antigens in CTCs as well as somatic mutations in CTC’s, in particular those that result in the production of neoantigens.
CTCs are highly heterogeneous, reflecting the genetic and phenotypic diversity of the primary tumor. There are numerous studies that demonstrate that gene expression and mutational analyses can greatly differ between different biopsy samples from the same tumor. While this tumor heterogeneity can play a significant role, e.g., in the treatment/resistance to treatment of a tumor, not all of the tumor heterogeneity is relevant in respect to metastasis since not all tumor cells are capable of leaving the tumor and penetrating surrounding tissue. In contrast, CTCs by definition have left the primary tumor and are thus relevant to the prognosis and treatment of metastasis.
In view of the high heterogeneity of CTCs, the characterization of a single CTC in a subject provides limited value and treatment designed to target a single CTC is highly unlikely to be successful since the treatment is unlikely to have an effect on other (uncharacterized) CTCs. One of the advantages of a CCC as described herein
is that it provides a platform for the deep molecular characterization of a multitude of CTCs.
As a skilled person will appreciate, the Circulating Cancer Catalog (CCC) described herein provides a number of utilities. For example, the CCC provides a snapshot of potential therapeutic targets in the tumor cells, and thus allows prognostic analysis, treatment monitoring and can guide treatment decisions. The CCC can also indicate alterations in therapeutic targets emerging during treatment as well as clinically relevant drug resistance alterations.
Firstly, the CCC can be used to guide treatment decision based on known treatment options. Some targets, e.g., HER2 or EGFR mutations, overexpressed antigens, or particular TAAs, may be known therapeutic targets for which treatment modalities are already available (e.g. known inhibitors of BRAF, HER2, EGFR, etc.). While current treatment decisions are generally based on a single-site, single-time sampling of a tumor, the CCC allows for a complete and updated view of the patient’s cancer to ensure the most tailored treatment decisions can be made.
Secondly, the CCC may also be used to design and manufacture personalized therapies, tailored to target antigens of the individual patient's tumor cells as discussed further herein.
Thirdly, applying deep multi-omics to a large collection of patient CCC’s will allow the identification of new therapeutic targets, e.g., with the assistance of artificial intelligence as discussed further herein.
In one embodiment of the disclosure, methods are provided for preparing a circulating cancer catalog (CCC) of a subject afflicted with or previously afflicted with cancer or is at elevated risk of developing cancer, comprising performing both whole genome sequencing and RNA sequencing from the same CTC. Such methods comprise:
- performing whole genome sequencing from a healthy sample from said subject, preferably wherein the healthy sample is white blood cells or buccal cells;
- providing a bodily fluid sample of the subject, preferably a blood sample, or saliva or urine;
- isolating a plurality of circulating tumor cells (CTCs) from the bodily fluid sample, preferably at least 8 or 48 CTCs;
- performing whole genome sequencing of said CTCs, preferably long-read whole genome sequencing;
- performing RNA sequencing on RNA of said CTCs, preferably long-read RNA sequencing or long-read sequencing of the corresponding cDNA, and wherein RNA is preferably poly-(A) selected mRNA;
- identifying somatic mutations in the genome sequence of said CTCs, said identification comprising for each CTC: (a) comparing the genome sequences from the healthy sample and the CTC to identify potential somatic mutations; and (b) identifying the presence of somatic mutations in said CTC if said potential somatic mutations also occur in the RNA of said CTC, and/or if said potential somatic mutations are identified in the genome sequence of at least one other CTC; and - identifying overexpressed tumor antigens and/or tumor -associated antigens (TAAs) expressed by the CTCs,
wherein said somatic mutations and overexpressed tumor antigens and/or TAAs form the CCC.
In one embodiment of the disclosure, methods are provided of preparing a circulating cancer catalog (CCC) of a subject afflicted with or previously afflicted with cancer or is at elevated risk of developing cancer, comprising performing whole genome sequencing of one CTC and RNA sequencing of another CTC. Such methods comprise:
- performing whole genome sequencing from a healthy sample from said subject, preferably wherein the healthy sample is white blood cells or saliva;
- providing a bodily fluid sample of the subject, preferably a blood sample, or saliva or urine;
- isolating a plurality of circulating tumor cells (CTCs) from said sample, preferably at least 8 or 48 CTCs;
- performing whole genome sequencing of a part of the plurality of CTCs, preferably at least 24 of said CTCs, preferably long-read whole genome sequencing, - performing RNA sequencing of a part of the plurality of CTCs, preferably at least 24 of said CTCs, preferably long-read RNA sequencing or long-read sequencing of the corresponding cDNA,
- identifying somatic mutations in the genome sequence of said CTCs, said step comprising: (a) comparing the genome sequences from the healthy sample and the CTC to identify potential somatic mutations; and (b) identifying the presence of somatic mutations in a CTC if said potential somatic mutations occur in the RNA of at least one other CTC, and/or if said potential somatic mutations are identified in the genome sequence of at least one other CTC; and
- identifying tumor- associated antigens (TAAs) and/or overexpressed tumor antigens expressed by the CTCs,
wherein said somatic mutations and TAAs and/or overexpressed tumor antigens form the CCC.
As used herein, the term “subject” refers to mammals, including both human and non-human mammals. In particular, the term “subject” includes but is not limited to humans, non-human primates, canines, murines, felines, bovines, equines, porcines and the like. Preferably, the mammal is a human. In some embodiments, the subject is afflicted with or has been previously afflicted with cancer.
The subject is afflicted or has been previously afflicted with a solid tumor. As used herein, the term “solid tumor” refers to a malignant neoplasm arising in solid organs or tissues, and excludes hematologic malignancies such as leukemias or lymphomas. Solid tumors include, for example, carcinomas, sarcomas, melanomas, and gliomas. Exemplary solid tumors include cancers of the breast, lung, colon, rectum, pancreas, prostate, liver, ovary, kidney, bladder, stomach, esophagus, head and neck, skin (melanoma), and brain (glioma). In some embodiments, the solid tumor is a metastatic tumor derived from any of the foregoing primary tumors.
As used herein, a subject “afflicted with or previously afflicted with cancer/solid tumor” refers to a subject who is currently diagnosed with, or has a known history of cancer, having been previously diagnosed and/or treated. While the subject may be afflicted with cancer, a diagnosis may not have been confirmed yet by a medical practitioner. In particular, a subject maybe suspected of being afflicted with or previously afflicted with cancer based on clinical symptoms, family history, or other diagnostic indicators. A subject afflicted with cancer may be characterized, e.g., as having stable disease (SD) or progressive disease (PD).
As is known to a skilled person, cancer remission refers to the absence of active disease for a period of at least one month. Complete remission (CR) refers to individuals who have no cancer symptoms and no evidence of cancer, e.g., detected by radiological tests. Partial remission (PR) refers to cancer that is still detectable but has decreased in size. “Recurrence” or “relapse” refer to the return of cancer after remission.
Individuals at elevated risk of developing cancer represent a significant population for early detection and preventive medicine. Such elevated risk may arise from various factors, including inherited genetic mutations (for example, pathogenic variants in BRCA1, BRCA2, TP53, MLH1, or APC), family history of cancer, exposure to environmental carcinogens, chronic inflammation, or lifestyle-related risk factors. In genetically predisposed individuals, oncogenic transformation may occur earlier or progress more rapidly than in the general population, emphasizing the importance of early molecular surveillance.
The methods described herein may be used at all stages, including those discussed above and at different timepoints. In particular, methods described herein may be used during cancer monitoring, during cancer treatment, post-remission, postrecovery, upon diagnosis, or when a subject is suspected of being afflicted with or previously afflicted with cancer.
Methods described herein may be part of routine examinations of cancer monitoring or during cancer treatment. Specifically, methods described herein may provide information on cancer prognosis and detect potential metastatic growth earlier compared to standard imaging and diagnostic methods. When used during cancer treatment, methods described herein may provide information on effectiveness of treatment.
Methods described herein may also be part of routine follow-up examinations after remission and/or recovery, occurring at intervals such as every 3, 4, 6 or 12 months. This may be useful for monitoring of residual CTCs and/or early detection of potential recurrence.
Methods described herein may also be used when a subject is suspected of being afflicted with or previously afflicted with cancer as explained above. In particular, CTCs may be released into bloodstream or body fluid even when no tumor is detectable by standard imagining or diagnostic methods. Accordingly, the methods described herein may allow early detection of cancer compared to standard imagining or diagnostic methods.
Methods described herein may also be used upon diagnosis of cancer. In particular, said methods may aid in staging the disease or determining the tumor burden.
Methods described herein may aid in providing a complete and updated characterization of a subject’s cancer to ensure the most tailored cancer treatment based on existing therapies, in identifying novel therapeutic targets (e.g. neoantigens and TAAs), and/or in generating personalized therapies such as cancer vaccines.
As used herein, a “healthy sample” refers to a non-tumorous sample, such as tissue or cells, obtained from a subject afflicted with or previously afflicted with cancer. As a skilled person will recognize, said healthy sample is from a non-tumor tissue or non-tumor cells, but may have other genetic, molecular or physiological defects or deficiencies. Suitable healthy samples include, but are not limited to, biopsy tissue sample from non-tumorous organs, white blood cells, buccal cells. A
preferred healthy sample is white blood cells or buccal cells. Said healthy sample serves as a reference for comparison with a bodily fluid sample.
As used herein, a “bodily fluid sample” refers to a fluid sample obtained from a subject afflicted with or previously afflicted with cancer. A bodily fluid sample may be selected from, for example, blood, urine, saliva, cerebrospinal fluid, lymphatic fluid, bone marrow, sputum, cyst fluid, pleural fluid, peritoneal fluid (Smit 2024, Mol Asp of Med 96, 101258; Lawrence 2023, Nat Rev Clin Oncol, 1-44). A preferred bodily fluid sample is blood, saliva or urine, more preferably blood. A blood sample can be a whole blood, more preferably peripheral blood or a peripheral blood cell fraction. A skilled person is aware of bodily fluid samples suitable for specific type of cancer. Saliva may especially be useful when providing a bodily fluid sample of a subject afflicted with or previously afflicted with head and neck cancer. Urine may especially be useful when providing a bodily fluid sample of a subject afflicted with or previously afflicted with bladder cancer.
CTCs may be isolated from the bodily fluid sample using various general methods known in the art, such as immunoaffinity-based methods, size-based filtration, density-based centrifugation, and microfluidic technologies (Banko et al 2019, J of Hemat and Oncol, 12:48; Edd et al 2022, iScience 25:104696). Immunoaffinity-based techniques utilize antibodies targeting CTC-specific surface markers for capturing CTCs. Examples of CTC-specific surface markers include, but are not limited to, EpCAM, EGFR, HER2, CDH11, and/or MET. Example of a commercially available and FDA approved immunoaffinity-based CTC isolation method is CellSearch (Menarini Silicon Biosystems), which detects CTCs based on a combination of markers EpCAM, CK8, CK18, and CK19 for positive selection of CTCs. Other immunoaffinity-based methods include Size-based methods leverage the larger size of CTCs relative to bodily fluid components, typically relative to other blood cells, separating them through filtration. Density-based centrifugation techniques, such as OncoQuick, rely on the difference in buoyancy between CTCs and other components of bodily fluids, (e.g. other blood components such as red blood cell, white blood cells). Microfluidic methods on the other hand often combine several separation principles, such as size-based and immuno affinity techniques, to enhance CTC isolation Example of a microfluidic system useful for CTC isolation is CTC-iChip. Alternatively, CTC may also be isolated based on their bioelectric properties using dielectrophoresis (Edd et al 2022, iScience 25:104696).
As explained herein, CTCs are highly heterogeneous. One of the advantages of the CCC as described herein is that it provides highly detailed information (e.g., whole genome sequence as well as expression values for a large number of mRNA molecules (e.g., more than one million) for each CTC and that information from
multiple CTCs in included in the catalog. A skilled person will appreciate that any plurality (e.g., 2 or more) of cells may be obtained. In some embodiments, at least 8 CTCs, at least 10 CTCs, at least 16 CTCs, at least 24 CTCs, at least 48 CTCs, at least 96 CTCs are isolated from a bodily fluid sample. In some embodiments, at least 20 CTCs, at least 40 CTCs, or at least 80 CTCs are isolated from a bodily fluid sample. Preferably, at least 48 CTCs are isolated from a bodily fluid. Preferably, at least 8 CTCs are isolated from a bodily fluid.
In some embodiments, whole genome sequencing and RNA sequencing may be performed on the same cell. Protocols suitable for single cell WGS combined with RNA sequencing are known to a skilled person and may include parallel nucleic acid isolation or split-sample approaches wherein the genomic DNA and RNA fractions are separately processed from a single captured cell. In some embodiments, the method comprising performing single cell WGS and single cell RNA sequencing on a plurality of CTCs.
In certain embodiments, the sequencing is performed as batch sequencing, meaning that multiple single-cell sequencing reactions are processed simultaneously in a shared sequencing run or analytical workflow. In an exemplary embodiment, each cell is uniquely barcoded prior to library pooling, ensuring that the sequence reads originating from each cell can be accurately demultiplexed and attributed to the corresponding individual cell after sequencing. In some embodiments, the sequencing may be performed as bulk sequencing or semi-bulk sequencing. As used herein, bulk sequencing refers to sequencing nucleic acids obtained from a pooled population of cells, such that the genetic material of multiple cells is combined prior to sequencing.
In certain embodiments, the sequencing may be performed as semi-bulk sequencing, wherein nucleic acids derived from small subsets of cells (for example, 2–100 cells per group) are pooled and sequenced together. Each subset may optionally be provided with a unique barcode or identifier, allowing the sequencing data to be attributed to the corresponding subset.
In some embodiments, whole genome sequencing will be performed on a part (e.g., a first part) of the plurality of CTCs while RNA sequencing will be performed on a part (e.g., a second part) of the plurality of CTCs. While the plurality may be split into two equal or nearly equal parts, e.g., 24 cells and 24 cells, a skilled person recognizes that other embodiments are also encompassed, such as 1/3 of the plurality is subjected to WGS and 2/3 of the plurality is subjected to RNA sequencing, or vice-versa.
The term “sequencing”, as used herein, refers to determining the order of nucleotides (base sequences) in a nucleic acid sample, e.g. DNA or RNA.
The expression “whole genome sequencing (WGS)”, as used herein, refers to analysing and determining the entire DNA sequence of a genome. WGS enables the detection of variations such as single nucleotide polymorphisms (SNPs), insertions, deletions, and larger structural rearrangements. WGS method typically involve fragmenting the genomic DNA into smaller segments, which are then sequenced and reassembled computationally to determine sequence of the entire genome. WGS may be performed using either short-read sequencing or long-read sequencing methods. Preferably, WGS is a long-read sequencing.
As used herein, "short read sequencing" refers to sequencing technique that generates relatively small fragments of DNA, typically ranging 50-300 base pairs in length, followed by computationally reconstructing the entire genome from said short reads. These methods are also known in the art as second- generation sequencing or next-generation sequencing. Examples of short-read sequencing methods include Illumina sequencing and Ion Torrent sequencing platforms. Shortread sequencing methods are advantageous for their high accuracy and costeffectiveness in detecting single nucleotide variants (SNVs) and small insertions or deletions (indels).
As used herein, “long-read sequencing” (also known as “third generation sequencing”) refers to a sequencing technology that produces significantly longer DNA fragments, generally longer than 1000 base pairs and often ranging from several kilobases up to tens of kilobases per read. Long-read sequencing platforms, such as those provided by Pacific Biosciences (PacBio) and Oxford Nanopore Technologies, are particularly advantageous for detecting large structural variants, phasing haplotypes, and analyzing repetitive regions within the genome. Long-read sequencing also offers greater accuracy in reconstructing complex genomic regions. Long-read sequencing is further compatible with RNA sequencing (RNA-seq) applications, particularly in single-cell analysis, where full-length transcript information may be required.
In embodiments where WGS and RNA sequencing are performed in the same cell, WGS methods compatible with RNA sequencing should be used. For example, WGS and RNA sequencing may be performed simultaneously. Examples of such simultaneous methods can be found in the paper by Rossi and Zamarchi (Rossi and Zamarchi 2019, Front Genet, 958), and include methods such as ‘genome and transcriptome sequencing (G& T-seq)’ (Macaulay et al 2016, Nat. Protoc. 11, 2081–2103), ‘DNA and RNA sequencing (DR-seq)’ (Dey et al 2015, Nat. Biotechnol. 33,
285–289), and ‘Simultaneous Isolation of genomic DNA and total RNA (SIDR seq)’ (Han et al 2018, Genome Res. 28, 75–87).
The expression “sequencing reads”, as used herein, refers to the piece of DNA that is sequenced (“read”) by a nucleic acid sequence.
As used herein, the term “sequencing depth” relates to the number of times a sequenced region is covered by the sequence reads. For example, an
average sequencing depth of 10-fold assumes that each nucleotide within the sequenced region is covered on average by 10 sequence reads.
In embodiments of the present disclosure, whole genome sequencing is performed to achieve at least 5x sequencing depth, preferably at least 10× or at least 15x. In some embodiments whole genome sequencing is performed to achieve between 5x to 30x sequencing depth.
In some embodiments, whole genome sequencing is performed to obtain at least 108 sequencing reads, such as 4 xlO8 sequencing reads, preferably at least 6 xlO8 sequencing reads. In some embodiments, whole genome sequencing is performed to obtain between 4 x108 and 8 x108 sequencing reads. In some embodiments, such reads are short reads such as less than 300bp. A skilled person will appreciate that when using long read sequencing fewer sequence reads may be necessary.
As is used herein, “RNA sequencing” also termed “RNA-Seq”, refers to a sequencing technique to characterize the quantity and/or sequence of RNA in a sample. RNA sequencing includes both direct RNA sequencing and sequencing based on corresponding cDNA. In some embodiments, the RNA is first reversed transcribed into cDNA and said cDNA is sequenced. In some embodiments, direct RNA sequencing is performed. Preferably, said RNA sequencing attempts to sequence essentially all RNA molecules (also referred to as total RNA/whole transcriptome sequencing) or essentially all mRNA molecules (also referred to as mRNA sequencing) and in this respect differs from targeted RNA sequencing in which only specific transcripts are sequenced.
In some embodiments, RNA sequencing is performed to obtain at least 103 sequencing reads, at least 104 sequencing reads, at least 105 sequencing reads, at least 106 sequencing reads, such as at least 6 xlO8 sequencing reads. In some embodiments, RNA sequencing is performed to obtain between 103 and 109 sequencing reads. In some embodiments, such reads are short reads such as less than 300bp. A skilled person will appreciate that when using long read sequencing fewer sequence reads may be necessary. As discussed in Example 1 and depicted in
Figure 6, more shallow RNA sequencing can be used when RNAseq is used simply to determine the presence and expression levels of highly expressed TAAs, whereas deeper sequencing can be employed in order to confirm that somatic DNA mutations are true mutations and are also present in mRNA.
Preferably, the RNA isolated for sequencing is cytosolic RNA, excluding tRNA or rRNA, more preferably the isolated RNA is mRNA. Preferably, the RNA is poly-(A) mRNA, more preferably poly-(A) mRNA. A skilled person is familiar with methods for isolating poly-(A)RNA, which typically involves binding total RNA to poly-(T) oligomers, retaining only the RNA that binds to said oligomers. In some embodiments, the RNA contains a 5’-cap. Protocols for selection 5’-capped RNA are well known to a skilled person (see e.g., Weiss B, Curran JA. CAP(+) selection: A combined chemical-enzymatic strategy for efficient eukaryotic messenger RNA enrichment via the 5' cap. Anal Biochem. 2015).
Methods of performing RNAseq are known to a skilled person and are also described, e.g., at Ura, H., et al. BMC Genomics 23, 303 (2022). Preferably, a single-cell RNAseq protocol is used such as described in Cheng et al. Cells. 2023 Aug; 12(15): 1970. More preferably, long-read single-cell RNAseq protocols are used such as the 10× Genomics Chromium system, the Oxford Nanopore Technology (ONT) platform PromethION and the PacBio system Sequel II system.
Identifying somatic mutations in the genome sequence of CTCs comprises comparing the genome sequences from the healthy sample and the CTC to identify potential somatic mutations. As used herein, “somatic mutations” refers to mutations in genomic DNA and includes, e.g., single nucleotide variants, inversions/deletions, and chromosomal rearrangements such as inversions, duplications, and translocations. There are a number of software programs available for identify mutations from sequence data such as VarScan2 and Mutect2.
As a skilled person will appreciate, sequencing errors can complicate the identification of actual mutations. Such errors may be introduced if a PCR amplification step is used (e.g., from polymerase misincorporation, template switching, DNA damage), while others are specific to the type of sequencing platform used.
In order to confirm whether the mutations identified are true mutations, one or more filtering steps can be performed. In some embodiments, a potential somatic mutation is identified as a true mutation if the mutation is also identified in the genomic sequence of at least one other CTC from said patient. In some
embodiments, a potential somatic mutation is identified as a true mutation if the mutation is also identified in the RNA sequence of at least one other CTC from said patient. In embodiments where both whole genome sequencing and RNA sequencing are performed in a single cell, a potential somatic mutation is identified as a true mutation if the mutation also occurs in the RNA sequences of the same CTC.
In some embodiments, the methods comprise identifying neoantigens (or rather, peptide sequences) encoded by the somatic mutations. Tumor- specific antigens (TSAs), also referred to herein as ‘neoantigens’, derive from tumor specific alterations such as mutations in the tumor genome. Tumor -specific antigens are recognized as non-self and can elicit an immune response. Neoantigens can arise from any number of different types of genomic mutations, in particular SNV and indels resulting in frameshift mutations. A skilled person is well-aware of how to identify neoantigen amino acid sequences based on the genomic sequence. Peptide predictions software is also available such as pVACtools and MuPeXI.
The disclosure provides methods for preparing a circulating cancer catalog (CCC) of a subject comprising antigens that are expressed or overexpressed in circulating tumor cells (CTCs), relative to their expression in normal tissue. These ‘tumor expressed antigens’ encompass tumor-specific antigens (TSAs or neoantigens) arising from somatic mutations, tumor-associated antigens (TAAs) such as differentiation antigens and cancer-testis antigens (CTAs), as well as overexpressed tumor antigens that are expressed at elevated levels in tumor cells compared to normal cells. In preferred embodiments, the CCC comprises TAAs and neoantigens arising from somatic mutations. In preferred embodiments, the CCC comprises i) TAAs and/or overexpressed tumor antigens and ii) neoantigens arising from somatic mutations.
In some embodiments the CCC comprises tumor- associated antigens (TAAs) expressed by the CTCs of the subject. In some embodiments, the TAAs are Cancer testis antigens (CT A). Tumor antigens include tumor-specific antigens (TSAs) and tumor- associated antigens (TAAs). TAAs refer to antigens that are abnormally expressed in tumor tissues as well as at low levels in normal tissue and/or during specific stages of differentiation. For example, “differentiation antigens” refers to TAAs that are derived from lineage-specific proteins that are expressed on tumor cells and nonmalignant cells of the same cell lineage during at least some stage of differentiation. TAAs are annotated in a number of different databases including the HPtaa database.
Cancer testis antigens (CTA) are a family of tumor-associated antigens expressed in human tumors, but not in normal tissues except for testis and
placenta. Examples of CTAs include, e.g., Acrosin binding protein (ACRBP), Testis-expressed protein 15 (TEX15), Nuclear RNA Export Factor 2 (NXF2). Additional CTAs are annotated at the CTA database on the world wide web at cta.lncc.br. In a preferred embodiment, the TAAs are CTAs.
Overexpressed tumor antigens represent an important category of tumor-expressed targets. Unlike neoantigens, which arise from somatic mutations, overexpressed antigens correspond to, in general, normal cellular proteins whose expression is significantly upregulated in malignant cells. Such overexpression may result from gene amplification, transcriptional dysregulation, or altered signaling pathways. Well-characterized examples include HER2 amplification in breast and gastric cancers, EGFR overexpression in non-small cell lung cancer and glioblastoma, MUC1 overexpression in pancreatic and breast cancer, and mesothelin in ovarian and mesothelioma tumors.
Overexpressed antigens may include, for example, oncogenic receptors and growth factors such as epidermal growth factor receptor (EGFR), HER2/ERBB2, MUC1, CEA (carcinoembryonic antigen), mesothelin, survivin (BIRC5), and others known in the art. Identification of such antigens may be achieved by comparing transcriptomic expression levels (e.g., RNA sequencing data) in CTCs against matched normal tissue or established expression databases such as GTEx or TCGA. In some embodiments, methods further comprise identifying overexpressed tumor antigens that are expressed in a CTC.
The methods described herein include sequencing the RNA of CTCs. In addition to providing information regarding potential somatic mutations, RNA sequencing provides information regarding the expression levels of RNA (e.g., mRNA). An exemplary workflow for carrying out single-cell RNAseq to determine expression is as follows. RNAs are reverse-transcribed into complementary DNAs (cDNAs), amplified and then subjected to NGS. The sequencing reads can be mapped on a reference genome and the number of reads aligned to each gene (i.e., ‘counts’) provides an indication of the gene expression level in the respective CTC. A skilled person will appreciate that other methods are also suitable. Preferably the singlecell RNA seq protocol is compatible with whole genome DNA sequencing.
In some embodiments, methods further comprise identifying TAAs that are expressed in a CTC. TAAs are known in the art and, as described above, a number of databases are available for determining whether a gene is considered a TAA. A large number of TAAs have also been described in the literature.
In some embodiments, the methods further comprise ranking the expression of TAAs in order to identify TAAs that are highly expressed in a CTC. Figure 6 depicts an exemplary embodiment in this respect. The X-axis plots each gene sequenced and ordered based on expression (Y-axis depicts the number of reads). In Figure 7, the 3 TAAs with the highest gene expression are PBK, MAGEA1, and MAGEA3.
As a skilled person will appreciate, knowledge as to the circulating cancer catalog (CCC) of a subject is useful to assist a medical practitioner for designing new therapeutics as well as predicting the subject’s response to known therapeutics. Accordingly, the disclosure further provides methods for identifying and/or preparing a treatment for a subject having cancer based on the subject’s CCC. In some embodiments, the methods comprising preparing a CCC from a subject as described further herein and preparing or selecting one or more vaccines for targeting
- one or more TAAs expressed by at least one CTC and/or
- one or neoantigens encoded by the somatic mutations in at least one CTCs.
In some embodiments, one or more vaccines is selected for targeting one or more overexpressed tumor antigens expressed by at least one CTC.
In a preferred embodiment, a vaccine described herein is a peptide vaccine. As known to a skilled person, peptide cancer vaccines are made up of amino acids that correspond to tumor antigens, such as TAAs or neoantigens. Short peptides, such as 8-12 amino acid maybe used. Longer peptides having at least 20 amino acids may also be used and are processed by antigen presenting cells to produce short peptides which can be presented to both HLA class I and HLA class II molecules. While not wishing to be bound by theory, it is though that the presentation of a tumor specific antigen (via a peptide cancer vaccine) to a T-cell leads to activation of T-cells which then increases the antitumor response.
In some embodiments, a peptide vaccine as described herein is at least 8 amino acids in length, between about 8 to about 100 amino acids in length, preferably from 8 to 50 amino acids, more preferably from 8 to 35 amino acids.
In some embodiments, vaccines are prepared or selected which target multiple CTCs. In an exemplary embodiment, three TAAs are targeted (each by three different vaccines, such as peptides) for each four different CTCs. In some embodiments, vaccines are prepared or selected which target at least 2 CTCs, at least 4 CTCs, at least 8 CTCs, at least 10 CTCs, or at least 24 CTCs. As explained further below, some vaccines may target more than one CTC.
While ideally one would want to target all overexpressed tumor antigents, TAAs and all neoantigens expressed in the CCC, a skilled person recognizes that there is a limit to the number of different vaccines can be administer. Therefore, a skilled person recognizes that some filtering may be needed in order to select the top candidates for targeting as well as for designing vaccines that target the tumor antigens.
As depicted in Figure 7, the tumor targets selected may be those that exhibit the highest expression levels. However, as discussed below, additional considerations may also be relevant. Therefore, the tumor targets are generally those with a relatively high level of expression, but are not necessarily those with the highest expression level.
In some embodiments, the tumor targets selected are selectively expressed in the CTCs of the subject and are not expressed in healthy tissue in the subject. Such targets may include, e.g., cancer testis antigens as well as other genes that may be expressed in tumor but not in, e.g., adult tissue.
In some embodiments, one or more CTCs from a subject may express the same TAA or the same neoantigen. Such ‘common tumor antigens’ are especially useful targets since one vaccine targets multiple CTCs. This has the advantage that fewer need to be prepared and administered in order to target a larger number of CTCs. Therefore, in some embodiments, common tumor antigens are preferentially selected for targeting. In some embodiments, the method comprises identifying (e.g., preparing or selecting) at least one vaccine that targets a TAA or a neoantigen expressed by at least two, at least three, at least five, or at least 10 CTCs from said subject. In some embodiments, only vaccines are selected which target a TAA or a neoantigen expressed by at least two, at least three, at least five, or at least 10 CTCs from said subject.
While common tumor antigens may be entirely specific to a particular patient, in some instances some TAA’s over overexpressed tumor antigens expressed by CTCs maybe ‘shared’ tumor antigens, i.e., tumor antigens that are shared among different cancer patients. The selection of ‘shared’ tumor antigens as targets may be useful if vaccines are already available.
In regards to the neoantigens included in a subject’s CCC, a preference may be made to those neoantigens that provide a large amount of novel sequences. For example, frameshift mutations may generate a neoantigens of 10, 20, 30 or more amino acids whereas a SNV will likely only result is a difference of one amino acid between the tumor sequence and the sequence in healthy cells.
In some embodiments, the method comprises selecting or preparing vaccines targeting at least two, preferably at least three tumor antigens (i.e., TAAs or overexpressed tumor antigens and neoantigens) expressed by a CTC. In some embodiments, the method comprises targeting at least two or more TAAs and/or at least two or more neoantigens from a CTC. See, e.g., Figure 7 which depicts as an example the production of three different peptides targeting PBK, three different peptides targeting MAGEA3, and three different peptides targeting PAGE1. While not wishing to be bound by theory, expression of particular tumor antigens may reduce over time. Therefore, targeting multiple different tumor antigens for each CTC is expected to increase the chance of effecting an immune response against each CTC.
In some embodiments, the targeting of a tumor antigen (i.e., a TAA or overexpressed tumor antigens or neoantigen) comprises selecting or preparing at least three different targeting said tumor antigen. As a skilled person will appreciate, this increases the chance of an immune response against said antigen.
Peptides targeting one or more TAAs or overexpressed tumor antigens and/or one or more neoantigens preferably bind major histocompatibility complex (MHC), preferably with a binding affinity of less than about 500 nM or less than 100 nM. As used herein, the term "affinity" refers to a measure of the strength of binding between the relevant peptide and MHC. Accordingly, a filtering step may also include selecting peptides based on their ability to bind MHC.
Affinity may be determined experimentally, for example by surface plasmon resonance (SPR) using commercially available Biacore SPR units. Affinity may also be predicted in silico. Suitable peptides may be predicted to bind MHC-I or MHC-II using any MHC-binding prediction algorithm known in the art, such as those freely available in the Immune Epitope Database (IEDB; www.iedb.org), NetMHCpan2.8, NetMHCpan3, NetMHCpan4, NetMHCcons, mhcflurry, mhcflurry pan, and MixMHCpred.
"Human Leukocyte Antigen" (HLA) refers to the human MHC. HLA class I genes include HLA- A, HLA-B, and HLA-C and HLA class II genes include HLA-DR, HLA-DQ, and HLA-DP. Characterization of the diversity of HLA alleles (i.e., HLA typing) can be performed by sequencing the alleles. As a skilled person will appreciate, the specific HLA- type of a subject influence whether a particular peptide vaccine will result in an immune response. In some embodiments, the HLA-type of the subject is determined. In some embodiments, the peptide vaccines are selected based on their ability to bind the HLA molecules of a subject, as an additional filtering step. In some embodiments, determining the HLA-type of a
subject comprises determining the sub-type of the A and B antigens; the A, B, and DR antigens; the A, B, and C antigens; or the A, B, C, and DR, DQ, and DP antigens. In some embodiments, the HLA-type of at least one MHC molecule is determined. In some embodiments, the HLA-type of at least one MHC class I molecule is determined. In some embodiments, the HLA-type of at least one MHC class II molecule is determined. In some embodiments, the HLA type of HLA-A is determined. In some embodiments, the HLA type of HLA-B is determined. In some embodiments, the HLA type of HLA-C is determined. In some embodiments, the HLA type of HLA-DR is determined. Preferably each subtype is determined to at least a 2-digit depth or at least a 4-digit depth. Once the HLA-type of the patient is established it can, e.g., be used as a docking model in silico to determine the binding affinity to a particular cancer vaccine. In some embodiments, peptide vaccines bind to at least one MHC class subtype of a subject with a binding affinity of less than about 500 nM or less than 100 nM. Preferably, HLA-typing is determined based on the whole genome sequencing from a healthy sample from said subject.
The vaccines, in particular the peptide vaccines, described herein are preferably "immunogenic" such that they bind HLA and induce a cell-mediated or humoral response, for example, cytotoxic T lymphocyte (CTL), helper T lymphocyte (HTL) and/or B lymphocyte response.
In some embodiments, suitable vaccines are selected based on their predicted stability, solubility, and or ease of manufacture. For example, cysteine has the propensity to form disulfide bridges. Therefore, peptide sequences maybe selected to minimize or avoid cysteines. In addition, cysteine residues can be substituted with a-amino butyric acid.
In some embodiments, suitable vaccines are selected based on the lack of sequences which correspond to sequences expressed in healthy tissue, or rather for a lack of ‘self-similarity’. For example, a frame shift mutation might lead to the production of 50 amino acids not encoded by the gene lacking the mutation. However, it is possible that by chance, eight consecutive amino acids are expressed by another gene in healthy tissue of said subject. In order to avoid negative side effects of a vaccine, potential vaccines can be screened in order to select sequences which are not expressed in healthy tissue. In some embodiments, the peptide vaccine does not comprise at least 5, preferably at least 8 amino acids expressed in healthy tissue of the subject.
In some embodiments, the vaccine may be chemically synthesized. For example, multiple peptide vaccines can be (chemically) linked together. In some
embodiments, a peptide vaccine maybe recombinantly produced, e.g., by expressing the corresponding DNA in a host cell. In such embodiments, several different peptides can be linked in a single polypeptide.
In some embodiments, peptide vaccines further comprise one or more modifications which increase in vivo half-life, cellular targeting, antigen uptake, antigen processing, MHC affinity, MHC stability, or antigen presentation.
The disclosure further provides methods of treating a subject afflicted with cancer or at risk of cancer comprising administering to a subject one or more peptide vaccines prepared as discussed herein. As a skilled person will recognize, such treatment can be used for any type of cancer so long as CTCs are available in the bodily fluid of the subject for analysis. In some embodiments, the methods comprise preparing a Circulating Cancer Catalog (CCC) as discussed herein, preparing a treatment comprising one or more peptide vaccines as described herein, and treating an individual as described herein. In some embodiments, peptide vaccines as described herein are provided for use in a method of treating cancer.
As a skilled person will appreciate, the methods described herein provide large amounts of datasets. For each patient, the entire genome and RNA will be sequenced for several CTCs. Artificial intelligence (AI) has the potential to aid in the prediction of the best choice of treatment and has been used to predict vaccine immunogenicity (see, e.g., Gonzalez-Dias P,. Hum Vaccin Immunother.
2020;16(2):269-276). Al refers to various techniques that allow computers to simulate (and in some cases surpass) human learning and includes, e.g., “machine learning” and “deep learning”. Al has also been implemented to aid in predicting personalized treatment.
The methods described herein provide extensive information regarding the different TAAs and neoantigens expressed by CTCs of a subject. Based on this data, a mixture of peptides can be designed to raise an immune response to kill the CTCs. The readout of such a treatment includes the immune response against each of the peptides in the cocktail as well as the disappearance of CTCs that express the TAAs/neoantigens.
This readout can ultimately provide a training set for AI to choose those peptides that have the highest efficacy in killing the CTCs. This is just one example of the potential application of AI to the datasets that are generated by the development of CCCs (Circulating Cancer Catalogs).
In an exemplary workflow of such a method, a prediction model (e.g., AI and models discussed in, e.g., Gonzalez-Dias P,. Hum Vaccin Immunother.
2020;16(2):269-276) is trained with data from a large number of patients, including the peptide sequences used to vaccinate each patient and the resulting immunogenicity score and CTC clearance score for each patient. The prediction model learns the appropriate patterns to predict which sequences are most likely to have an effect. The prediction model can then be used to input new neoantigens/TAAs in order to predict peptide sequences that are likely to have an effect.
In addition to peptide vaccines, the circulating cancer catalog (CCC) described herein can be used to guide the design and preparation of other therapeutic modalities directed against tumor -associated antigens (TAAs), overexpressed tumor antigens and neoantigens identified in circulating tumor cells (CTCs). Such modalities include, without limitation, engineered T cells such as chimeric antigen receptor T cells (CAR-T) and TCR-engineered T cells, natural killer (NK) cell therapies, antibody-based therapies, nucleic acid-based vaccines, and oncolytic viruses. Each of these approaches can be customized using the CCC to ensure that the therapy is tailored to the individual subject’s tumor antigen repertoire. The same considerations regarding selecting the optimal peptide targets described above, apply similar to these therapeutics.
In some embodiments, the therapeutic modality comprises preparing CAR-T cells engineered to recognize and target one or more TAAs, or overexpressed tumor antigens or neoantigens identified in the CCC. Autologous T cells maybe obtained from the subject by leukapheresis, activated, and transduced with a viral or non-viral vector encoding a chimeric antigen receptor specific for a selected TAA, or overexpressed tumor antigens, or neoantigen. The CAR construct may include an extracellular single-chain variable fragment (scFv) for antigen recognition, a hinge and transmembrane domain, and intracellular signaling domains such as CD3ζ, optionally combined with one or more co-stimulatory domains (e.g., CD28, 4- IBB). The CAR-T cells may be expanded ex vivo and administered back to the subject in an effective amount to selectively eliminate CTCs expressing the corresponding TAA or neoantigen.
In some embodiments, the therapeutic modality comprises preparing T cells engineered to express recombinant T-cell receptors (TCRs) specific for TAAs or neoantigens identified in the CCC. Such TCRs may be obtained, e.g., by cloning naturally occurring high-affinity receptors or by engineering receptors with enhanced affinity or specificity. The engineered T cells may be expanded ex vivo
and administered to the subject to mediate cytotoxicity against CTCs presenting the targeted peptide-MHC complexes.
In some embodiments, the therapeutic modality comprises preparing natural killer (NK) cells that are engineered or expanded to target TAAs, or overexpressed tumor antigens or neoantigens identified in the CCC. NK cells may be obtained from the subject or from an allogeneic source and may be expanded ex vivo using feeder cells or cytokines such as IL-2 and IL- 15. In certain embodiments, NK cells are modified to express CARs (CAR-NK cells) that target TAAs, or overexpressed tumor antigens or neoantigens identified in the CCC. Expanded or engineered NK cells may then be administered to the subject to promote clearance of CTCs.
In some embodiments, the therapeutic modality comprises antibody-based therapies, such as antibodies and antigen binding fragments thereof. Suitable antibodies include monoclonal antibodies, bispecific antibodies, or antibody-drug conjugates that specifically bind to TAAs, or overexpressed tumor antigens or neoantigens identified in the CCC. Monoclonal antibodies may be generated by any method known in the art, such as hybridoma technology, phage display, or recombinant antibody engineering.
In some embodiments, the therapeutic modality comprises nucleic acid-based vaccines. For example, an mRNA vaccine may be prepared that encodes one or more TAAs, or overexpressed tumor antigens, or neoantigens identified in the CCC. The mRNA may be formulated in lipid nanoparticles or other delivery vehicles and administered to the subject to induce expression of the target antigen and stimulate an immune response. In other embodiments, nucleic acid therapeutics such as small interfering RNAs (siRNAs) or antisense oligonucleotides may be designed to silence oncogenic transcripts expressed by CTCs, thereby reducing tumor cell viability. Reference herein to ‘peptide vaccines’ should be understood to also include RNA vaccines encoding said peptides.
In some embodiments a method is provided which a computer implemented method is provided.
Such method comprises i) preparing a CCC for a plurality of subjects afflicted with or previously afflicted with cancer, e.g., at least 10, 20, 50, 100, 500 or more subjects.
In some embodiments, the methods may further comprise collecting data on the subjects related to the outcome of individual treatments based on an individual’s CCC; for example, said individual treatments may be vaccines as disclosed herein for targeting TAAs and/or or overexpressed tumor antigens and/or neoantigens expressed in the CCC of a subject. As used herein, “outcome” refers to the ability of
a treatment to induce an immune response and/or kill a CTC (e.g., CTCs expressing the tumor antigen targeted by a peptide vaccine). This data (treatment choice (e.g., peptide vaccine targeting a particular antigen) linked to outcome) can be used as a training set for AI in order to determine which peptide vaccines are most likely to have a positive effect.
Such method comprises ii) providing data related to the plurality of subjects to a prediction model (e.g., a neural network or other machine learning model). Each subject has been administered a cancer treatment, such as one or more peptide vaccines targeting a TAA, overexpressed tumor antigen, or neoantigen expressed by one or more CTCs of said subject.
The data comprises for each subject
- the HLA-type of the subject (preferably the method comprises determining the HLA-type of the subject),
- the amino acid sequences of the one or more peptide vaccines administered to said subject (preferably the method comprises administering said peptide vaccines to said subject), and
- the immunogenicity and/or clearance of CTCs expressing the target of said one or more peptide vaccines, following administration of said peptide vaccines.
Immune response and CTC characterization can be monitored at various stages during and after treatment. For example, CTCs may be isolated at various stages during and after treatment and expression of TAAs/neoantigens/overexpressed tumor antigen targeted by treatment can be measured. Expression can be measured by any means known in the art including (targeted) RNA sequencing. Immune response, or immunogenicity, can be determined by, e.g., measuring antibody titer or T-cell response. Suitable assays include, e.g., interferon-γ (INFγ) ELISpot assay and T Cell Cytotoxicity Assays. Immunogenicity and/or clearance can be measured after one day, one week, two weeks, one month, two months, three months, six months or longer after the first treatment with said peptide vaccines.
Such method comprises iii) training the prediction model with said data. Various training models known to a skilled person can be applied. See, e.g., those described in Gonzalez-Dias P,. Hum Vaccin Immunother. 2020;16(2):269-276. The training allows the system to learn to predict amino acid sequences that are likely to have a positive therapeutic effect, e.g., increased immunogenicity and/or clearance of CTCs.
In some embodiments, the artificial intelligence engine/pre diction model may be used to identify patterns correlating a particular peptide vaccine with a desired outcome (e.g., reduction of CTCs expressing the corresponding tumor antigen). The machine learning models may then match a pattern between the characteristics of
the new patient (e.g., HLA-type and tumor antigen to be targeted) and a particular peptide vaccine.
In some embodiments, the method further comprises iv) providing to the trained prediction model, the HLA-type of an individual and the sequence of at least one TAA or neoantigen expressed in one or more CTC’s of said individual and v) obtaining from the trained prediction model, one or more peptide sequences predicted to be effective in increasing the immunogenicity and/or clearing CTCs expressing said TAA, overexpressed tumor antigen, or neoantigen. In some embodiments, the method comprises determining the HLA-type of said individual. In some embodiments, the method comprises identifying TAAs and/or neoantigens expressed in one or more CTC’s of said individual, e.g., by determining the CCC of said individual as described herein.
The screening for circulating tumor DNA (ctDNA) has become an important tool in cancer early detection, prognosis and even therapy choice and/or design. Cell-free DNA (cfDNA) refers to nucleic acid fragments released into the circulation, typically as a consequence of cell death through apoptosis or necrosis. In oncology, the tumor-derived fraction of cfDNA is referred to as circulating tumor DNA (ctDNA).
Tumor-naive assays rely on predefined mutation panels covering common oncogenes and tumor suppressors. Commercial examples include Guardant360. Because such assays scan for mutations without prior knowledge of a patient’s specific tumor profile, their sensitivity is inherently limited, particularly in early disease or minimal residual disease (MRD) settings, where ctDNA levels are very low.
Tumor-informed assays use prior knowledge of the patient’s tumor to design personalized detection panels. In these methods, the exome or genome of a primary tumor biopsy is sequenced, and somatic mutations are identified by comparison with matched normal tissue. A customized panel of clonal, tumor- specific mutations is then constructed, enabling detection of ctDNA in plasma. One example is the Signatera™ assay (Natera, Inc.).
While tumor-informed approaches provide improved sensitivity and specificity compared to tumor-naive methods, they are limited by their reliance on a tissue biopsy. The disclosure provides an improved method of screening ctDNA that does not rely on invasive tissue biopsies. Specifically, ctDNA is detected by identifying somatic mutations in cfDNA that are found in the patient’s CCC.
In a further aspect, the disclosure provides detecting circulating tumor DNA (ctDNA) in a bodily fluid sample of a subject afflicted with or previously afflicted with a solid tumor, the method comprising:
- preparing a circulating cancer catalog (CCC) of said subject, as described herein, from a plurality of circulating tumor cells (CTCs), said CCC comprising tumor-associated antigens (TAAs), overexpressed tumor antigens, and/or neoantigens expressed by said CTCs; and
- detecting, in the cfDNA of a bodily fluid sample of the subject, the presence of mutations also found in the subject’s CCC. Preferably, the method also comprises sequencing the cfDNA.
In a further aspect, the disclosure provides a method of monitoring a subject afflicted with or previously afflicted with cancer, the method comprising:
- preparing a circulating cancer catalog (CCC) of said subject, as described herein, from a plurality of circulating tumor cells (CTCs), said CCC comprising tumor-associated antigens (TAAs), overexpressed tumor antigens, and/or neoantigens expressed by said CTCs;
- following treatment of a primary tumor in said subject, analyzing the cfDNA from a bodily fluid sample from said subject for the presence of circulating tumor DNA (ctDNA). ctDNA is detected when a somatic mutation present in the patients’s CCC is identified.
In some embodiments, upon detecting ctDNA, the method comprises preparing and preferably administering to said subject a therapeutic agent, as disclosed herein, that specifically targets at least one TAA, overexpressed tumor antigen, or neoantigen identified in said CCC.
In some embodiments, the CCC of a subject is at the time of initial diagnosis or before, during, or shortly after primary treatment. Following treatment of the primary tumor, such as by surgery, radiation therapy, or chemotherapy, the subject may be monitored periodically, for example every 3, 4, 6 or 12 months, or twice per year, for signs of recurrence or residual disease.
In some embodiments, monitoring may include analyzing circulating tumor DNA (ctDNA) in a bodily fluid sample, preferably blood. Detection of ctDNA can be performed using next-generation sequencing, droplet digital PCR, or other sensitive molecular methods. The presence of ctDNA in the circulation indicates residual or recurrent tumor activity.
In some embodiments, once ctDNA is detected during monitoring, a therapy such as once disclosed herein is prepared and administered based on the CCC that was
previously established for the subject. Thus, the CCC need not be prepared again at each monitoring time point. This provides a practical and less invasive approach to patient management: the CCC is prepared once, but monitoring can be continued using, e.g., only blood plasma for ctDNA analysis.
In some embodiments, the CCC may be repeated at a later stage, for example if ctDNA levels indicate disease recurrence and there is clinical concern that tumor evolution may have altered the antigenic landscape. A new CCC can then be prepared from freshly obtained CTCs to identify additional TAAs, overexpressed tumor antigens, or neoantigens, ensuring that therapies remain tailored to the subject’s cancer as it evolves.
EXAMPLES
Example 1. Detection and isolation of CTCs
CTC isolation from blood may be performed as outlined below. Here, a blood sample or blood derivative sample is collected either in blood collection tubes containing anticoagulants, or through a (diagnostic) leukapheresis (DLA) procedure, in which the mononuclear cell fraction is targeted.
CTC Enrichment and Staining
A blood or DLA aliquot containing CTCs is processed using an immunomagnetic enrichment protocol. The aliquot is incubated with magnetic beads bound to antibodies targeting CTCs to separate them from leukocytes, for instance using a BD IMag™ magnet or FETCH magnetic separation system.
The enriched fraction containing the captured CTCs, is thereafter fluorescently stained with staining reagents for the identification of CTCs, such as for example anti-CD45 (a leukocyte marker), anti-cytokeratin, anti-EpCAM, anti-HER2, anti-EGFR, anti-PSMA or any other cancer or hemopoietic marker used to cell identity identification. Additionally, vitality, cytoplastic or membrane markers may be added such as for example Calcein-FITC (a viability stain), or Dil (a membrane marker). Staining mix may further optionally include cancer-specific markers, such as anti-PSMA, anti-HER2, anti-EGFR, anti DLL2, anti-B7H3 or any other treatment related marker. Stained cells are washed using a magnetic separation to remove unbound reagents, leaving a purified, stained cell suspension.
CTC identification
The stained cell suspension is then imaged using an inverted fluorescence microscope or flowcytometer for the identification of CTC based on the selected identification markers and if desired quantification of the absence or presence of the selected treatment markers.
CTC isolation
Following or directly during identification those cells fitting the criteria for suspected CTC are isolated for further downstream processing (e.g. genomic or
transcriptomic analysis) as single cells, or in groups, using the sorting capability of a FACS flow cytometer or using a single cell isolation technology such as a magnetic isolation needle or (automated) single cell pipetting.
Example 2. Secretion analysis of CTCs
CTC isolation from blood may be performed as outlined below. See also the schematic example workflow in figure 3. Here, a blood sample or blood derivative sample is collected either in blood collection tubes containing anticoagulants, or through a (diagnostic) leukapheresis (DLA) procedure, in which the mononuclear cell fraction is targeted.
CTC Enrichment and Staining
A blood or DLA aliquot containing CTCs is processed using an immunomagnetic enrichment protocol. The aliquot is incubated with magnetic beads bound to antibodies targeting CTCs in order to separate them from leukocytes, for instance using a BD IMag™ magnet or FETCH magnetic separation system. The enriched fraction containing the captured CTCs, is thereafter fluorescently stained with staining reagents for the identification of CTCs, such as for example anti-CD45 (a leukocyte marker), anti-cytokeratin, anti-EpCAM, anti-HER2, anti-EGFR, anti-PSMA or any other cancer or hemopoietic marker used to cell identity identification. Additionally, vitality, cytoplastic or membrante markers may be added such as for example Calcein-FITC (a viability stain), or DiLL (a membrane marker).. Staining mix may further optionally include cancer-specific markers, such as anti-PSMA, anti-HER2, anti-EGFR, anti DLL2, anti-B7H3 or any other treatment related marker. Stained cells are washedusing a magnetic separation to remove unbound reagents, leaving a purified, stained cell suspension.
CTC Sorting, Seeding, and Imaging
The stained cell suspension is optionally first processed using a FACSAria II flow cytometer. The viability positive and leukocyte marker-negative fraction is gated to isolate viable CTCs while excluding leukocytes. The sorted cells are subsequently seeded into nanowells chips, containing a port in each well capable of trapping a single cell. This seeding process ensures the isolation of individual CTCs into separate nanowells, allowing for single -cell analysis.
The nanowell array is thereafter scanned using an inverted microscope with different fluorescence channels (e.g. PE, FITC, APC, BV421). Multi-channel imaging enables the identification and confirmation of CTCs based on their staining profile.
PSA Capture and Spot Analysis
CTCs are incubated overnight in the nanowell array with a membrane pre-coated with a capture antibody such as for instance anti-PSA (Prostate-Specific Antigen) antibody to capture PSA secreted by the CTCs during incubation.
After the overnight incubation, secondary antibody anti-IgG-PE (Phycoerythrin-conjugated anti-IgG) is added to medium to allow for imprint detection of PSA secretion. The nanowell chip is thereafter disconnected from the membrane.
The PSA spots on the membrane are imaged again using an inverted microscope in the PE and FITC channels. Custom software is used to analyse the PSA spots, differentiating between secreting and non-secreting CTCs based on the IgG imprint.
CTC Isolation
Identified CTCs are isolated from their individual nanowells by for instance by micro-pipetting or magnetic needle for subsequent molecular analyses, e.g. WGS and/or RNA analysis.
Example 3. Single-cell whole genome sequencing and single-cell transcriptome sequencing of LNCaP cells.
Methods as described herein were first tested on human prostate adenocarcinoma cell line (LNCaP cells, Lymph Node Carcinoma of the Prostate cells) to establish experimental conditions for preparation of RNA and DNA libraries from the same cell. In particular, the effect of different buffers and larger sample volume on library quality was evaluated. Genome sequencing depth was also evaluated in relation to genome coverage.
Methods
LNCaP cells were collected and downstream single-cell library preparation was performed using the Bioskrybr Resolveome kit.
Downstream analyses were performed on the following samples:
Cells were lysed using lysis buffer. RNA was reverse transcribed to cDNA.
Subsequently, whole genome amplification (WGA) was performed based on the PTA (Primary Template Directed Amplification) amplification method.
Following amplification, DNA and RNA (cDNA) libraries were prepared using standard next-generation sequencing (NGS) protocols (Bioskrybr Resolveome). Briefly, for both DNA and cDNA, samples were fragmented and adapters were ligated to both ends of fragments. Thereafter libraries were subjected to high-throughput Illumina sequencing.
Genome (DNA) sequencing data was aligned against the reference human genome (hg38) using minimap2, somatic structural variants were called using GRIDSS, somatic copy-number alterations were called using aneufinder, and somatic singlenucleotide variants and short indels were calling using bcf tools and fings. RNA sequencing data was aligned against the human reference genome (hg38) using STAR, and transcript quantification was performed using RSEM against the primary gene assembly of GENCODE.
Results
Sequencing of the DNA and RNA libraries was successful. All three tested samples produced a desirable single peak (data not shown). Additionally, tested samples and negative control show no contaminations or errors. The obtained results thus confirmed good compatibility of the DNA and RNA library preparation protocol with both the X buffer and a larger sample volume.
Furthermore, Figure 5 shows how much genome coverage is gained by sequencing more reads. Due to uneven genome coverage caused by DNA amplification bias, there are diminishing returns in sequencing more reads.
Expression of genes in tested cells was evaluated and their distribution was ranked by expression, as show in Figure 6. It was observed that tumor-associated antigens, specifically cancer-testis antigens (CTAs) are notably expressed in said LNCaP cell at significant levels, making them promising targets for a (peptide) vaccine.
For example, a vaccine targeting said CTAs may be generated based on the most expressed CTAs or CTAs shared between individual cells. The top three most expressed CTAs were PBK, MAGEA3, and PAGE1. Accordingly, three peptide vaccines of, e.g. 25 amino acids in length, may be generated per individual CTA (Figure 7).
Example 4. Whole Genome Sequencing of prostate cancer CTCs Methods as described herein were tested on CTCs obtained from a patient with prostate cancer. In brief, CTCs were selected from blood of a patient with prostate cancer as described herein (Example 1). Next, individual CTCs were subject to DNA library preparation using the BioSkrybr ResolveOME kit (PTA). DNA libraries were sequenced at high coverage depth (>30X) on Illumina sequencing instruments. As a control, white blood cells of the same individual with cancer were sequenced.
Whole-genome sequencing data was aligned against the reference human genome (hg38) using minimap2, somatic structural variants were called using GRIDSS, somatic copy-number alterations were called using aneufinder, and somatic singlenucleotide variants and short indels were calling using bcf tools and fings.
The somatic structural variants for each of the four individual CTCs isolated from a patient with prostate cancer were compared, depicting both overlapping somatic genetic variants, as well as genetic variants unique to a single CTC (Figure 9).
Example 5. Short-read and Long-read RNA sequencing of single MCF7 cancer cells
Methods as described herein were tested on freshly harvested MCF7 cells. Cells were sorted into Takara lysis buffer and stored at -80 °C until further processing. cDNA synthesis was performed using the Takara SMART-Seq® mRNA Long- Read Kit (Cat. No. 634377). Sequencing libraries were prepared either with the KAPA HyperPlus Kit, employing enzymatic fragmentation (KAPA Cat. No. KK8514 I 07962428001) according to the manufacturer’s instructions, or with the Nanopore Ligation Sequencing Kit V14 without fragmentation. Equimolar amounts of
libraries were pooled and sequenced on either an Illumina or Nanopore PromethION platform. Genes with a minimum coverage of 10× were included for quantification in both datasets (Figure 10).
Long-read transcriptome data were analyzed to find out what fraction of long sequencing reads cover the full length of the GAPDH gene. Figure 11 demonstrates the strength of long-read RNA sequencing to faithfully identify mRNA (isoform) structure from single tumor cells.
Example 6. Measuring efficiency of capturing cancer cells from blood Methods as described herein were tested to show the ability to capture tumour cells from blood as an example bodily fluid. Therefore, we measured the recovery of cultured tumour cells from blood using four prostate cancer cell lines with different EpCAM-expression (LNCaP = very high, 22Rv1= medium high, PC3-9 = medium low and PC3 = low).
Three independent experiments were performed. In each experiment, using microscopy counting an exact number of tumour cells (range 72-384) from one of the four cell lines was spiked into a different tube from the same donor. Each sample was then supplemented to 14 mL using a buffer consisting of PBS supplemented with BSA, casein and mouse serum (from here on referred to as CellBuffer). After centrifugation at 800RCF the plasma portion was removed and the remaining ~6 mL of sample was supplemented to 9 mL using CellBuffer.
Samples were then incubated with anti-EpCAM coated magnetic particles (Menarini, Bologna, Italy) and subsequently enriched by passage of the sample though a flow chamber (Ibidi, Grafeling, Germany) positioned against a magnetic array. Sample speed was controlled at 0.5 mL/min using a syringe pump (Harvard Apparatus, Holliston, USA). After the sample passed, the flow chamber was washed using 1 mL of CellBuffer at the same flow speed to remove residual uncaptured cells. Next, the magnetic array was removed, and the collected sample was retrieved by flushing the chamber with 2mL CellBuffer and air using a manually operated syringe.
Enriched samples were collected in a 12x75 mm conical tube and magnetically separated for 5 min in an iMAG magnet (BD, Franklin Lakes, USA). The unbound fraction was then aspirated using a glass Pasteur-pipet and the syringe pump set to 2 mL/min.
Samples were washed using 1 mL CellBuffer, separated again and resuspended in 300 μL staining buffer, consisting of 50 μL nuclear stain (Menarini), 50 μL staining reagent (Menarini), 50 μL permeabilization reagent (Menarini), 75 μL CellBuffer and 75 μL PBS.
After 20 minutes of staining at 37°C, 700 μL CellBuffer was added and the sample was magnetically separated for 5 min in an iMAG magnet. After aspiration of the unbound fraction, the sample was resuspended 325 μL buffer consisting of 150 μL CellFix (Menarini) together with 175 μL PBS and placed into an imaging cartridge from the FDA-cleared CellSearch system.
The CellSearch cartridge was then scanned using the CellTracks system, part of the FDA-cleared CellSearch system for CTC enumeration. Tumour cells were identified and counted using the standard software and protocol for tumour cell enumeration (Figure 8).
Claims
Claims
1. A method of preparing a circulating cancer catalog (CCC) of a subject afflicted with or previously afflicted with a solid tumor or is at elevated risk of developing cancer, the method comprising:
performing whole genome sequencing from a healthy sample from said subject, preferably wherein the healthy sample is white blood cells or buccal cells;
providing a bodily fluid sample of the subject, preferably blood, saliva or urine;
isolating a plurality of circulating tumor cells (CTCs) from the bodily fluid sample, preferably at least 2 CTCs;
performing whole genome sequencing of said CTCs, preferably long-read whole genome sequencing;
performing RNA sequencing on RNA of said CTCs, preferably long-read RNA sequencing or long-read sequencing of the corresponding cDNA, and wherein RNA is preferably poly-(A) selected mRNA;
identifying somatic mutations in the genome sequences of said CTCs, said identification comprising for each CTC:
a) comparing the genome sequences from the healthy sample and the CTC to identify potential somatic mutations; and
b) identifying the presence of somatic mutations in said CTC if said potential somatic mutations also occur in the RNA of one or more of said CTCs, and/or if said potential somatic mutations are identified in the genome sequence of at least one other CTC; and
identifying overexpressed tumor antigens and/or tumor-associated antigens (TAAs) expressed by the CTCs,
wherein said somatic mutations and TAAs and/or overexpressed tumor antigens form the CCC.
2. A method of preparing a circulating cancer catalog of a subject afflicted with or previously afflicted with a solid tumor or is at elevated risk of developing cancer, the method comprising the steps of:
performing whole genome sequencing from a healthy sample from said subject, preferably wherein the healthy sample is white blood cells or saliva; providing a bodily fluid sample of the subject, preferably blood, saliva or urine;
isolating a plurality of circulating tumor cells (CTCs) from said sample, preferably at least 2 CTCs;
performing whole genome sequencing of a part of the plurality of CTCs, preferably at least 2 of said CTCs, preferably long-read whole genome sequencing,
performing RNA sequencing of a part of the plurality of CTCs, preferably at least 2 of said CTCs, preferably long-read RNA sequencing or long-read sequencing of the corresponding cDNA,
identifying somatic mutations in the genome sequence of said CTCs, said step comprising:
a) comparing the genome sequences from the healthy sample and the CTC to identify potential somatic mutations; and
b) identifying the presence of somatic mutations in a CTC if said potential somatic mutations occur in the RNA of at least one other CTC, and/or if said potential somatic mutations are identified in the genome sequence of at least one other CTC; and
identifying overexpressed tumor antigens and/or tumor-associated antigens (TAAs) expressed by the CTCs,
wherein said somatic mutations and TAAs and/or overexpressed tumor antigens form the CCC.
3. The method of claim 1 or 2, wherein whole genome sequencing is performed to achieve between 5x to 30x sequencing depth.
4. The method of any one of the preceding claims wherein the identification of TAAs expressed by CTCs comprises comparing the RNA levels of a CTC against the RNA levels of TAAs from a database.
5. The method of any one of the preceding claims further comprising identifying neoantigens encoded by the somatic mutations.
6. The method of any one of the preceding claims, wherein at least 48 CTCs are isolated from a bodily fluid sample.
7. The method of any one of the preceding claims, comprising determining the HLA-type of at least one MHC molecule in said subject.
8. The method of any one of the preceding claims, wherein said method comprises performing single-cell whole genome sequencing and single-cell RNA sequencing of said CTCs.
9. The method of any one of the preceding claims, wherein said method further comprises identifying a treatment for said subject, comprising selecting one or more vaccines for targeting
- one or more TAAs expressed by at least one CTC of the CCC and/or
- one or more overexpressed tumor antigens expressed by at least one CTC of the CCC and/or
- one or more neoantigens encoded by the somatic mutations in at least one CTCs of the CCC.
10. The method of claim 9, wherein the method comprises selecting at least one vaccine that targets a TAA or a neoantigen expressed by at least two CTCs from said subject.
11. The method of any one of claims 1-10, wherein said method further comprises identifying a treatment for said subject, wherein the treatment comprises one or more of the following:
- preparing a chimeric antigen receptor (CAR)-T cell specific for a TAA or neoantigen identified in the CCC;
- preparing natural killer (NK) cells engineered to target one or more TAAs or neoantigens identified in the CCC;
- preparing an antibody or antigen binding fragment thereof specific for one or more TAAs or neoantigens identified in the CCC;
- preparing a nucleic acid vaccine encoding one or more TAAs or neoantigens identified in the CCC.
12. A method for preparing a treatment for a subject having a solid tumor, said method comprising preparing a circulating cancer catalog (CCC) of the subject from a plurality of circulating tumor cells (CTCs) according to any one of claims 1-11 and preparing one or more vaccines for targeting
- one or more TAAs expressed by at least one CTC of the CCC and/or
- one or more overexpressed tumor antigens expressed by at least one CTC of the CCC and/or
- one or more neoantigens encoded by the somatic mutations in at least one CTCs of the CCC; preferably comprising preparing at least one vaccine that targets a TAA, overexpressed antigen or a neoantigen expressed by at least two CTCs from said subject.
13. The method of any one of claims 9-12, wherein the method comprises selecting one or more vaccines for targeting at least two or more TAAs and/or at least two or more neoantigens expressed and/or at least two or more overexpressed tumor antigens by at least one CTC.
14. The method of any one of claims 9-13, wherein the targeting of a TAA, overexpressed tumor antigen or neoantigen comprises identifying and/or preparing at least two different vaccines targeting said TAA, overexpressed tumor antigen or neoantigen, respectively.
15. The method of any one of the preceding claims wherein said tumor-associated antigens (TAAs) are not expressed in healthy tissue of said subject and/or said TAAs are cancer testis antigens.
16. The method of any one of the preceding claims wherein cancer vaccines are identified and/or prepared which target at least 2 CTCs from said subject.
17. A method implemented by one or more processors executing computer program instructions that, when executed, perform the method, the method comprising:
i) preparing a CCC according to any one of the preceding claims for a plurality of subjects afflicted with or previously afflicted with cancer,
ii) providing data related to the plurality of subjects to a prediction model (e.g., a neural network or other machine learning model),
wherein each subject has been administered one or more vaccines targeting a TAA, overexpressed tumor antigen, or neoantigen expressed by one or more CTCs of said subject,
wherein said data comprises for each subject
- the HLA-type of the subject,
- the amino acid sequences of the one or more peptide vaccines administered to said subject,
- the immunogenicity and/or clearance of CTCs expressing the target of said one or more peptide vaccines, following administration of said peptide vaccines; and iii) training the prediction model with said data,
preferably wherein said method further comprises
iv) providing to the trained prediction model, the HLA-type of an individual and the sequence of at least one TAA, overexpressed tumor antigen, or neoantigen expressed in one or more CTC’s of said individual and
v) obtaining from the trained prediction model, one or more peptide sequences predicted to be effective in increasing the immunogenicity and/or clearing CTCs expressing said TAA or neoantigen.
18. A method for detecting circulating tumor DNA (ctDNA) in a bodily fluid sample of a subject afflicted with or previously afflicted with a solid tumor, the method comprising:
- preparing a circulating cancer catalog (CCC) of said subject according to a method of any one of claims 1-11, and
- detecting, in the cfDNA of a bodily fluid sample of the subject, the presence of somatic mutations also found in the subject’s CCC.
19. A method of treating a subject afflicted with or previously afflicted with a solid tumor, comprising:
- preparing a circulating cancer catalog (CCC) of said subject from a plurality of circulating tumor cells (CTCs), said CCC comprising tumor-associated antigens (TAAs), overexpressed tumor antigens, and/or neoantigens expressed by said CTCs according to a method of any one of claims 1-11; and
- preparing and administering to said subject a therapeutic agent that specifically targets at least one TAA, overexpressed tumor antigen or neoantigen identified in said CCC.
20. The method of claim 18 or 19, wherein said therapeutic agent is selected from the group consisting of:
(i) a peptide vaccine;
(ii) a chimeric antigen receptor T (CAR-T) cell;
(iii) a T-cell receptor (TCR)-engineered T cell;
(iv) a natural killer (NK) cell engineered or expanded to target said TAA or neoantigen;
(v) an antibody or antigen binding fragment thereof; and
(vi) a nucleic acid-based vaccine.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24208649.4 | 2024-10-24 | ||
| EP24208649 | 2024-10-24 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2026089611A1 true WO2026089611A1 (en) | 2026-04-30 |
Family
ID=93284062
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/NL2025/050538 Pending WO2026089611A1 (en) | 2024-10-24 | 2025-10-23 | Circulating cancer catalog |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2026089611A1 (en) |
-
2025
- 2025-10-23 WO PCT/NL2025/050538 patent/WO2026089611A1/en active Pending
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12018336B2 (en) | Methods for sequencing samples | |
| Smith et al. | Endogenous retroviral signatures predict immunotherapy response in clear cell renal cell carcinoma | |
| Ott et al. | An immunogenic personal neoantigen vaccine for patients with melanoma | |
| US11098121B2 (en) | “Immune checkpoint intervention” in cancer | |
| Vos et al. | Nivolumab plus ipilimumab in advanced salivary gland cancer: a phase 2 trial | |
| US12461110B2 (en) | Methods for treating breast cancer and for identifying breast cancer antigens | |
| Altorki et al. | Global evolution of the tumor microenvironment associated with progression from preinvasive invasive to invasive human lung adenocarcinoma | |
| Wang et al. | Interactions between LAMP3+ dendritic cells and T-cell subpopulations promote immune evasion in papillary thyroid carcinoma | |
| CN106062561A (en) | Genotypic and phenotypic analysis of circulating tumor cells to monitor tumor evolution in prostate cancer patients | |
| KR20180014769A (en) | Compositions and methods for screening T cells as antigens for a particular population | |
| Feng et al. | Heterogeneity of tumor-infiltrating lymphocytes ascribed to local immune status rather than neoantigens by multi-omics analysis of glioblastoma multiforme | |
| Williamson et al. | Clinical response to nivolumab in an INI1-deficient pediatric chordoma correlates with immunogenic recognition of brachyury | |
| CN109081866B (en) | T cell subsets and their signature genes in cancer | |
| Liu et al. | Immune-featured stromal niches associate with response to neoadjuvant immunotherapy in oral squamous cell carcinoma | |
| Sahin et al. | Individualized mRNA vaccines evoke durable T cell immunity in adjuvant TNBC | |
| Willemsen et al. | Changes in AXL and/or MITF melanoma subpopulations in patients receiving immunotherapy | |
| US20190195858A1 (en) | Separation of Rare Cells and Genomic Analysis Thereof | |
| JP2023510113A (en) | Methods for treating glioblastoma | |
| WO2026089611A1 (en) | Circulating cancer catalog | |
| Ager et al. | Neoadjuvant androgen deprivation therapy with or without Fc-enhanced non-fucosylated anti-CTLA-4 (BMS-986218) in high risk localized prostate cancer: A randomized phase 1 trial | |
| Marc Najjar et al. | 34th Annual Meeting & Pre-Conference Programs of the Society for Immunotherapy of Cancer (SITC 2019): part | |
| WO2021077094A1 (en) | Discovering, validating, and personalizing transposable element cancer vaccines | |
| Obradovic | Discovering Master Regulators of Single-Cell Transcriptional States in the Tumor Immune Microenvironment to Reveal Immuno-Therapeutic Targets and Synergistic Treatments | |
| Roller et al. | 27 Tumor agnostic CD8 immune-phenotype related gene signature defines clinical outcome across early and late phase clinical trials | |
| Thiele | Morphological and Genomic Profiling of Circulating Tumor Cells in Metastatic Colorectal Cancer |