---
title: "De Novo Transcriptome Characterization"
canonical: "https://help.biobam.com/space/OBD/2377941020/De%20Novo%20Transcriptome%20Characterization"
format: markdown
---
# Introduction

Transcriptome Analysis of *Monilinia laxa.*

### Dataset Description

Transcriptomes of *Monilinia fructicola*, ***Monilinia laxa****,* and *Monilinia fructigena*, the causal agents of brown rot of stone and pome fruits. For this tutorial, only the data of *Monilinia laxa* is used. This dataset comprises paired-end reads that were corresponding to mycelium grown in the dark for 4 days, mycelium grown in the dark for 2 days, and then exposed to light for 2 days, as well as in germinating conidia (2 replicates per each condition). 

- Organism: *[Monilinia laxa](https://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi?mode=Info&id=61186&lvl=3&lin=f&keep=1&srchmode=1&unlock)*
- Instrument: Illumina HiScanSQ
- Layout: Paired-end

### Publication

[De Miccolis Angelini RM, Abate D, Rotolo C, Gerin D, Pollastro S, Faretra F. De novo assembly and comparative transcriptome analysis of Monilinia fructicola, Monilinia laxa and Monilinia fructigena, the causal agents of brown rot on stone fruits. BMC Genomics. 2018 Jun 5;19(1):436. doi: 10.1186/s12864-018-4817-4. PMID: 29866047; PMCID: PMC5987419.](https://bmcgenomics.biomedcentral.com/articles/10.1186/s12864-018-4817-4)*[Monilinia fructicola](https://bmcgenomics.biomedcentral.com/articles/10.1186/s12864-018-4817-4)*[, ](https://bmcgenomics.biomedcentral.com/articles/10.1186/s12864-018-4817-4)*[Monilinia laxa](https://bmcgenomics.biomedcentral.com/articles/10.1186/s12864-018-4817-4)*[ and ](https://bmcgenomics.biomedcentral.com/articles/10.1186/s12864-018-4817-4)*[Monilinia fructigena](https://bmcgenomics.biomedcentral.com/articles/10.1186/s12864-018-4817-4)*[, the causal agents of brown rot on stone fruits.](https://bmcgenomics.biomedcentral.com/articles/10.1186/s12864-018-4817-4)

<details>
<summary>Abstract</summary>

Brown rots are important fungal diseases of stone and pome fruits. They are caused by several *Monilinia* species but *M. fructicola*, *M. laxa,* and *M. fructigena* are the most common all over the world. Although they have been intensively studied, the availability of genomic and transcriptomic data in public databases is still scant. We sequenced, assembled, and annotated the transcriptomes of the three pathogens using mRNA from germinating conidia and actively growing mycelia of two isolates of opposite mating types per species for comparative transcriptome analyses.
</details>

### Original Data

- NCBI BioProject: [PRJNA419302](https://www.ncbi.nlm.nih.gov/bioproject/PRJNA419302)
- SRA Experiments: [SRR6312174, SRR6312175, SRR6312181, SRR6312182, SRR6312187, and SRR6312190](https://www.ncbi.nlm.nih.gov/Traces/study/?acc=SRP125382&o=acc_s%3Aa).
- NCBI Nucleotide: [TSA: ](https://www.ncbi.nlm.nih.gov/nuccore/GGAL00000000.1)*[Monilinia laxa](https://www.ncbi.nlm.nih.gov/nuccore/GGAL00000000.1)*[, transcriptome shotgun assembly.](https://www.ncbi.nlm.nih.gov/nuccore/GGAL00000000.1)

# Bioinformatic Analysis

## 1- RNA-Seq *de novo *Assembly

### Application

[RNA-Seq ](https://biobam.atlassian.net/wiki/spaces/OBD/pages/1420197922)*[de novo](https://biobam.atlassian.net/wiki/spaces/OBD/pages/1420197922)*[ Assembly](https://biobam.atlassian.net/wiki/spaces/OBD/pages/1420197922) (Transcriptomics). 

### Input

- [Sequencing data in FASTQ format (SRR6312174, SRR6312175, SRR6312181, SRR6312182, SRR6312187, and SRR6312190). ](https://drive.google.com/drive/u/1/folders/1v7H0B3TmUOhcHqLb_B-gb03MGy1ThIi_)

#### Parameters

### Execution Time

~ 2 hours. 

### Output

- [transcripts.box: Sequence project containing assembled transcripts](https://drive.google.com/file/d/1LFyv0cNGuI3970wL3SNRDlt-zLH3GJIz/view?usp=sharing).
- [rna_seq_de_novo_assembly_report.box: Report about the RNA-Seq de novo assembly results](https://drive.google.com/file/d/18lDgxkRC83u128N3Is7et_AEl8PH-zGO/view?usp=sharing).
- [transcript_to_gene_map.txt: Tab-delimited file with the information to map from transcript (isoform) identifiers to gene identifiers](https://drive.google.com/file/d/1Zo2fAqHGHvAcwYkySmkxDBRFrJvECgAj/view?usp=sharing).

## 2- Completeness Assessment

### Application

[Completeness Assessment](https://biobam.atlassian.net/wiki/spaces/OBD/pages/761528371) (Transcriptomics).

### Input

- [Assembled transcripts](https://drive.google.com/file/d/1LFyv0cNGuI3970wL3SNRDlt-zLH3GJIz/view?usp=sharing) (from the [1- RNA-Seq de novo Assembly](https://biobam.atlassian.net/wiki/spaces/OBD/pages/2377941020/De+Novo+Transcriptome+Characterization#1--RNA-Seq-de-novo-Assembly) step).

### Parameters

- Lineage: Helotiales (order).
- Mode: Transcriptome.
- Blast e-value: 1.0E-3.

### Execution Time

10-15 minutes.

### Output

- [ca_results_transcripts: Project containing BUSCO results.](https://drive.google.com/file/d/1MriGv9oqI-WmjBpmaZB0yl5wE9i-aW7P/view?usp=sharing)

## 3- Clustering

### Application

[Clustering ](https://biobam.atlassian.net/wiki/spaces/OBD/pages/1117388844)(Transcriptomics). 

### Input

- [Assembled transcripts](https://drive.google.com/file/d/1LFyv0cNGuI3970wL3SNRDlt-zLH3GJIz/view?usp=sharing) (from the [1- RNA-Seq de novo Assembly](https://biobam.atlassian.net/wiki/spaces/OBD/pages/2377941020/De+Novo+Transcriptome+Characterization#1--RNA-Seq-de-novo-Assembly) step).

### Parameters

### Execution Time

10-15 minutes.

### Output

- [clustering_transcripts.box: Sequence project containing the representative sequence of each cluster.](https://drive.google.com/file/d/1Ha9s8D1ekS9kOok6SJrnxynKgQRt3aQP/view?usp=sharing)
- [clustering_results.box: Report about the clustering results.](https://drive.google.com/file/d/1q5erabjqiszY949PtVzv2LMTMlM0ie5g/view?usp=sharing)
- [cluster_distribution_transcripts.box: Bar plot showing the number of clusters of each cluster size. ](https://drive.google.com/file/d/1BXKKu7MOrhhVFk8uoMet_2iB3hZeri1p/view?usp=sharing)
- [clusters.txt: This is a text file generated by CD-HIT containing information about each cluster.](https://drive.google.com/file/d/1DP-1tZLw09Y5zlp2WdT_t2RrWSka0nVm/view?usp=sharing)
- [ca_results_clustering_transcripts.box: BUSCO assessment for the results of the clustering step.](https://drive.google.com/file/d/1Rm2whQGHHTF57fPzWyOBZEMNjltKuRaG/view?usp=sharing)

## 4- Predict Coding Regions

### Application

[Predict Coding Regions](https://biobam.atlassian.net/wiki/spaces/OBD/pages/761692214) (Transcriptomics). 

### Input

- [Clustered transcripts](https://drive.google.com/file/d/1Ha9s8D1ekS9kOok6SJrnxynKgQRt3aQP/view?usp=sharing) (from the [3- Clustering](https://biobam.atlassian.net/wiki/spaces/OBD/pages/2377941020/De+Novo+Transcriptome+Characterization#3--Clustering) step).

### Parameters

- Genetic Code: Universal
- Minimum Protein Length: 100
- Strand Specific: false
- Provide Gene-Transcript Relationships: false
- Pfam Search: true (recommended, but time-consuming)
- Retain Long ORFs Mode: Dynamic
- Single Best Only: true
- No Refine Starts: false
- Top Longest ORFs for Training: 500

### Execution Time

45 minutes (10-15 minutes without Pfam Search).

### Output

- [cds.box: Sequence project containing nucleotide sequences for coding regions of the final candidate ORFs.](https://drive.google.com/file/d/11IB6kMwkaACU_wM9szvr1Z9dCZbxK8id/view?usp=sharing)
- [protein.box: Sequence project containing peptide sequences for the final candidate ORFs.](https://drive.google.com/file/d/1Vh5aSm4PurPTiMP2MbjpmZ6PdK551FM6/view?usp=sharing)
- [coordinates.box: GFF project that contains positions within the target transcripts of the final selected ORFs.](https://drive.google.com/file/d/1yqkWfITl_QL3wRAi2PGbQTdWnfjNrsNB/view?usp=sharing)
- [orf_types.box: Pie chart that shows the percentage of ORFs that have been predicted as Complete, 5' Partial, 3' Partial, and Internal.](https://drive.google.com/file/d/12cLCA5vOPnZtPkAVyeZiF8252e5bfqaY/view?usp=sharing)
- [predict_coding_regions_results.box: Report about the ‘predict coding regions’ results.](https://drive.google.com/file/d/1r1qHhys8uA121Rz1nbnDMsVeS0_u4u6U/view?usp=sharing)
- [ca_results_protein.box: BUSCO assessment for the results of the predict coding regions step (proteins).](https://drive.google.com/file/d/1OXWdMuexBZHLvItsBIgDNvC3RFrXExg5/view?usp=sharing)

## 5- Functional Annotation

### Application

[Functional Annotation Pipeline](https://biobam.atlassian.net/wiki/spaces/OBD/pages/598278164) (Functional Analysis). 

### Input

- [Predicted proteins](https://drive.google.com/file/d/1Vh5aSm4PurPTiMP2MbjpmZ6PdK551FM6/view?usp=sharing) (from the [4- Predict Coding Regions](https://biobam.atlassian.net/wiki/spaces/OBD/pages/2377941020/De+Novo+Transcriptome+Characterization#4--Predict-Coding-Regions) step).

### Parameters

#### Merge InterProScan GOs to Annotation

### Execution Time

3 hours with IPS Scan, less than 1 hour without IPS Scan. 

### Output

- [protein.box: Sequence project containing annotated CDS sequences.](https://drive.google.com/file/d/1LBBweuwUdyrqlhKrZu2jy_Hsu57YSNhG/view?usp=sharing)
- [merge_interpro_annotation.box: Bar plot that summarizes the results after merging IPS GOs to annotation. ](https://drive.google.com/file/d/1M7aMDr6EGVmsUJkstV7ixHDitNKaelKh/view?usp=sharing)

## Workflow

![image](media://c96f6ef0-35bf-4952-bc3a-d2be2e5b222e)