---
title: "Long-Read Isoform Identification with FLAIR"
canonical: "https://help.biobam.com/space/OED0324/3525084750/Long-Read%20Isoform%20Identification%20with%20FLAIR"
format: markdown
---
# Introduction

Isoform Identification in *[Apostichopus japonicus](https://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi?id=307972)**.*

### Dataset Description

This dataset contains one file of long reads sequenced by PacBio Sequel technology and two BAM files generated using [STAR](https://manual.omicsbox.biobam.com/user-manual/omicsbox-modules/module-transcriptomics/rna-seq-alignment/rna-seq-star/) from two pairs of paired-end FASTQ files with short reads sequenced by Illumina HiSeq 2500 technology.

- Organism: *[Apostichopus japonicus](https://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi?id=307972)**.*
- Instrument: PacBio Sequel and Illumina HiSeq 2500.

### Publication

[Wang, Y., Yin, Y., Cong, X., Storey, K.B. and Chen, M., 2022. PacBio Isoform Sequencing and Illumina RNA Sequencing Provide Novel Insights on Responses to Acute Heat Stress in Apostichopus japonicus Coelomocytes. Frontiers in Marine Science, 8, p.815109.](https://www.frontiersin.org/articles/10.3389/fmars.2021.815109/full)

<details>
<summary>Abstract</summary>

Significant increases in global sea surface temperatures are expected with climate change and may cause a serious challenge for marine organisms cultured in aquatic environments that are characterized by short and long-term fluctuations in water temperatures. *Apostichopus japonicus*, a sea cucumber with high nutritional value and pharmacological properties, is an important economic species that is widely raised in aquaculture in China. In recent years, continuous extreme high temperatures (up to 30°C) have occurred frequently in summer leading to mass mortality of sea cucumbers cultured in semi-open shallow regions seriously restricting the sustainable development of sea cucumber aquaculture. In the present study, we combined RNA-seq and PacBio single-molecular real-time (SMRT) sequencing technology to unveil the potential mechanisms of response to acute heat stress in *A. japonicus* coelomocytes. A total of 1,375 differentially expressed genes (DEGs) were identified in a comparison of control and 48 h heat stress (HS) groups.
</details>

### Original Data

- PacBio NCBI Project: [PRJNA785124](https://www.ncbi.nlm.nih.gov/bioproject/PRJNA785124).
- Illumina NCBI Project: [PRJNA687597](https://www.ncbi.nlm.nih.gov/bioproject/PRJNA687597).
- NCBI Genome and Annotation:[ Apostichopus japonicus](https://www.ncbi.nlm.nih.gov/genome/12044?genome_assembly_id=352579).

# Bioinformatic Analysis

## 1- Long-Read Alignment using Minimap2

### Input

- [PacBio Long-Reads dataset](https://drive.google.com/file/d/15zSdbAD9ObGOBAcQ1KshIRBAMs9ONYAF/view?usp=share_link) in FASTQ format.
- [NCBI genome](https://drive.google.com/file/d/1TPsY_fRUKQbl6xgdFLVYtn46xSkbADHy/view?usp=share_link) in FASTA format.

### Parameters

#### Execution Time

3 hours 23 minutes.

### Output

- [SRR17083776.bam](https://drive.google.com/file/d/1RlrsQzPTJIwfu7A_yIentqvrRoCQXXM0/view?usp=sharing): Alignment file.
- [long_read_alignment_with_minimap2.box](https://drive.google.com/file/d/12rUjF42rfQ2DevtrC7OAn1fvnZKG_6Lm/view?usp=drive_link): Summary report of the reference genome and the alignment file.
- [alignments_per_category.box](https://drive.google.com/file/d/1Z1SVGPIYDYZ2MylTV3czsi9hkYI-LPDr/view?usp=drive_link): Aligned vs unaligned reads per sample.
- [relative_alignments_per_category.box](https://drive.google.com/file/d/10JJg46lQMCH1TqWmsOhQLha_Cz-lW90-/view?usp=sharing): The same but with relative values.

## 2- Long-Read Isoform Identification using FLAIR

### Application

Long-Read Isoform Identification (FLAIR).

### Input

- [PacBio Long-Reads dataset](https://drive.google.com/file/d/15zSdbAD9ObGOBAcQ1KshIRBAMs9ONYAF/view?usp=share_link) in FASTQ format.
- [Aligned Long Reads in BAM format](https://drive.google.com/file/d/1RlrsQzPTJIwfu7A_yIentqvrRoCQXXM0/view?usp=sharing).
- [Aligned Short Reads](https://drive.google.com/drive/folders/1vAFF741dM549LV4lEyc8Pg3AJi3vWgPN?usp=share_link) in BAM format using [STAR](https://manual.omicsbox.biobam.com/user-manual/omicsbox-modules/module-transcriptomics/rna-seq-alignment/rna-seq-star/).
- [NCBI genome](https://drive.google.com/file/d/1TPsY_fRUKQbl6xgdFLVYtn46xSkbADHy/view?usp=share_link) in FASTA format.
- [Annotation File](https://drive.google.com/file/d/1KygTICVKhMZrl37i4Otobe8PATwXgoVm/view?usp=share_link) in GTF format.
- [Reads Manifest](https://drive.google.com/file/d/1ZM1zQdiNmOyfdJnj9XNrwNzz7e-syyMl/view?usp=share_link) in TSV format to quantify final isoforms.

### Parameters

#### Execution Time

2 hours 23 minutes.

### Output

- [flair.transcriptome.gtf](https://drive.google.com/file/d/1P_fS798xS4-wiw18ZIwa5zFLTRvvSdqM/view?usp=drive_link): Transcriptome Annotation in GTF format. It can be the input to SQANTI3.
- [flair.transcriptome.fa](https://drive.google.com/file/d/1CSB2KWKk8ORU8PqFRw-r_PswnJXZXAr8/view?usp=drive_link): Transcriptome Sequences in FASTA format.
- [flair.map.txt](https://drive.google.com/file/d/13G_pw1YuF_42b6S0c_QfdrWhbNK12Dmn/view?usp=drive_link): Isoform-Read relationships.
- [flair.counts.tsv](https://drive.google.com/file/d/1clxGAJ9p7JEJow3_T9YLBCQVKT9mXS8H/view?usp=drive_link): Counts File of each discovered isoform.
- [flair_report.box](https://drive.google.com/file/d/1y9qDB8tZpf5Pa9UShhG-jlxCm89PM2aW/view?usp=sharing): Summary Report
- [isoforms_length.box](https://drive.google.com/file/d/1Z37EmJlWP2TMsnzWCmmkuAHH1lfuuUwW/view?usp=drive_link): Isoform Length Distribution