---
title: "Long Read Prokaryotic Analysis"
canonical: "https://help.biobam.com/space/OBD/2434465807/Long%20Read%20Prokaryotic%20Analysis"
format: markdown
---
# Introduction

Genome analysis of *Bacillus subtilis.*

### Dataset Description

DNA sequencing data from different *Bacillus subtilis* experiments. This dataset comprises a long read sequencing library coming from PacBio technology and a short read sequencing library coming from Illumina technology. 

- Organism: *[Bacillus subtilis](https://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi?mode=Info&id=1423&lvl=3&lin=f&keep=1&srchmode=1&unlock)**.*
- Instrument: Illumina HiSeq 1500 and PacBio RS II.
- Layout: Paired-end (Illumina) and long single reads (PacBio)

### Publication

[Borriss R, Danchin A, Harwood CR, Médigue C, Rocha EPC, Sekowska A, Vallenet D. Bacillus subtilis, the model Gram-positive bacterium: 20 years of annotation refinement. Microb Biotechnol. 2018 Jan;11(1):3-17. doi: 10.1111/1751-7915.13043. PMID: 29280348; PMCID: PMC5743806.](https://pubmed.ncbi.nlm.nih.gov/29280348/)

<details>
<summary>Abstract</summary>

Genome annotation is, nowadays, performed via automatic pipelines that cannot discriminate between right and wrong annotations. Given their importance in increasing the accuracy of the genome annotations of other organisms, it is critical that the annotations of model organisms reflect the current annotation gold standard. The genome of Bacillus subtilis strain 168 was sequenced twenty years ago. Using a combination of inductive, deductive and abductive reasoning, we present a unique, manually curated annotation, essentially based on experimental data. This reveals how this bacterium lives in a plant niche, while carrying a paleome operating system common to Firmicutes and Tenericutes. Dozens of new genomic objects and an extensive literature survey have been included for the sequence available at the INSDC (AccNum AL009126.3). We also propose an extension to Demerec's nomenclature rules that will help investigators connect to this type of curated annotation via the use of common gene names.
</details>

> ℹ️ Note that each sequencing sample comes from a different experiment. The article cited here provides information about the latest advances in the genome of *B. subtillis *genome but is not related to any of the sequencing libraries of this dataset.

### Original Data

- Illumina ENA Run: [ERR2935851](https://www.ebi.ac.uk/ena/browser/view/ERR2935851).
- PacBio ENA Run: [SRR7498042](https://www.ebi.ac.uk/ena/browser/view/SRR7498042).
- NCBI Genome: [Bacillus subtilis subsp. subtilis str. 168](https://www.ncbi.nlm.nih.gov/genome/665?genome_assembly_id=300274).

# Bioinformatic Analysis

## 1- DNA-Seq *de novo* Assembly

### Application

[DNA-Seq ](https://biobam.atlassian.net/wiki/spaces/OBD/pages/2489516151)*[de novo](https://biobam.atlassian.net/wiki/spaces/OBD/pages/2489516151)*[ Assembly](https://biobam.atlassian.net/wiki/spaces/OBD/pages/2489516151) (Flye).

### Input

- [PacBio sequencing data in FASTQ Format (SRR7498042.fastq.gz).](https://drive.google.com/file/d/1QsHPsiUo3Y_h6X3FUWHuyWrFKUjI_txQ/view?usp=sharing)

### Parameters

- Input Reads: [PacBio Raw] SRR7498042.fastq.gz
- Provide Genome Size: false
- Automatic Minimum Overlap: true
- Polishing: true
- Number of Polishing Iterations: 1
- Plasmids: false
- Keep Haplotypes: false
- Trestle: true
- Assembly Fasta: assembly.fasta
- Save Graph File: true
- Graph File: assembly_graph.gfa

### Execution Time

30-40 minutes.

### Output

- [assembly.fasta: FASTA file containing assembled sequences.](https://drive.google.com/file/d/10npIwv16aU9Fri7WFFUv8qSYUd6Kkp-L/view?usp=sharing)
- [assembly_graph.gfa: Text file containing the repeat graph in GFA format.](https://drive.google.com/file/d/1uQ6_jqD80eRrRo0yGq9p_5vSF9b39Hr9/view?usp=sharing)
- [nx_plot_flye.box: Line chart about DNA-Seq](https://drive.google.com/file/d/1C589z8G5GLuBTlpKw4196l174eVQQD-Q/view?usp=sharing)*[ de novo ](https://drive.google.com/file/d/1C589z8G5GLuBTlpKw4196l174eVQQD-Q/view?usp=sharing)*[Assembly results.](https://drive.google.com/file/d/1C589z8G5GLuBTlpKw4196l174eVQQD-Q/view?usp=sharing)
- [report_flye.box: Report about DNA-Seq ](https://drive.google.com/file/d/1TKJIsbf49wHDsotY_jHrDGVFedFzpsTg/view?usp=sharing)*[de novo ](https://drive.google.com/file/d/1TKJIsbf49wHDsotY_jHrDGVFedFzpsTg/view?usp=sharing)*[Assembly results.](https://drive.google.com/file/d/1TKJIsbf49wHDsotY_jHrDGVFedFzpsTg/view?usp=sharing)

## 2 - DNA-Seq Alignment

### Application

[DNA-Seq Alignment](https://biobam.atlassian.net/wiki/spaces/OBD/pages/2488008844) (BWA).

### Input

- Illumina sequencing data in FASTQ format: [ERR2935851_1.fastq.gz](https://drive.google.com/file/d/1oPr8L5P8ss8NKZHwt7oBA-5T1AEGWVW4/view?usp=sharing) and [ERR2935851_2.fastq.gz](https://drive.google.com/file/d/1-g_jV2YHWgoGMnyo9fnppEfVIeOgjfSj/view?usp=sharing).
- [Assembled genome](https://drive.google.com/file/d/10npIwv16aU9Fri7WFFUv8qSYUd6Kkp-L/view?usp=sharing) in FASTA format (from the [1- DNA-Seq de novo Assembly](https://biobam.atlassian.net/wiki/spaces/OBD/pages/2434465807/Long+Read+Prokaryotic+Analysis#1--DNA-Seq-de-novo-Assembly) step).

### Parameters

### Execution Time

15-20 minutes.

### Output

- [ERR2935851.bam: BAM file containing short read alignments.](https://drive.google.com/file/d/14z8RNvHCV95A05jhrdo4f6ai-DRn3o0h/view?usp=sharing)
- [alignments_per_category_bwa.box: Stacked bar plot about the DNA-Seq Alignments results.](https://drive.google.com/file/d/1ucHq9Q3iTF8sIxr1KdYUyRsNxZ08QFRM/view?usp=sharing)
- [relative_alignments_per_category.box: Stacked bar plot about the DNA-Seq Alignment results (percentages).](https://drive.google.com/file/d/1i5KNkNw5j_sZzM-DmOdfIMmjsd0jhlMA/view?usp=sharing)
- [report_bwa.box: Report about the DNA-Seq Alignment results.](https://drive.google.com/file/d/1DrsEHATeLFVN20nRmdCW1hoyqvP0cmnL/view?usp=sharing)

## 3- Polishing

### Application

[DNA-Seq Polishing](https://biobam.atlassian.net/wiki/spaces/OBD/pages/1117356121).

### Input

- [Assembled genome](https://drive.google.com/file/d/10npIwv16aU9Fri7WFFUv8qSYUd6Kkp-L/view?usp=sharing) (from the [1- DNA-Seq de novo Assembly](https://biobam.atlassian.net/wiki/spaces/OBD/pages/2434465807/Long+Read+Prokaryotic+Analysis#1--DNA-Seq-de-novo-Assembly) step).
- [DNA-Seq Alignments](https://drive.google.com/file/d/14z8RNvHCV95A05jhrdo4f6ai-DRn3o0h/view?usp=sharing) in BAM format (from the [2- DNA-Seq Alignment](https://biobam.atlassian.net/wiki/spaces/OBD/pages/2434465807/Long+Read+Prokaryotic+Analysis#2---DNA-Seq-Alignment) step).

### Parameters

### Execution Time

10-15 minutes.

### Output

- [polished_sequences.fasta: FASTA file containing polished sequences.](https://drive.google.com/file/d/1wqcIiPoHih4j7rzMROIfaZtDUPepZAUO/view?usp=sharing)
- [changes.txt: Text file containing a space-delimited record of every change made in the assembly. ](https://drive.google.com/file/d/1R-5HX84wSkb4aBBohMi41X0F3tEnj793/view?usp=sharing)
- [dna_seq_polishing_results.box: Report about DNA-Seq Polishing results.](https://drive.google.com/file/d/1b_TZzVev8f7Y9mf2RoBeuiUFNwi4LyYD/view?usp=sharing)
- [nx_plot_polishing.box: Line chart about DNA-Seq Polishing results.](https://drive.google.com/file/d/171ttWp0CgEjjqEFSKSI0ToLaZI4eqfZ2/view?usp=sharing)
- [fix_type_distribution.box: Pie chart that summarizes the fix types performed during polishing.](https://drive.google.com/file/d/1SPub2aw2rsNfOGHXpf8PWQo5F_ytKGw7/view?usp=sharing)

## 4- Gene Finding

### Application

[Prokaryotic Gene Finding by Glimmer](https://biobam.atlassian.net/wiki/spaces/OBD/pages/598147189).

### Input

- [Polished assembly](https://drive.google.com/file/d/1wqcIiPoHih4j7rzMROIfaZtDUPepZAUO/view?usp=sharing) in FASTA format (from the [3- Polishing](https://biobam.atlassian.net/wiki/spaces/OBD/pages/2434465807/Long+Read+Prokaryotic+Analysis#3--Polishing) step).

### Parameters

- Input Sequences:polished_sequences.fasta
- Select the genetic code: The Bacterial, Archaeal and Plant Plastid Code
- Minimum gene length: 110
- Maximum gene overlap: 30
- Minimum gene score: 30
- Select the genome shape: Circular
- Choose ICM use option: Create new ICM model
- Set advanced ICM parameters: false
- Save ICM model: false
- Select the run model: Iterated
- Define the start codons: false
- Define the stop codons: false
- Define GC content: false

### Execution Time

~ 5 min

### Output

- [seqs_pgf.box: Sequence project containing predicted gene sequences.](https://drive.google.com/file/d/12a1UMmpmyuWNBeGAvcQcO2WHwidSOjdG/view?usp=sharing)
- [gff_pgf.box: GFF project containing predicted gene coordinates.](https://drive.google.com/file/d/1za0VSa69r5oOMq37LayLotUJoLwa0V54/view?usp=sharing)
- [report_pgf.box: Report about Prokaryotic Gene Finding results.](https://drive.google.com/file/d/1W8s6bbizZrAiUS1giUnVDa2qQM-qeV_a/view?usp=sharing)

## 5- BLAST & InterProScan

### Application:

[CloudBLAST](https://biobam.atlassian.net/wiki/spaces/OBD/pages/3335684097) & [InterProScan Annotation](https://biobam.atlassian.net/wiki/spaces/OBD/pages/598048967).

### Input:

- [Predicted gene sequences](https://drive.google.com/file/d/12a1UMmpmyuWNBeGAvcQcO2WHwidSOjdG/view?usp=sharing) in an OmicsBox project (from the [4- Gene Finding](https://biobam.atlassian.net/wiki/spaces/OBD/pages/2434465807/Long+Read+Prokaryotic+Analysis#4--Gene-Finding) step).

### Parameters

### Execution Time

~1 hour.

### Output

- [blast_ips_genes.box: Project containing predicted gene sequences along with Blast and IPS results.](https://drive.google.com/file/d/1yQ-ukXqXKpSErq2AfsftSBlcg_Il9t-t/view?usp=sharing)

## Workflow


![image](media://f2309e58-b91b-46ed-9969-2bb10b3256ca)