独立重比对工具
纵向分割方法
Cut the input sequences into substrings, and once these shorter substrings are realigned, they stitch their alignments together.
1. Refin-Align (2019)
In each iteration:
- Extract the blocks from the initial alignment based on the columns with the same bases.
- Delete all gaps in each block and realign it with Promalign.
- Calculate the SP score of the initial block and the new block.
- If the score of the new one is higher than the initial, the new block will be the input of the next iteration.
Until no longer improve the SP score of each block.
- Test data: Protein (BALIBASE, PREFAB, OXBENCH, HOMSTRAD)
- The source code is not publicized.
2. SpliVert (2020)
- Split the initial alignment vertically into 3 parts.
- Remove the gap characters in the middle part and realign it alone.
- Splice the realigned middle parts with the other two initial pieces to obtain a realigned alignment.
- Test data: Protein (BALIBASE, OXBENCH, SABmark)
- We did not find the source code.
3. RPfam (2022)
RPfam employs the Simulated Annealing algorithm in iteration.
In each iteration:
- Scan the current MSA for badly aligned blocks and compute score of the blocks.
- Randomly pick a badly aligned block.
- Seek the worst aligned fragment in the block.
- Realign the fragment utilizing DP algorithm.
- Update the current MSA.
Until the temperature is down to the setting value the iteration ends.
- Test data: Protein (PFAM)
- The source code is not publicized.
横向分割方法
Split an alignment into groups of whole sequences, which are then merged together by realigning between groups, possibly using each group’s induced subalignment.
1. ReAligner (1997)
In each iteration has a traversal of all sequence:
- Pick a sequence
- Realign the sequence with the profile constructed with the remaining sequences.
- If the new alignment's quality is improved, the new one will replace the old one.
Until the alignment score saturates to a stable value.
- Test data: DNA/RNA
- The source code link is unavailable.
2. Remove First method (2005)
In each iteration:
- Randomly pick a sequence.
- Realign the sequence with the profile constructed with the remaining sequences.
- Update the alignment if the accuracy improved.
Until the score converges or the reach the limit of iterations.
- Test data: Protein (HOMSTRAD)
- Source code
3. TreeRefiner (2005)
This realigner is based on three-dimensional alignment rather than using an iterative approach.
- Test data: Nucleotide sequences that were generated by the Rose program
- TreeRefiner's homepage
4. REFINER (2006)
This realigner is similar to RF method but aims to realign the sequences with the family block models representing conserved sequence/structure regions.
In each iteration:
- Randomly pick a sequence.
- Realign the sequence with the PSSM constructed with the remaining sequences.
- Update the alignment.
Until the alignment score saturates to a stable value or until the iteration cycle terminates.
- Test data: Protein (CDD and PFAM)
- FTP server
5. ReformAlign (2014)
In each iteration:
- Build a profile from the initial alignment.
- All sequences realign with the profile to obtain the reformed alignment. (There has a profile fine-tuning model.)
- The alignment with better quality will be the input of next iteration.
Until the alignment between two successive runs remains unchanged or a predefined maximum number of iterations is reached.
- Test data: DNA/RNA (BRAliBase 2.1 RNA alignment database and DNA SMART database)
- ReformAlign's homepage
横纵组合方法
1. RASCAL (2003)
The alignment is initially divided horizontally and vertically to identify well-aligned, reliable regions. Potential alignment errors are detected by comparing statistical models of the reliable regions. RASCAL then performs a single realignment of each badly aligned region using an algorithm similar to that implemented in ClustalW.
Partitioning steps:
- The alignment is divided horizontally into sequence subfamilies using Secator
- Then the alignment is divided vertically into 'core block' regions that are reliably aligned in the majority of the sequences.
These global core blocks are determined using the mean distance column scores implemented in the NorMD OF. Local core block regions are also determined for each subfamily individually, using the same method as for the complete family.
- Test data: Protein (BALIBASE, ProDom)
- The source code link is unavailable.
2. Crumble and Prune (2011)
[vertical] Crumble deals with problems that are large in sequence length. It breaks up long alignment problems into shorter problems.
[horizontal] Prune handles problems with a large number of sequences. It cuts up deep alignment problems into sub-problems with fewer sequences.
Job-tree combine Crumble and Prune to align long and deep sequences.
- Test data: human genome, Rfam
- The source code link is unavailable.