KEDROS Log in Database

Phylogenetic Tree

Filter proteins by Pfam and taxonomy, run alignment and tree inference, view results.

Start from a protein ? Paste a UniProt accession and KEDROS shows the Pfam domains that protein carries, in order along the sequence. Tick the ones you want and they fill the Pfam box below. Useful when you know the protein but not its domain accessions.

Tick two or more and set Multi-Pfam logic to and to require proteins carrying that whole combination.
One or more, comma separated.
1 · Sequences ? A tree can be built from a database query, from your own FASTA, or from both merged into one tree. Fill in whichever you have.

An uploaded FASTA can be either of two things.

Selected sequences — already named {taxid}.{accession}, e.g. >9606.P04637. Used as given.

A whole proteome from an organism that need not be in the database. Give its taxon ID and KEDROS runs hmmsearch over it with the Pfam above, at the same gathering thresholds the database was built with, and keeps only the hits. Protein sequences only — translate a transcriptome first.

Sequences the query already returned are not duplicated.
Headers: {taxid}.{accession}
Optional if you uploaded a FASTA.
2 · Taxonomy filter
Or upload a file (one ID per line)
3 · Alignment & tree
4 · Viewer

Clade actions ? Pick a branch, then choose what to do with it. To pick one, right-click any branch in the ETE4 explorer above and choose “Use branch for divergence” — it is highlighted and appears here. You can select more than one.

Build a tree first. Fill in the form on the left and press Run Tree Pipeline. These actions work on a branch of the tree you build. To pick branches by right-clicking, choose the ETE4 server viewer.
No branch selected — actions below will use the whole tree.

MSA divergence analysis ? Which positions in this family's alignment behave differently between subfamilies? Three scores, computed from the alignment this tree was built from:

Conservation (JSD) — the position is unusually fixed across the whole family. Needs no groups.
Group specificity (GroupSim) — fixed within each subfamily but different between them: the signature of a position that determines what a subfamily does.
Diagnostic power (MI) — how well the residue tells you which subfamily a sequence belongs to. 1.00 means with certainty.

All three down-weight near-identical sequences, so a clade sequenced thirty times does not count thirty times.