Here is described how to compare two bacterial typing methods. As an example, the workflow compares whole-genome sequencing (WGS) based cgMLST clustering with clustering results generated by IR Biotyper. The aim is to calculate the discriminatory indices for each typing method and assess concordance between the methods using MBioSEQ Ridom Typer.

In brief, the workflow involves separately importing the groups assignments generated by the IR Biotyper and the cluster assignments derived from the cgMLST Minimum Spanning Tree analysis. The comparison and statistical analysis are then performed using the tools available within MBioSEQ Ridom Typer.


Prerequisites

This comparison requires that the same isolates are present in both typing method datasets using identical sample IDs. Matching sample identifiers are essential because the import and comparison process relies on directly linking isolates between the two datasets.

Required input data:

  • A dataset containing isolate cluster assignments generated by the first typing method, i.e., IR Biotyper, provided as either a CSV file or an Excel spreadsheet.
  • A dataset containing isolate cluster assignments generated by the second typing method, i.e., WGS-based cgMLST clustering. The sample data should already be imported into MBioSEQ Ridom Typer and cgMLST analysis should already be completed.


Importing group assignments from first typing method, i.e., IR Biotyper

Group assignments can be imported using either a CSV file or an Excel spreadsheet containing at least two columns: unique sample IDs and the corresponding group assignments generated by the IR Biotyper. The file can be imported using the Import Epi Metadata function. During the import process, please assign the IR Biotyper cluster information to the Cluster/Outbreak (Epi Source) field.


The Import Epi Metadata function is accessible from the menu bar, as shown in the figure below.

 


Fn import epimetadata.png

 


Choose the file that should be imported here and click Open.

 


File epimetadata import.png

 


The data should be imported into the project containing the cgMLST samples by selecting the correct project from Import into project drop-down menu. The IR-BT cluster assignments can be mapped to the Cluster/Outbreak (Epi Source) and sample ID fields. Select the option indicating that the first row contains column headers. Once all done, click OK.

 


Column mapping clusters epi metadata.png

 

Group assignments from second typing method, i.e., WGS-based cgMLST clustering

Group assignments can be generated from cgMLST data by generating minimum spanning tree (MST). To begin, create a comparison table containing the samples that should be analyzed.

 


Comparison table acco.png

 


Next, generate the MST from the comparison table. In the MST settings, define the clustering distance, which is used to determine how groups are formed. By default, the CT threshold distance is applied to define clusters. This threshold can be adjusted in the MST settings under Maximum distance in MST cluster, specify the appropriate clustering distance (e.g. CT distance). Additionally, set the Minimum number of Samples/Nodes parameter to 1, to ensure that even 'singletons' get assigned to a cluster.

 


{{{2}}}

 

 


Mst options.png

 


In the MST menu, select View > Create Groups for Clusters. The clustering information will be added to the table, and the rows will be color-coded according to the MST clusters. Export these clusters using the menu options (File > Export Table Data). In the export dialog, please select the Remove Target Columns and Export Color Grouping Column options and click Save.

 


Export comparison table data.png

 

 


Export comparison table options.png

 

The clusters from the MST analysis are now included in the exported file. The file can next be dragged into the MBioSEQ Ridom Typer Client window to open it as a comparison table. In the import dialog, a column name for MST cluster information which is generally the far right column can be assigned, for example MST clusters. This name will be used as the header of a new column in the comparison table. In this way, the imported clustering or metadata information becomes directly available within the comparison table and can be used for further analysis.

 


Drag table data into RT.png

 

 


Drag table data into RT col name.png

 

Statistical analysis on the two typing methods

Calculating Discriminatory Indices

The discriminatory index (DI) for each typing method can be calculated by selecting Tools > Calculate Discriminatory Index from the menu of the comparison table saved in previous step. A discriminatory index value closer to 1 indicates that the typing system has greater discriminatory power, while a value closer to 0 indicates lower discriminatory power. These values can be directly compared between typing systems, with the higher value representing the more discriminatory typing method. Please note that usually only the DIs relative to each other and not the absolute value should be judged.

 


Calc discriminatory index.png

 


Please select the appropriate column (e.g., MST clusters for the first or Cluster/Outbreak (Source) column for the second typing method) that can be used to calculate discriminatory index. Before doing so, make sure to click Select None to deselect all other columns.

 


Options calc discriminatory index.png

 


The calculated discriminatory index for a typing method is shown including the 95% confidence intervals. The details on these calculations are on this page.

 


Discriminatory index result.png

 

Calculating Typing System Concordances

The typing system concordance can be calculated by selecting Tools > Calculate Typing System Concordance from the menu of the comparison table. Please go through the detailed documentation on how to interpret these values here.

 


Calc typing cocord acco.png

 

The concordance test is directional. Please choose the two columns with the group assignments in the correct order: first, the gold standard group (here the cgMLST-derived clusters from the minimum spanning tree labeled MST clusters), followed by the IR Biotyper results under Cluster/Outbreak from the previous step. This will calculate the concordance between the two clustering methods (cgMLST vs. IR Biotyper).

 


Calc typing cocord choose columns.png

 


The concordance between the two clustering methods is now displayed. Ensure that the results do not display the message Rejected X samples because of incomplete information.

 


Concordance result.png

 

References

Simpson, E.H. Measurement of diversity. Nature 1949, 163:688 [Nature 163688a0]

Hunter, P.R., and Gaston, M.A. Numerical index of the discriminatory ability of typing systems: an application of Simpson’s index of diversity. J. Clin. Microbiol. 1988, 26: 2465–2466 [PubMed 3069867]

Grundmann, H., Hori, S., Tanner, G. Determining confidence intervals when measuring genetic diversity and the discriminatory abilities of typing methods for microorganisms. J. Clin. Microbiol. 2001, 39: 4190-4192 [PubMed 11682558]

Carriço J.A., Silva-Costa C., Melo-Cristino J., Pinto F.R., de Lencastre H., Almeida J.S., Ramirez M. Illustration of a common framework for relating multiple typing methods by application to macrolide-resistant Streptococcus pyogenes J. Clin. Microbiol. 2006, 44: 2524-32 [PubMed 16825375]

 
FOR RESEARCH USE ONLY. NOT FOR USE IN CLINICAL DIAGNOSTIC PROCEDURES.