Skip to contents

Complete Results

These results are based on Stanley (2017) data-generating mechanism with a total of 324 conditions.

Average Performance

Method performance measures are aggregated across all simulated conditions to provide an overall impression of method performance. However, keep in mind that a method with a high overall ranking is not necessarily the “best” method for a particular application. To select a suitable method for your application, consider also non-aggregated performance measures in conditions most relevant to your application.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Mean Rank Rank Method Mean Rank
1 RoBMA (PSMA) 5.571 1 RoBMA (PSMA) 5.951
2 MMPH (default) 7.951 2 AK (AK1) 8.151
3 AK (AK1) 8.429 3 SM (3PSM) 8.429
4 SM (3PSM) 8.543 4 MMPH (default) 9.093
5 FMA (default) 10.204 5 FMA (default) 10.040
6 WLS (default) 10.216 6 WLS (default) 10.052
7 WAAPWLS (default) 10.299 7 WAAPWLS (default) 10.108
8 puniform (star) 11.228 8 puniform (star) 11.096
9 WILS (default) 11.269 9 WILS (default) 11.123
10 SM (4PSM) 11.361 10 SM (4PSM) 11.225
11 RMA (default) 12.148 11 trimfill (default) 12.102
12 trimfill (default) 12.287 12 RMA (default) 12.151
13 PETPEESE (default) 12.420 13 PETPEESE (default) 12.256
14 PEESE (default) 12.475 14 PEESE (default) 12.293
15 EK (default) 13.525 15 EK (default) 13.568
16 PET (default) 13.586 16 PET (default) 13.630
17 AK (AK2) 14.031 17 AK (AK2) 13.784
18 MAIVE (default) 14.302 18 MAIVE (default) 14.235
19 pcurve (default) 15.377 19 pcurve (default) 15.367
20 MAN (default) 16.003 20 RTMA (relaxed) 15.972
21 RTMA (relaxed) 16.173 21 puniform (default) 16.086
22 puniform (default) 16.176 22 MAN (default) 16.488
23 mean (default) 16.540 23 mean (default) 16.713
24 MAIVE (WAIVE) 17.293 24 MAIVE (WAIVE) 17.494

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average RMSE is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of RMSE values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Mean Rank Rank Method Mean Rank
1 SM (3PSM) 7.821 1 SM (3PSM) 7.728
2 SM (4PSM) 8.238 2 SM (4PSM) 8.302
3 RoBMA (PSMA) 8.833 3 RoBMA (PSMA) 9.222
4 AK (AK1) 9.549 4 AK (AK1) 9.454
5 MAIVE (default) 9.917 5 puniform (star) 10.090
6 puniform (star) 10.127 6 MAIVE (default) 10.093
7 PETPEESE (default) 10.414 7 PETPEESE (default) 10.562
8 EK (default) 10.528 8 EK (default) 10.735
9 PET (default) 10.537 9 PET (default) 10.744
10 WAAPWLS (default) 11.309 10 WAAPWLS (default) 11.309
11 PEESE (default) 11.719 11 PEESE (default) 11.744
12 MMPH (default) 12.454 12 WLS (default) 12.623
13 WLS (default) 12.481 13 FMA (default) 12.639
14 FMA (default) 12.497 14 AK (AK2) 12.920
15 MAIVE (WAIVE) 12.849 15 MAIVE (WAIVE) 13.040
16 WILS (default) 13.312 16 MMPH (default) 13.503
17 AK (AK2) 13.778 17 WILS (default) 13.552
18 RMA (default) 14.423 18 RTMA (relaxed) 14.117
19 puniform (default) 14.731 19 RMA (default) 14.750
20 trimfill (default) 14.948 20 puniform (default) 14.809
21 pcurve (default) 15.648 21 trimfill (default) 15.090
22 RTMA (relaxed) 16.713 22 pcurve (default) 15.840
23 mean (default) 16.886 23 mean (default) 17.312
24 MAN (default) 18.025 24 MAN (default) 17.559

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Methods are compared using condition-wise ranks. Direct comparison using the average bias is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Mean Rank Rank Method Mean Rank
1 RMA (default) 3.827 1 RMA (default) 3.531
2 WLS (default) 4.821 2 WLS (default) 4.540
3 FMA (default) 4.827 3 FMA (default) 4.546
4 AK (AK1) 6.691 4 AK (AK1) 6.799
5 MMPH (default) 6.753 5 MMPH (default) 7.898
6 WAAPWLS (default) 8.466 6 WAAPWLS (default) 8.256
7 trimfill (default) 8.491 7 trimfill (default) 8.377
8 RoBMA (PSMA) 8.769 8 RoBMA (PSMA) 9.102
9 mean (default) 9.472 9 mean (default) 9.241
10 SM (3PSM) 11.676 10 SM (3PSM) 11.512
11 MAN (default) 12.062 11 PEESE (default) 12.988
12 PEESE (default) 13.154 12 pcurve (default) 13.093
13 pcurve (default) 13.262 13 WILS (default) 13.302
14 WILS (default) 13.512 14 MAN (default) 13.633
15 puniform (default) 14.272 15 puniform (star) 14.083
16 puniform (star) 14.293 16 puniform (default) 14.164
17 AK (AK2) 15.565 17 AK (AK2) 15.898
18 SM (4PSM) 16.136 18 SM (4PSM) 15.929
19 PETPEESE (default) 16.509 19 PETPEESE (default) 16.420
20 MAIVE (default) 17.852 20 RTMA (relaxed) 17.636
21 RTMA (relaxed) 18.015 21 MAIVE (default) 17.802
22 EK (default) 18.574 22 EK (default) 18.417
23 PET (default) 18.623 23 PET (default) 18.469
24 MAIVE (WAIVE) 21.892 24 MAIVE (WAIVE) 21.880

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Methods are compared using condition-wise ranks. Direct comparison using the average empirical standard error is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of empirical standard error values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 4.515 1 RoBMA (PSMA) 4.932
2 SM (3PSM) 6.519 2 SM (3PSM) 6.636
3 puniform (star) 7.377 3 puniform (star) 7.380
4 MMPH (default) 7.917 4 AK (AK1) 7.852
5 AK (AK1) 8.123 5 SM (4PSM) 8.410
6 SM (4PSM) 8.821 6 MMPH (default) 9.117
7 WAAPWLS (default) 11.414 7 WAAPWLS (default) 11.488
8 EK (default) 11.772 8 MAIVE (default) 11.840
9 MAIVE (default) 11.901 9 EK (default) 11.843
10 RMA (default) 12.133 10 RMA (default) 12.127
11 WLS (default) 12.213 11 WLS (default) 12.262
12 trimfill (default) 12.778 12 trimfill (default) 12.673
13 PETPEESE (default) 12.864 13 PETPEESE (default) 12.753
14 PET (default) 12.972 14 AK (AK2) 12.769
15 PEESE (default) 12.978 15 PEESE (default) 12.864
16 AK (AK2) 13.080 16 PET (default) 13.028
17 WILS (default) 14.346 17 WILS (default) 14.327
18 MAIVE (WAIVE) 14.389 18 MAIVE (WAIVE) 14.565
19 puniform (default) 14.657 19 puniform (default) 14.639
20 FMA (default) 14.870 20 RTMA (relaxed) 14.895
21 MAN (default) 15.997 21 FMA (default) 14.935
22 RTMA (relaxed) 16.262 22 MAN (default) 16.290
23 mean (default) 17.531 23 mean (default) 17.806
24 pcurve (default) 24.000 24 pcurve (default) 24.000

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average Interval Score is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of empirical standard error values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.902 1 RoBMA (PSMA) 0.896
2 SM (4PSM) 0.845 2 SM (4PSM) 0.833
3 AK (AK2) 0.801 3 MAIVE (default) 0.796
4 MAIVE (default) 0.796 4 AK (AK2) 0.783
5 RTMA (relaxed) 0.772 5 puniform (star) 0.765
6 puniform (star) 0.765 6 MAIVE (WAIVE) 0.737
7 SM (3PSM) 0.743 7 SM (3PSM) 0.736
8 MAIVE (WAIVE) 0.737 8 EK (default) 0.715
9 EK (default) 0.715 9 PET (default) 0.688
10 PET (default) 0.688 10 PETPEESE (default) 0.681
11 MMPH (default) 0.688 11 MMPH (default) 0.649
12 PETPEESE (default) 0.681 12 RTMA (relaxed) 0.634
13 AK (AK1) 0.608 13 AK (AK1) 0.607
14 puniform (default) 0.541 14 puniform (default) 0.542
15 PEESE (default) 0.524 15 PEESE (default) 0.524
16 WAAPWLS (default) 0.510 16 WAAPWLS (default) 0.510
17 trimfill (default) 0.497 17 trimfill (default) 0.497
18 WILS (default) 0.494 18 WILS (default) 0.494
19 RMA (default) 0.493 19 RMA (default) 0.493
20 WLS (default) 0.481 20 WLS (default) 0.481
21 FMA (default) 0.380 21 FMA (default) 0.380
22 mean (default) 0.366 22 mean (default) 0.366
23 MAN (default) 0.238 23 MAN (default) 0.304
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Mean Rank Rank Method Mean Rank
1 FMA (default) 2.438 1 FMA (default) 2.435
2 WILS (default) 3.315 2 WILS (default) 3.256
3 WLS (default) 4.123 3 WLS (default) 4.049
4 WAAPWLS (default) 6.349 4 WAAPWLS (default) 6.315
5 trimfill (default) 7.448 5 trimfill (default) 7.420
6 RMA (default) 7.509 6 RMA (default) 7.423
7 mean (default) 9.272 7 mean (default) 9.336
8 PEESE (default) 9.972 8 RoBMA (PSMA) 10.090
8 RoBMA (PSMA) 9.972 9 PEESE (default) 10.111
10 AK (AK1) 10.123 10 AK (AK1) 10.370
11 MMPH (default) 11.657 11 MMPH (default) 11.164
12 SM (3PSM) 11.969 12 SM (3PSM) 12.247
13 PETPEESE (default) 12.802 13 MAN (default) 12.775
14 puniform (default) 13.133 14 PETPEESE (default) 13.046
15 MAN (default) 13.438 15 puniform (default) 13.182
16 puniform (star) 14.052 16 puniform (star) 14.497
17 AK (AK2) 16.157 17 SM (4PSM) 16.198
18 SM (4PSM) 16.247 18 AK (AK2) 16.312
19 PET (default) 16.997 19 PET (default) 17.299
20 EK (default) 18.546 20 EK (default) 18.941
21 MAIVE (default) 18.701 21 MAIVE (default) 19.059
22 MAIVE (WAIVE) 20.429 22 RTMA (relaxed) 19.105
23 RTMA (relaxed) 20.778 23 MAIVE (WAIVE) 20.799
24 pcurve (default) 24.000 24 pcurve (default) 24.000

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average CI width is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of CI width values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RoBMA (PSMA) 5.355 1 RoBMA (PSMA) 4.807
2 AK (AK2) 2.939 2 AK (AK2) 2.293
3 SM (4PSM) 2.219 3 SM (4PSM) 2.238
4 puniform (star) 2.163 4 puniform (star) 2.163
5 MAIVE (WAIVE) 2.146 5 MAIVE (WAIVE) 2.146
6 MMPH (default) 2.136 6 SM (3PSM) 1.984
7 SM (3PSM) 1.998 7 MAIVE (default) 1.956
8 MAIVE (default) 1.956 8 EK (default) 1.840
9 EK (default) 1.840 9 PET (default) 1.838
10 PET (default) 1.838 10 PETPEESE (default) 1.834
11 PETPEESE (default) 1.834 11 RTMA (relaxed) 1.801
12 puniform (default) 1.728 12 MMPH (default) 1.745
13 RTMA (relaxed) 1.651 13 puniform (default) 1.667
14 AK (AK1) 1.493 14 AK (AK1) 1.445
15 WILS (default) 1.370 15 WILS (default) 1.370
16 RMA (default) 1.161 16 RMA (default) 1.161
17 PEESE (default) 1.109 17 PEESE (default) 1.109
18 WAAPWLS (default) 1.056 18 WAAPWLS (default) 1.056
19 WLS (default) 1.019 19 WLS (default) 1.019
20 trimfill (default) 0.953 20 trimfill (default) 0.953
21 MAN (default) 0.879 21 MAN (default) 0.939
22 mean (default) 0.849 22 mean (default) 0.849
23 FMA (default) 0.800 23 FMA (default) 0.800
24 pcurve (default) NaN 24 pcurve (default) NaN

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 SM (3PSM) -4.898 1 AK (AK2) -6.215
2 RoBMA (PSMA) -4.722 2 SM (3PSM) -4.947
3 PETPEESE (default) -4.721 3 SM (4PSM) -4.828
4 AK (AK2) -4.622 4 RoBMA (PSMA) -4.737
5 EK (default) -4.570 5 PETPEESE (default) -4.721
6 PET (default) -4.570 6 EK (default) -4.570
7 SM (4PSM) -4.438 7 PET (default) -4.570
8 MAIVE (default) -4.415 8 MAIVE (default) -4.415
9 puniform (star) -4.287 9 MMPH (default) -4.336
10 MMPH (default) -4.071 10 puniform (star) -4.287
11 PEESE (default) -3.817 11 PEESE (default) -3.817
12 WILS (default) -3.661 12 WILS (default) -3.661
13 puniform (default) -3.597 13 puniform (default) -3.603
14 trimfill (default) -3.502 14 RTMA (relaxed) -3.554
15 AK (AK1) -3.458 15 trimfill (default) -3.502
16 WLS (default) -3.393 16 AK (AK1) -3.461
17 RMA (default) -3.312 17 WLS (default) -3.393
18 FMA (default) -3.220 18 RMA (default) -3.312
19 WAAPWLS (default) -3.095 19 FMA (default) -3.220
20 mean (default) -2.700 20 WAAPWLS (default) -3.095
21 MAIVE (WAIVE) -2.422 21 mean (default) -2.700
22 RTMA (relaxed) -2.209 22 MAN (default) -2.536
23 MAN (default) -1.791 23 MAIVE (WAIVE) -2.422
24 pcurve (default) NaN 24 pcurve (default) NaN

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.022 1 RoBMA (PSMA) 0.031
2 AK (AK2) 0.082 2 MAIVE (WAIVE) 0.112
3 MAIVE (WAIVE) 0.112 3 SM (4PSM) 0.125
4 SM (4PSM) 0.112 4 AK (AK2) 0.156
5 PET (default) 0.237 5 PET (default) 0.237
6 EK (default) 0.237 6 EK (default) 0.237
7 MAIVE (default) 0.238 7 MAIVE (default) 0.238
8 puniform (star) 0.242 8 puniform (star) 0.242
9 RTMA (relaxed) 0.251 9 PETPEESE (default) 0.269
10 PETPEESE (default) 0.269 10 SM (3PSM) 0.288
11 SM (3PSM) 0.282 11 WILS (default) 0.373
12 MMPH (default) 0.349 12 RTMA (relaxed) 0.379
13 WILS (default) 0.373 13 MMPH (default) 0.416
14 PEESE (default) 0.541 14 PEESE (default) 0.541
15 puniform (default) 0.544 15 puniform (default) 0.542
16 AK (AK1) 0.556 16 AK (AK1) 0.557
17 MAN (default) 0.556 17 WAAPWLS (default) 0.573
18 WAAPWLS (default) 0.573 18 MAN (default) 0.592
19 RMA (default) 0.603 19 RMA (default) 0.603
20 WLS (default) 0.612 20 WLS (default) 0.612
21 trimfill (default) 0.615 21 trimfill (default) 0.615
22 mean (default) 0.688 22 mean (default) 0.688
23 FMA (default) 0.720 23 FMA (default) 0.720
24 pcurve (default) NaN 24 pcurve (default) NaN

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 FMA (default) 0.990 1 FMA (default) 0.990
2 WLS (default) 0.981 2 WLS (default) 0.981
3 RMA (default) 0.980 3 RMA (default) 0.980
4 trimfill (default) 0.979 4 trimfill (default) 0.979
5 AK (AK1) 0.970 5 AK (AK2) 0.977
6 mean (default) 0.965 6 AK (AK1) 0.970
7 WAAPWLS (default) 0.956 7 mean (default) 0.965
8 PEESE (default) 0.951 8 WAAPWLS (default) 0.956
9 AK (AK2) 0.949 9 RTMA (relaxed) 0.954
10 SM (3PSM) 0.936 10 PEESE (default) 0.951
11 MMPH (default) 0.925 11 SM (3PSM) 0.945
12 puniform (default) 0.911 12 MMPH (default) 0.927
13 WILS (default) 0.898 13 puniform (default) 0.913
14 PETPEESE (default) 0.884 14 MAN (default) 0.905
15 puniform (star) 0.876 15 SM (4PSM) 0.900
16 SM (4PSM) 0.869 16 WILS (default) 0.898
17 MAIVE (default) 0.866 17 PETPEESE (default) 0.884
18 EK (default) 0.852 18 puniform (star) 0.876
19 PET (default) 0.851 19 MAIVE (default) 0.866
20 MAN (default) 0.832 20 EK (default) 0.852
21 RoBMA (PSMA) 0.832 21 PET (default) 0.851
22 MAIVE (WAIVE) 0.659 22 RoBMA (PSMA) 0.834
23 RTMA (relaxed) 0.648 23 MAIVE (WAIVE) 0.659
24 pcurve (default) NaN 24 pcurve (default) NaN

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Conditional on Method Convergence)

The results below are conditional on method convergence. Note that the methods might differ in convergence rate and are therefore not compared on the same data sets.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average RMSE is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of RMSE values on the corresponding outcome scale.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Methods are compared using condition-wise ranks. Direct comparison using the average bias is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Methods are compared using condition-wise ranks. Direct comparison using the empirical standard error is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the interval score is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average 95% CI width is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of 95% CI width values on the corresponding outcome scale.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Replacement in Case of Non-Convergence)

The results below incorporate method replacement to handle non-convergence. If a method fails to converge, its results are replaced with the results from a simpler method (e.g., random-effects meta-analysis without publication bias adjustment). This emulates what a data analyst may do in practice in case a method does not converge. However, note that these results do not correspond to “pure” method performance as they might combine multiple different methods. See Method Replacement Strategy for details of the method replacement specification.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average RMSE is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of RMSE values on the corresponding outcome scale.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Methods are compared using condition-wise ranks. Direct comparison using the average bias is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Methods are compared using condition-wise ranks. Direct comparison using the empirical standard error is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the interval score is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average 95% CI width is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of 95% CI width values on the corresponding outcome scale.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

Subset: Standardized Mean Difference Effect Sizes

These results are based on Stanley (2017) data-generating mechanism with a total of 1 conditions.

Average Performance

Method performance measures are aggregated across all simulated conditions to provide an overall impression of method performance. However, keep in mind that a method with a high overall ranking is not necessarily the “best” method for a particular application. To select a suitable method for your application, consider also non-aggregated performance measures in conditions most relevant to your application.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.057 1 RoBMA (PSMA) 0.057
2 AK (AK2) 0.067 2 AK (AK2) 0.083
3 MMPH (default) 0.092 3 WILS (default) 0.095
4 WILS (default) 0.095 4 MMPH (default) 0.098
5 PEESE (default) 0.105 5 PEESE (default) 0.105
6 WAAPWLS (default) 0.108 6 WAAPWLS (default) 0.108
7 PETPEESE (default) 0.109 7 PETPEESE (default) 0.109
8 trimfill (default) 0.111 8 trimfill (default) 0.111
9 MAIVE (default) 0.112 9 MAIVE (default) 0.112
10 FMA (default) 0.112 10 FMA (default) 0.112
10 WLS (default) 0.112 10 WLS (default) 0.112
12 EK (default) 0.119 12 EK (default) 0.119
13 PET (default) 0.119 13 PET (default) 0.119
14 RMA (default) 0.123 14 RMA (default) 0.123
15 SM (3PSM) 0.130 15 SM (3PSM) 0.126
16 mean (default) 0.146 16 RTMA (relaxed) 0.130
17 AK (AK1) 0.163 17 mean (default) 0.146
18 MAIVE (WAIVE) 0.173 18 AK (AK1) 0.158
19 RTMA (relaxed) 0.190 19 MAIVE (WAIVE) 0.173
20 SM (4PSM) 0.202 20 MAN (default) 0.185
21 MAN (default) 0.207 21 SM (4PSM) 0.203
22 pcurve (default) 0.467 22 pcurve (default) 0.420
23 puniform (default) 0.648 23 puniform (default) 0.545
24 puniform (star) 90.749 24 puniform (star) 90.749

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.006 1 RoBMA (PSMA) 0.006
2 WILS (default) 0.007 2 WILS (default) 0.007
3 AK (AK2) 0.010 3 AK (AK2) 0.011
4 PET (default) 0.015 4 PET (default) 0.015
5 EK (default) 0.015 5 EK (default) 0.015
6 SM (4PSM) -0.023 6 SM (4PSM) -0.022
7 SM (3PSM) 0.023 7 SM (3PSM) 0.025
8 PETPEESE (default) 0.036 8 PETPEESE (default) 0.036
9 MAIVE (default) 0.040 9 MAIVE (default) 0.040
10 MAIVE (WAIVE) -0.042 10 MAIVE (WAIVE) -0.042
11 MMPH (default) 0.057 11 RTMA (relaxed) 0.051
12 PEESE (default) 0.057 12 PEESE (default) 0.057
13 AK (AK1) 0.060 13 AK (AK1) 0.061
14 trimfill (default) 0.069 14 MMPH (default) 0.062
15 WAAPWLS (default) 0.076 15 trimfill (default) 0.069
16 FMA (default) 0.085 16 WAAPWLS (default) 0.076
16 WLS (default) 0.085 17 FMA (default) 0.085
18 puniform (default) 0.088 17 WLS (default) 0.085
19 RTMA (relaxed) -0.090 19 MAN (default) -0.096
20 RMA (default) 0.103 20 RMA (default) 0.103
21 mean (default) 0.125 21 puniform (default) 0.103
22 MAN (default) -0.174 22 mean (default) 0.125
23 pcurve (default) 0.322 23 pcurve (default) 0.260
24 puniform (star) -9.762 24 puniform (star) -9.762

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 MAN (default) 0.034 1 RMA (default) 0.038
2 RMA (default) 0.038 2 mean (default) 0.042
3 mean (default) 0.042 3 FMA (default) 0.043
4 MMPH (default) 0.042 3 WLS (default) 0.043
5 FMA (default) 0.043 5 MMPH (default) 0.045
5 WLS (default) 0.043 6 RoBMA (PSMA) 0.045
7 RoBMA (PSMA) 0.045 7 trimfill (default) 0.046
8 trimfill (default) 0.046 8 WAAPWLS (default) 0.047
9 WAAPWLS (default) 0.047 9 PEESE (default) 0.058
10 PEESE (default) 0.058 10 MAN (default) 0.058
11 AK (AK2) 0.059 11 WILS (default) 0.063
12 WILS (default) 0.063 12 AK (AK2) 0.076
13 PETPEESE (default) 0.079 13 PETPEESE (default) 0.079
14 MAIVE (default) 0.081 14 MAIVE (default) 0.081
15 EK (default) 0.095 15 RTMA (relaxed) 0.090
16 PET (default) 0.095 16 EK (default) 0.095
17 RTMA (relaxed) 0.104 17 PET (default) 0.095
18 SM (3PSM) 0.111 18 SM (3PSM) 0.107
19 AK (AK1) 0.113 19 AK (AK1) 0.109
20 MAIVE (WAIVE) 0.137 20 MAIVE (WAIVE) 0.137
21 SM (4PSM) 0.200 21 SM (4PSM) 0.201
22 pcurve (default) 0.290 22 pcurve (default) 0.286
23 puniform (default) 0.550 23 puniform (default) 0.448
24 puniform (star) 89.867 24 puniform (star) 89.867

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.271 1 RoBMA (PSMA) 0.271
2 AK (AK2) 0.508 2 AK (AK2) 0.641
3 puniform (star) 0.914 3 SM (4PSM) 0.891
4 EK (default) 1.027 4 puniform (star) 0.914
5 SM (3PSM) 1.058 5 EK (default) 1.027
6 PET (default) 1.097 6 SM (3PSM) 1.042
7 SM (4PSM) 1.140 7 PET (default) 1.097
8 MAIVE (default) 1.358 8 MAIVE (default) 1.358
9 MAIVE (WAIVE) 1.399 9 MAIVE (WAIVE) 1.399
10 PETPEESE (default) 1.536 10 PETPEESE (default) 1.536
11 MMPH (default) 1.575 11 WILS (default) 1.658
12 WILS (default) 1.658 12 MMPH (default) 1.765
13 PEESE (default) 1.834 13 PEESE (default) 1.834
14 WAAPWLS (default) 2.206 14 RTMA (relaxed) 1.916
15 trimfill (default) 2.263 15 WAAPWLS (default) 2.206
16 WLS (default) 2.545 16 trimfill (default) 2.263
17 RTMA (relaxed) 2.722 17 WLS (default) 2.545
18 RMA (default) 2.805 18 RMA (default) 2.805
19 FMA (default) 3.110 19 FMA (default) 3.110
20 AK (AK1) 3.353 20 AK (AK1) 3.157
21 puniform (default) 4.061 21 puniform (default) 3.890
22 mean (default) 4.067 22 mean (default) 4.067
23 MAN (default) 5.443 23 MAN (default) 4.835
24 pcurve (default) NaN 24 pcurve (default) NaN

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.950 1 RoBMA (PSMA) 0.950
2 SM (4PSM) 0.924 2 SM (4PSM) 0.911
3 AK (AK2) 0.902 3 AK (AK2) 0.904
4 RTMA (relaxed) 0.839 4 puniform (star) 0.830
5 puniform (star) 0.830 5 MAIVE (default) 0.827
6 MAIVE (default) 0.827 6 SM (3PSM) 0.802
7 SM (3PSM) 0.807 7 EK (default) 0.744
8 EK (default) 0.744 8 MAIVE (WAIVE) 0.724
9 MMPH (default) 0.727 9 PETPEESE (default) 0.718
10 MAIVE (WAIVE) 0.724 10 PET (default) 0.716
11 PETPEESE (default) 0.718 11 MMPH (default) 0.698
12 PET (default) 0.716 12 RTMA (relaxed) 0.690
13 AK (AK1) 0.666 13 AK (AK1) 0.665
14 PEESE (default) 0.573 14 PEESE (default) 0.573
15 WAAPWLS (default) 0.571 15 WAAPWLS (default) 0.571
16 trimfill (default) 0.550 16 trimfill (default) 0.550
17 RMA (default) 0.546 17 RMA (default) 0.546
18 WLS (default) 0.537 18 WLS (default) 0.537
19 puniform (default) 0.531 19 puniform (default) 0.532
20 WILS (default) 0.524 20 WILS (default) 0.524
21 FMA (default) 0.414 21 FMA (default) 0.414
22 mean (default) 0.376 22 mean (default) 0.376
23 MAN (default) 0.203 23 MAN (default) 0.296
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 FMA (default) 0.079 1 FMA (default) 0.079
2 mean (default) 0.105 2 mean (default) 0.105
3 WILS (default) 0.130 3 WILS (default) 0.130
4 WLS (default) 0.134 4 WLS (default) 0.134
5 WAAPWLS (default) 0.149 5 MAN (default) 0.139
6 MAN (default) 0.159 6 WAAPWLS (default) 0.149
7 trimfill (default) 0.167 7 trimfill (default) 0.167
8 PEESE (default) 0.171 8 PEESE (default) 0.171
9 RMA (default) 0.173 9 RMA (default) 0.173
10 RoBMA (PSMA) 0.191 10 RoBMA (PSMA) 0.191
11 MMPH (default) 0.214 11 MMPH (default) 0.212
12 PETPEESE (default) 0.235 12 PETPEESE (default) 0.235
13 AK (AK2) 0.255 13 puniform (star) 0.278
14 puniform (star) 0.278 14 PET (default) 0.300
15 PET (default) 0.300 15 SM (3PSM) 0.321
16 MAIVE (default) 0.360 16 MAIVE (default) 0.360
17 SM (3PSM) 0.366 17 EK (default) 0.367
18 EK (default) 0.367 18 MAIVE (WAIVE) 0.451
19 MAIVE (WAIVE) 0.451 19 puniform (default) 0.453
20 puniform (default) 0.588 20 RTMA (relaxed) 0.461
21 SM (4PSM) 0.992 21 AK (AK2) 0.464
22 AK (AK1) 1.982 22 SM (4PSM) 0.655
23 RTMA (relaxed) 2.048 23 AK (AK1) 1.781
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RoBMA (PSMA) 5.324 1 RoBMA (PSMA) 5.324
2 AK (AK2) 3.027 2 AK (AK2) 2.486
3 SM (4PSM) 2.431 3 SM (4PSM) 2.441
4 MAIVE (WAIVE) 2.388 4 MAIVE (WAIVE) 2.388
5 puniform (star) 2.007 5 puniform (star) 2.007
6 MAIVE (default) 1.911 6 MAIVE (default) 1.911
7 SM (3PSM) 1.871 7 SM (3PSM) 1.852
8 MMPH (default) 1.863 8 PET (default) 1.818
9 PET (default) 1.818 9 EK (default) 1.818
10 EK (default) 1.818 10 PETPEESE (default) 1.677
11 PETPEESE (default) 1.677 11 MMPH (default) 1.632
12 AK (AK1) 1.339 12 RTMA (relaxed) 1.589
13 RTMA (relaxed) 1.295 13 AK (AK1) 1.313
14 WILS (default) 1.130 14 WILS (default) 1.130
15 RMA (default) 1.060 15 RMA (default) 1.060
16 PEESE (default) 0.977 16 PEESE (default) 0.977
17 puniform (default) 0.962 17 puniform (default) 0.966
18 WAAPWLS (default) 0.946 18 WAAPWLS (default) 0.946
19 WLS (default) 0.882 19 WLS (default) 0.882
20 trimfill (default) 0.825 20 trimfill (default) 0.825
21 mean (default) 0.651 21 mean (default) 0.651
22 FMA (default) 0.566 22 FMA (default) 0.566
23 MAN (default) 0.391 23 MAN (default) 0.523
24 pcurve (default) NaN 24 pcurve (default) NaN

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RoBMA (PSMA) -5.344 1 AK (AK2) -6.371
2 MAIVE (default) -5.046 2 RoBMA (PSMA) -5.344
3 PETPEESE (default) -5.031 3 SM (4PSM) -5.289
4 AK (AK2) -4.939 4 MAIVE (default) -5.046
5 PET (default) -4.925 5 PETPEESE (default) -5.031
6 EK (default) -4.925 6 SM (3PSM) -4.949
7 SM (3PSM) -4.896 7 PET (default) -4.925
8 SM (4PSM) -4.804 8 EK (default) -4.925
9 MMPH (default) -4.241 9 MMPH (default) -4.364
10 puniform (star) -4.121 10 puniform (star) -4.121
11 puniform (default) -3.885 11 puniform (default) -3.892
12 PEESE (default) -3.691 12 PEESE (default) -3.691
13 WAAPWLS (default) -3.494 13 WAAPWLS (default) -3.494
14 trimfill (default) -3.454 14 trimfill (default) -3.454
15 AK (AK1) -3.453 15 AK (AK1) -3.453
16 WLS (default) -3.274 16 WLS (default) -3.274
17 WILS (default) -3.233 17 WILS (default) -3.233
18 RMA (default) -3.200 18 RMA (default) -3.200
19 FMA (default) -3.049 19 RTMA (relaxed) -3.129
20 mean (default) -3.038 20 FMA (default) -3.049
21 MAIVE (WAIVE) -2.985 21 mean (default) -3.038
22 RTMA (relaxed) -1.509 22 MAIVE (WAIVE) -2.985
23 MAN (default) -0.712 23 MAN (default) -1.608
24 pcurve (default) NaN 24 pcurve (default) NaN

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.025 1 RoBMA (PSMA) 0.025
2 AK (AK2) 0.073 2 SM (4PSM) 0.109
3 SM (4PSM) 0.095 3 MAIVE (WAIVE) 0.115
4 MAIVE (WAIVE) 0.115 4 AK (AK2) 0.136
5 PET (default) 0.259 5 PET (default) 0.259
6 EK (default) 0.259 6 EK (default) 0.259
7 MAIVE (default) 0.259 7 MAIVE (default) 0.259
8 puniform (star) 0.268 8 puniform (star) 0.268
9 RTMA (relaxed) 0.279 9 PETPEESE (default) 0.296
10 PETPEESE (default) 0.296 10 SM (3PSM) 0.314
11 SM (3PSM) 0.308 11 WILS (default) 0.410
12 MMPH (default) 0.391 12 RTMA (relaxed) 0.416
13 WILS (default) 0.410 13 MMPH (default) 0.436
14 PEESE (default) 0.568 14 PEESE (default) 0.568
15 AK (AK1) 0.579 15 AK (AK1) 0.579
16 WAAPWLS (default) 0.590 16 WAAPWLS (default) 0.590
17 puniform (default) 0.615 17 puniform (default) 0.612
18 MAN (default) 0.621 18 RMA (default) 0.623
19 RMA (default) 0.623 19 WLS (default) 0.634
20 WLS (default) 0.634 20 trimfill (default) 0.638
21 trimfill (default) 0.638 21 MAN (default) 0.657
22 mean (default) 0.723 22 mean (default) 0.723
23 FMA (default) 0.753 23 FMA (default) 0.753
24 pcurve (default) NaN 24 pcurve (default) NaN

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 mean (default) 0.999 1 mean (default) 0.999
2 FMA (default) 0.999 2 FMA (default) 0.999
3 puniform (default) 0.997 3 puniform (default) 0.997
4 AK (AK1) 0.990 4 AK (AK1) 0.990
5 WLS (default) 0.989 5 WLS (default) 0.989
6 RMA (default) 0.989 6 RMA (default) 0.989
7 trimfill (default) 0.988 7 trimfill (default) 0.988
8 WAAPWLS (default) 0.972 8 AK (AK2) 0.983
9 PEESE (default) 0.969 9 WAAPWLS (default) 0.972
10 AK (AK2) 0.968 10 PEESE (default) 0.969
11 MMPH (default) 0.963 11 SM (3PSM) 0.965
12 SM (3PSM) 0.960 12 MMPH (default) 0.963
13 MAIVE (default) 0.945 13 RTMA (relaxed) 0.962
14 PETPEESE (default) 0.931 14 MAIVE (default) 0.945
15 EK (default) 0.905 15 SM (4PSM) 0.938
15 PET (default) 0.905 16 PETPEESE (default) 0.931
17 SM (4PSM) 0.903 17 EK (default) 0.905
18 WILS (default) 0.901 17 PET (default) 0.905
19 RoBMA (PSMA) 0.898 19 WILS (default) 0.901
20 puniform (star) 0.886 20 RoBMA (PSMA) 0.898
21 MAN (default) 0.814 21 MAN (default) 0.892
22 MAIVE (WAIVE) 0.767 22 puniform (star) 0.886
23 RTMA (relaxed) 0.596 23 MAIVE (WAIVE) 0.767
24 pcurve (default) NaN 24 pcurve (default) NaN

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Conditional on Method Convergence)

The results below are conditional on method convergence. Note that the methods might differ in convergence rate and are therefore not compared on the same data sets.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Replacement in Case of Non-Convergence)

The results below incorporate method replacement to handle non-convergence. If a method fails to converge, its results are replaced with the results from a simpler method (e.g., random-effects meta-analysis without publication bias adjustment). This emulates what a data analyst may do in practice in case a method does not converge. However, note that these results do not correspond to “pure” method performance as they might combine multiple different methods. See Method Replacement Strategy for details of the method replacement specification.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

Subset: Log Odd Ratio Effect Sizes

These results are based on Stanley (2017) data-generating mechanism with a total of 1 conditions.

Average Performance

Method performance measures are aggregated across all simulated conditions to provide an overall impression of method performance. However, keep in mind that a method with a high overall ranking is not necessarily the “best” method for a particular application. To select a suitable method for your application, consider also non-aggregated performance measures in conditions most relevant to your application.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 MAN (default) 0.172 1 MAN (default) 0.210
2 RoBMA (PSMA) 0.206 2 RoBMA (PSMA) 0.238
3 MMPH (default) 0.243 3 EK (default) 0.263
4 EK (default) 0.263 4 PET (default) 0.264
5 PET (default) 0.264 5 SM (3PSM) 0.273
6 AK (AK2) 0.265 6 WILS (default) 0.274
7 WILS (default) 0.274 7 PETPEESE (default) 0.297
8 SM (3PSM) 0.274 8 SM (4PSM) 0.308
9 PETPEESE (default) 0.297 9 PEESE (default) 0.313
10 PEESE (default) 0.313 10 MMPH (default) 0.320
11 SM (4PSM) 0.316 11 MAIVE (default) 0.337
12 MAIVE (default) 0.337 12 trimfill (default) 0.348
13 RTMA (relaxed) 0.338 13 RTMA (relaxed) 0.348
14 trimfill (default) 0.348 14 WAAPWLS (default) 0.352
15 WAAPWLS (default) 0.352 15 FMA (default) 0.372
16 FMA (default) 0.372 16 WLS (default) 0.372
17 WLS (default) 0.372 17 RMA (default) 0.391
18 RMA (default) 0.391 18 MAIVE (WAIVE) 0.420
19 MAIVE (WAIVE) 0.420 19 mean (default) 0.501
20 mean (default) 0.501 20 AK (AK2) 1.061
21 pcurve (default) 1.293 21 pcurve (default) 1.127
22 AK (AK1) 1.581 22 puniform (default) 1.339
23 puniform (default) 1.713 23 AK (AK1) 1.429
24 puniform (star) 157.696 24 puniform (star) 157.696

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 MAIVE (WAIVE) 0.093 1 puniform (default) -0.069
2 MAN (default) 0.144 2 MAIVE (WAIVE) 0.093
3 SM (4PSM) 0.160 3 SM (4PSM) 0.166
4 EK (default) 0.171 4 EK (default) 0.171
5 PET (default) 0.171 5 PET (default) 0.171
6 RoBMA (PSMA) 0.184 6 MAN (default) 0.181
7 MAIVE (default) 0.191 7 MAIVE (default) 0.191
8 MMPH (default) 0.201 8 RoBMA (PSMA) 0.204
9 SM (3PSM) 0.214 9 PETPEESE (default) 0.220
10 PETPEESE (default) 0.220 10 SM (3PSM) 0.222
11 AK (AK2) 0.231 11 WILS (default) 0.240
12 puniform (default) -0.240 12 AK (AK1) 0.253
13 WILS (default) 0.240 13 AK (AK2) 0.259
14 AK (AK1) 0.248 14 MMPH (default) 0.280
15 RTMA (relaxed) 0.277 15 PEESE (default) 0.284
16 PEESE (default) 0.284 16 RTMA (relaxed) 0.315
17 trimfill (default) 0.329 17 trimfill (default) 0.329
18 WAAPWLS (default) 0.335 18 WAAPWLS (default) 0.335
19 FMA (default) 0.356 19 FMA (default) 0.356
20 WLS (default) 0.356 20 WLS (default) 0.356
21 RMA (default) 0.375 21 RMA (default) 0.375
22 mean (default) 0.464 22 mean (default) 0.464
23 pcurve (default) 0.933 23 pcurve (default) 0.698
24 puniform (star) -3.661 24 puniform (star) -3.661

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 MAN (default) 0.048 1 FMA (default) 0.058
2 FMA (default) 0.058 2 WLS (default) 0.058
3 WLS (default) 0.058 3 MAN (default) 0.059
4 trimfill (default) 0.060 4 trimfill (default) 0.060
5 RMA (default) 0.061 5 RMA (default) 0.061
6 RoBMA (PSMA) 0.064 6 WAAPWLS (default) 0.067
7 WAAPWLS (default) 0.067 7 WILS (default) 0.075
8 MMPH (default) 0.070 8 MMPH (default) 0.088
9 WILS (default) 0.075 9 PEESE (default) 0.088
10 AK (AK2) 0.075 10 RoBMA (PSMA) 0.091
11 PEESE (default) 0.088 11 SM (3PSM) 0.095
12 SM (3PSM) 0.100 12 RTMA (relaxed) 0.108
13 mean (default) 0.113 13 mean (default) 0.113
14 RTMA (relaxed) 0.133 14 PETPEESE (default) 0.134
15 PETPEESE (default) 0.134 15 EK (default) 0.141
16 EK (default) 0.141 16 PET (default) 0.142
17 PET (default) 0.142 17 SM (4PSM) 0.170
18 SM (4PSM) 0.179 18 MAIVE (default) 0.209
19 MAIVE (default) 0.209 19 MAIVE (WAIVE) 0.375
20 MAIVE (WAIVE) 0.375 20 pcurve (default) 0.843
21 pcurve (default) 0.812 21 AK (AK2) 0.861
22 AK (AK1) 1.373 22 puniform (default) 1.175
23 puniform (default) 1.516 23 AK (AK1) 1.222
24 puniform (star) 156.858 24 puniform (star) 156.858

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 MAIVE (WAIVE) 3.195 1 MAIVE (WAIVE) 3.195
2 EK (default) 3.581 2 EK (default) 3.581
3 PET (default) 3.689 3 PET (default) 3.689
4 MAN (default) 3.772 4 RoBMA (PSMA) 4.717
5 RoBMA (PSMA) 3.887 5 SM (4PSM) 4.768
6 SM (4PSM) 4.723 6 MAN (default) 4.784
7 MAIVE (default) 4.790 7 MAIVE (default) 4.790
8 MMPH (default) 5.314 8 SM (3PSM) 5.735
9 SM (3PSM) 5.619 9 puniform (star) 5.831
10 puniform (star) 5.831 10 puniform (default) 6.525
11 AK (AK2) 5.959 11 PETPEESE (default) 6.582
12 PETPEESE (default) 6.582 12 WILS (default) 7.531
13 WILS (default) 7.531 13 PEESE (default) 7.667
14 PEESE (default) 7.667 14 RTMA (relaxed) 8.175
15 RTMA (relaxed) 7.863 15 MMPH (default) 8.278
16 puniform (default) 7.877 16 trimfill (default) 9.544
17 trimfill (default) 9.543 17 WAAPWLS (default) 9.694
18 WAAPWLS (default) 9.694 18 FMA (default) 10.660
19 FMA (default) 10.660 19 WLS (default) 10.776
20 WLS (default) 10.776 20 RMA (default) 10.804
21 RMA (default) 10.804 21 AK (AK2) 12.194
22 mean (default) 13.014 22 mean (default) 13.014
23 AK (AK1) 44.040 23 AK (AK1) 32.961
24 pcurve (default) NaN 24 pcurve (default) NaN

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 MAIVE (WAIVE) 0.804 1 MAIVE (WAIVE) 0.804
2 RoBMA (PSMA) 0.661 2 MAIVE (default) 0.643
3 MAIVE (default) 0.643 3 RoBMA (PSMA) 0.625
4 puniform (default) 0.592 4 puniform (default) 0.588
5 EK (default) 0.574 5 EK (default) 0.574
6 PET (default) 0.548 6 PET (default) 0.548
7 PETPEESE (default) 0.496 7 PETPEESE (default) 0.496
8 MMPH (default) 0.486 8 SM (4PSM) 0.443
9 SM (4PSM) 0.454 9 puniform (star) 0.440
10 RTMA (relaxed) 0.440 10 SM (3PSM) 0.407
11 puniform (star) 0.440 11 MMPH (default) 0.400
12 SM (3PSM) 0.423 12 RTMA (relaxed) 0.359
13 MAN (default) 0.399 13 WILS (default) 0.343
14 AK (AK2) 0.394 14 MAN (default) 0.340
15 WILS (default) 0.343 15 AK (AK1) 0.319
16 AK (AK1) 0.321 16 mean (default) 0.314
17 mean (default) 0.314 17 AK (AK2) 0.297
18 PEESE (default) 0.277 18 PEESE (default) 0.277
19 trimfill (default) 0.231 19 trimfill (default) 0.231
20 RMA (default) 0.226 20 RMA (default) 0.226
21 FMA (default) 0.213 21 FMA (default) 0.213
22 WAAPWLS (default) 0.205 22 WAAPWLS (default) 0.205
23 WLS (default) 0.200 23 WLS (default) 0.200
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 WILS (default) 0.221 1 MAN (default) 0.219
2 MAN (default) 0.228 2 WILS (default) 0.221
3 WLS (default) 0.235 3 WLS (default) 0.235
4 WAAPWLS (default) 0.249 4 WAAPWLS (default) 0.249
5 FMA (default) 0.250 5 FMA (default) 0.250
6 RoBMA (PSMA) 0.271 6 RoBMA (PSMA) 0.271
7 trimfill (default) 0.274 7 trimfill (default) 0.274
8 RMA (default) 0.298 8 RMA (default) 0.298
9 PEESE (default) 0.299 9 PEESE (default) 0.299
10 AK (AK2) 0.322 10 SM (3PSM) 0.332
11 SM (3PSM) 0.344 11 MMPH (default) 0.373
12 MMPH (default) 0.390 12 puniform (star) 0.407
13 puniform (star) 0.407 13 PETPEESE (default) 0.457
14 PETPEESE (default) 0.457 14 SM (4PSM) 0.503
15 PET (default) 0.536 15 PET (default) 0.536
16 SM (4PSM) 0.569 16 EK (default) 0.658
17 EK (default) 0.658 17 RTMA (relaxed) 0.717
18 MAIVE (default) 1.017 18 MAIVE (default) 1.017
19 MAIVE (WAIVE) 1.251 19 MAIVE (WAIVE) 1.251
20 mean (default) 1.505 20 mean (default) 1.505
21 RTMA (relaxed) 1.658 21 puniform (default) 2.751
22 puniform (default) 4.072 22 AK (AK2) 5.859
23 AK (AK1) 38.095 23 AK (AK1) 26.987
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RoBMA (PSMA) 5.471 1 puniform (default) 4.298
2 puniform (default) 4.600 2 RoBMA (PSMA) 2.869
3 MMPH (default) 3.192 3 puniform (star) 2.747
4 RTMA (relaxed) 2.987 4 RTMA (relaxed) 2.599
5 puniform (star) 2.747 5 SM (3PSM) 2.476
6 AK (AK2) 2.708 6 PETPEESE (default) 2.425
7 SM (3PSM) 2.473 7 WILS (default) 2.268
8 PETPEESE (default) 2.425 8 MAN (default) 2.253
9 MAN (default) 2.422 9 MMPH (default) 2.174
10 WILS (default) 2.268 10 MAIVE (default) 2.125
11 MAIVE (default) 2.125 11 AK (AK1) 1.940
12 AK (AK1) 2.070 12 EK (default) 1.922
13 EK (default) 1.922 13 PET (default) 1.913
14 PET (default) 1.913 14 AK (AK2) 1.782
15 FMA (default) 1.678 15 FMA (default) 1.678
16 PEESE (default) 1.605 16 PEESE (default) 1.605
17 mean (default) 1.593 17 mean (default) 1.593
18 RMA (default) 1.539 18 RMA (default) 1.539
19 WLS (default) 1.530 19 WLS (default) 1.530
20 WAAPWLS (default) 1.469 20 SM (4PSM) 1.478
21 trimfill (default) 1.432 21 WAAPWLS (default) 1.469
22 SM (4PSM) 1.425 22 trimfill (default) 1.432
23 MAIVE (WAIVE) 1.239 23 MAIVE (WAIVE) 1.239
24 pcurve (default) NaN 24 pcurve (default) NaN

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 WILS (default) -5.267 1 AK (AK2) -5.801
2 MAN (default) -5.195 2 MAN (default) -5.468
3 puniform (star) -4.910 3 WILS (default) -5.267
4 SM (3PSM) -4.904 4 RTMA (relaxed) -5.147
5 RTMA (relaxed) -4.832 5 SM (3PSM) -4.940
6 PEESE (default) -4.288 6 puniform (star) -4.910
7 FMA (default) -3.859 7 PEESE (default) -4.288
8 WLS (default) -3.837 8 MMPH (default) -4.228
9 AK (AK2) -3.785 9 FMA (default) -3.859
10 RMA (default) -3.730 10 WLS (default) -3.837
11 trimfill (default) -3.682 11 RMA (default) -3.730
12 PETPEESE (default) -3.556 12 trimfill (default) -3.682
13 AK (AK1) -3.477 13 PETPEESE (default) -3.556
14 MMPH (default) -3.414 14 AK (AK1) -3.488
15 EK (default) -3.238 15 EK (default) -3.238
16 PET (default) -3.236 16 PET (default) -3.236
17 SM (4PSM) -3.063 17 SM (4PSM) -3.100
18 puniform (default) -2.518 18 puniform (default) -2.517
19 RoBMA (PSMA) -2.388 19 RoBMA (PSMA) -2.460
20 MAIVE (default) -2.051 20 MAIVE (default) -2.051
21 WAAPWLS (default) -1.596 21 WAAPWLS (default) -1.596
22 mean (default) -1.432 22 mean (default) -1.432
23 MAIVE (WAIVE) -0.308 23 MAIVE (WAIVE) -0.308
24 pcurve (default) NaN 24 pcurve (default) NaN

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.003 1 puniform (default) 0.017
2 puniform (default) 0.009 2 puniform (star) 0.053
3 MMPH (default) 0.031 3 PETPEESE (default) 0.072
4 RTMA (relaxed) 0.036 4 PET (default) 0.075
5 puniform (star) 0.053 5 EK (default) 0.076
6 PETPEESE (default) 0.072 6 RoBMA (PSMA) 0.080
7 PET (default) 0.075 7 MAIVE (default) 0.084
8 EK (default) 0.076 8 MAIVE (WAIVE) 0.088
9 MAIVE (default) 0.084 9 SM (3PSM) 0.096
10 MAIVE (WAIVE) 0.088 10 WILS (default) 0.099
11 SM (3PSM) 0.089 11 RTMA (relaxed) 0.104
12 MAN (default) 0.091 12 MAN (default) 0.133
13 WILS (default) 0.099 13 SM (4PSM) 0.238
14 AK (AK2) 0.130 14 AK (AK2) 0.257
15 SM (4PSM) 0.237 15 MMPH (default) 0.259
16 PEESE (default) 0.334 16 PEESE (default) 0.334
17 AK (AK1) 0.385 17 AK (AK1) 0.387
18 mean (default) 0.420 18 mean (default) 0.420
19 trimfill (default) 0.444 19 trimfill (default) 0.444
20 WAAPWLS (default) 0.445 20 WAAPWLS (default) 0.445
21 WLS (default) 0.446 21 WLS (default) 0.446
22 RMA (default) 0.450 22 RMA (default) 0.450
23 FMA (default) 0.468 23 FMA (default) 0.468
24 pcurve (default) NaN 24 pcurve (default) NaN

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 FMA (default) 0.958 1 FMA (default) 0.958
2 WLS (default) 0.950 2 AK (AK2) 0.958
3 RMA (default) 0.948 3 WLS (default) 0.950
4 trimfill (default) 0.943 4 RMA (default) 0.948
5 WAAPWLS (default) 0.896 5 MAN (default) 0.947
6 AK (AK1) 0.893 6 trimfill (default) 0.943
7 MAN (default) 0.888 7 RTMA (relaxed) 0.925
8 WILS (default) 0.888 8 WAAPWLS (default) 0.896
9 AK (AK2) 0.886 9 AK (AK1) 0.894
10 PEESE (default) 0.882 10 WILS (default) 0.888
11 SM (3PSM) 0.850 11 PEESE (default) 0.882
12 puniform (star) 0.842 12 SM (3PSM) 0.871
13 RTMA (relaxed) 0.840 13 puniform (star) 0.842
14 mean (default) 0.835 14 mean (default) 0.835
15 MMPH (default) 0.777 15 MMPH (default) 0.789
16 SM (4PSM) 0.742 16 SM (4PSM) 0.759
17 PETPEESE (default) 0.708 17 PETPEESE (default) 0.708
18 EK (default) 0.654 18 EK (default) 0.654
19 PET (default) 0.652 19 PET (default) 0.652
20 puniform (default) 0.591 20 puniform (default) 0.597
21 RoBMA (PSMA) 0.584 21 RoBMA (PSMA) 0.596
22 MAIVE (default) 0.572 22 MAIVE (default) 0.572
23 MAIVE (WAIVE) 0.256 23 MAIVE (WAIVE) 0.256
24 pcurve (default) NaN 24 pcurve (default) NaN

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Conditional on Method Convergence)

The results below are conditional on method convergence. Note that the methods might differ in convergence rate and are therefore not compared on the same data sets.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Replacement in Case of Non-Convergence)

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

Session Info

This report was compiled on Wed Sep 30 15:50:51 2026 (UTC) using the following computational environment

## R version 4.6.1 (2026-06-24)
## Platform: x86_64-pc-linux-gnu
## Running under: Ubuntu 24.04.5 LTS
## 
## Matrix products: default
## BLAS:   /usr/lib/x86_64-linux-gnu/openblas-pthread/libblas.so.3 
## LAPACK: /usr/lib/x86_64-linux-gnu/openblas-pthread/libopenblasp-r0.3.26.so;  LAPACK version 3.12.0
## 
## locale:
##  [1] LC_CTYPE=C.UTF-8       LC_NUMERIC=C           LC_TIME=C.UTF-8       
##  [4] LC_COLLATE=C.UTF-8     LC_MONETARY=C.UTF-8    LC_MESSAGES=C.UTF-8   
##  [7] LC_PAPER=C.UTF-8       LC_NAME=C              LC_ADDRESS=C          
## [10] LC_TELEPHONE=C         LC_MEASUREMENT=C.UTF-8 LC_IDENTIFICATION=C   
## 
## time zone: UTC
## tzcode source: system (glibc)
## 
## attached base packages:
## [1] stats     graphics  grDevices utils     datasets  methods   base     
## 
## other attached packages:
## [1] scales_1.4.0                   ggdist_3.3.3                  
## [3] ggplot2_4.0.3                  PublicationBiasBenchmark_0.3.0
## 
## loaded via a namespace (and not attached):
##  [1] gtable_0.3.6         xfun_0.61            bslib_0.12.0        
##  [4] htmlwidgets_1.6.4    lattice_0.22-9       vctrs_0.7.3         
##  [7] tools_4.6.1          Rdpack_2.6.6         generics_0.1.4      
## [10] curl_8.0.0           sandwich_3.1-3       tibble_3.3.1        
## [13] pkgconfig_2.0.3      RColorBrewer_1.1-3   S7_0.2.2            
## [16] desc_1.4.3           distributional_0.9.0 lifecycle_1.0.5     
## [19] compiler_4.6.1       farver_2.1.2         stringr_1.6.0       
## [22] textshaping_1.0.5    htmltools_0.5.9      sass_0.4.10         
## [25] clubSandwich_0.7.0   yaml_2.3.12          pillar_1.11.1       
## [28] pkgdown_2.2.1        jquerylib_0.1.4      cachem_1.1.0        
## [31] tidyselect_1.2.1     digest_0.6.39        stringi_1.8.9       
## [34] dplyr_1.2.1          purrr_1.2.2          labeling_0.4.3      
## [37] fastmap_1.2.0        grid_4.6.1           cli_3.6.6           
## [40] magrittr_2.0.5       triebeard_0.4.1      crul_1.6.0          
## [43] osfr_0.2.9           withr_3.0.3          rmarkdown_2.32      
## [46] httr_1.4.9           otel_0.2.0           ragg_1.5.2          
## [49] zoo_1.9-1            kableExtra_1.4.1     memoise_2.0.1       
## [52] evaluate_1.0.5       knitr_1.52           rbibutils_2.4.1     
## [55] viridisLite_0.4.3    rlang_1.3.0          urltools_1.7.3.1    
## [58] Rcpp_1.1.2           glue_1.8.1           httpcode_0.3.0      
## [61] xml2_1.6.0           svglite_2.2.2        rstudioapi_0.19.0   
## [64] jsonlite_2.0.0       R6_2.6.1             systemfonts_1.3.2   
## [67] fs_2.1.0