Skip to contents

Complete Results

These results are based on Stanley (2017), Alinaghi (2018), Bom (2019), and Carter (2019) data-generating mechanisms with a total of 1665 conditions.

Average Performance

Method performance measures are aggregated across all simulated conditions to provide an overall impression of method performance. However, keep in mind that a method with a high overall ranking is not necessarily the “best” method for a particular application. To select a suitable method for your application, consider also non-aggregated performance measures in conditions most relevant to your application.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Mean Rank Rank Method Mean Rank
1 AK (AK1) 7.428 1 AK (AK1) 7.354
2 RoBMA (PSMA) 7.895 2 RoBMA (PSMA) 8.124
3 MMPH (default) 8.014 3 WAAPWLS (default) 8.741
4 WAAPWLS (default) 8.856 4 MMPH (default) 9.138
5 FMA (default) 9.500 5 FMA (default) 9.331
6 WLS (default) 9.508 6 WLS (default) 9.338
7 trimfill (default) 10.020 7 trimfill (default) 9.882
8 SM (3PSM) 10.321 8 PEESE (default) 10.348
9 PEESE (default) 10.402 9 SM (3PSM) 10.401
10 PETPEESE (default) 10.993 10 PETPEESE (default) 10.974
11 WILS (default) 11.020 11 WILS (default) 11.055
12 puniform (star) 11.350 12 puniform (star) 11.350
13 RMA (default) 12.280 13 RMA (default) 12.289
14 EK (default) 13.177 14 AK (AK2) 12.706
15 PET (default) 13.313 15 EK (default) 13.277
16 AK (AK2) 13.354 16 PET (default) 13.417
17 SM (4PSM) 13.489 17 SM (4PSM) 13.535
18 pcurve (default) 13.906 18 pcurve (default) 13.953
19 MAIVE (default) 15.156 19 MAIVE (default) 15.225
20 puniform (default) 15.682 20 puniform (default) 15.669
21 mean (default) 16.743 21 mean (default) 17.020
22 RTMA (relaxed) 17.996 22 RTMA (relaxed) 17.083
23 MAN (default) 18.072 23 MAN (default) 18.156
24 MAIVE (WAIVE) 19.172 24 MAIVE (WAIVE) 19.280

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average RMSE is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of RMSE values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Mean Rank Rank Method Mean Rank
1 WAAPWLS (default) 8.816 1 WAAPWLS (default) 8.881
2 PETPEESE (default) 9.034 2 PETPEESE (default) 9.216
3 AK (AK1) 9.250 3 AK (AK1) 9.356
4 PEESE (default) 9.570 4 PEESE (default) 9.686
5 SM (3PSM) 9.687 5 SM (3PSM) 9.802
6 RoBMA (PSMA) 9.905 6 RoBMA (PSMA) 10.064
7 EK (default) 10.433 7 EK (default) 10.695
8 PET (default) 10.517 8 PET (default) 10.780
9 puniform (star) 10.650 9 puniform (star) 10.799
10 WLS (default) 10.978 10 WLS (default) 11.080
11 FMA (default) 10.982 11 FMA (default) 11.084
12 SM (4PSM) 11.252 12 SM (4PSM) 11.435
13 MMPH (default) 11.729 13 AK (AK2) 11.959
14 WILS (default) 11.862 14 WILS (default) 12.162
15 trimfill (default) 12.370 15 trimfill (default) 12.506
16 MAIVE (default) 12.636 16 MMPH (default) 12.578
17 AK (AK2) 12.969 17 MAIVE (default) 12.873
18 RMA (default) 14.439 18 RTMA (relaxed) 14.564
19 puniform (default) 15.265 19 RMA (default) 14.877
20 pcurve (default) 15.362 20 puniform (default) 15.422
21 MAIVE (WAIVE) 15.634 21 pcurve (default) 15.557
22 RTMA (relaxed) 17.377 22 MAIVE (WAIVE) 15.820
23 mean (default) 17.782 23 mean (default) 18.242
24 MAN (default) 19.333 24 MAN (default) 18.393

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Methods are compared using condition-wise ranks. Direct comparison using the average bias is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Mean Rank Rank Method Mean Rank
1 RMA (default) 4.332 1 RMA (default) 4.047
2 AK (AK1) 5.870 2 AK (AK1) 6.001
3 WLS (default) 6.395 3 WLS (default) 6.151
4 FMA (default) 6.399 4 FMA (default) 6.156
5 MMPH (default) 7.181 5 trimfill (default) 7.695
6 trimfill (default) 7.842 6 MMPH (default) 8.443
7 mean (default) 8.887 7 mean (default) 8.727
8 pcurve (default) 9.139 8 pcurve (default) 9.055
9 RoBMA (PSMA) 9.753 9 WAAPWLS (default) 9.761
10 WAAPWLS (default) 9.980 10 RoBMA (PSMA) 10.438
11 PEESE (default) 12.220 11 PEESE (default) 12.094
12 SM (3PSM) 12.908 12 SM (3PSM) 12.689
13 puniform (default) 12.937 13 puniform (default) 12.841
14 MAN (default) 13.605 14 WILS (default) 13.959
15 WILS (default) 14.178 15 puniform (star) 14.265
16 puniform (star) 14.526 16 MAN (default) 14.887
17 PETPEESE (default) 15.353 17 PETPEESE (default) 15.226
18 AK (AK2) 15.818 18 AK (AK2) 15.801
19 SM (4PSM) 16.634 19 SM (4PSM) 16.415
20 EK (default) 17.553 20 EK (default) 17.448
21 PET (default) 17.632 21 PET (default) 17.532
22 MAIVE (default) 17.773 22 MAIVE (default) 17.744
23 RTMA (relaxed) 19.031 23 RTMA (relaxed) 18.565
24 MAIVE (WAIVE) 21.789 24 MAIVE (WAIVE) 21.798

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Methods are compared using condition-wise ranks. Direct comparison using the average empirical standard error is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of empirical standard error values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 6.388 1 RoBMA (PSMA) 6.557
2 AK (AK1) 7.422 2 AK (AK1) 7.303
3 MMPH (default) 7.462 3 SM (3PSM) 8.056
4 SM (3PSM) 7.960 4 puniform (star) 8.498
5 puniform (star) 8.427 5 MMPH (default) 8.754
6 WAAPWLS (default) 9.323 6 WAAPWLS (default) 9.277
7 trimfill (default) 10.377 7 SM (4PSM) 10.149
8 SM (4PSM) 10.420 8 trimfill (default) 10.255
9 PEESE (default) 11.228 9 PEESE (default) 11.202
10 WLS (default) 11.392 10 WLS (default) 11.297
11 PETPEESE (default) 11.605 11 AK (AK2) 11.408
12 AK (AK2) 11.893 12 PETPEESE (default) 11.581
13 EK (default) 12.303 13 RMA (default) 12.335
14 RMA (default) 12.319 14 EK (default) 12.414
15 MAIVE (default) 12.971 15 MAIVE (default) 13.022
16 PET (default) 13.273 16 PET (default) 13.381
17 WILS (default) 13.534 17 WILS (default) 13.578
18 FMA (default) 14.292 18 FMA (default) 14.283
19 puniform (default) 14.546 19 puniform (default) 14.508
20 MAIVE (WAIVE) 15.691 20 RTMA (relaxed) 15.613
21 RTMA (relaxed) 17.089 21 MAIVE (WAIVE) 15.768
22 MAN (default) 17.283 22 MAN (default) 17.623
23 mean (default) 18.352 23 mean (default) 18.688
24 pcurve (default) 23.974 24 pcurve (default) 23.974

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average Interval Score is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of empirical standard error values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.800 1 RoBMA (PSMA) 0.798
2 AK (AK2) 0.795 2 AK (AK2) 0.769
3 SM (4PSM) 0.765 3 SM (4PSM) 0.760
4 RTMA (relaxed) 0.758 4 puniform (star) 0.733
5 puniform (star) 0.733 5 SM (3PSM) 0.728
6 SM (3PSM) 0.733 6 MAIVE (default) 0.695
7 MAIVE (default) 0.695 7 MAIVE (WAIVE) 0.647
8 MMPH (default) 0.677 8 EK (default) 0.641
9 MAIVE (WAIVE) 0.647 9 PETPEESE (default) 0.629
10 EK (default) 0.641 10 MMPH (default) 0.621
11 PETPEESE (default) 0.629 11 PET (default) 0.620
12 PET (default) 0.620 12 AK (AK1) 0.609
13 AK (AK1) 0.609 13 WAAPWLS (default) 0.582
14 WAAPWLS (default) 0.582 14 RTMA (relaxed) 0.575
15 trimfill (default) 0.544 15 trimfill (default) 0.543
16 PEESE (default) 0.526 16 PEESE (default) 0.526
17 WILS (default) 0.504 17 WILS (default) 0.504
18 puniform (default) 0.484 18 puniform (default) 0.484
19 WLS (default) 0.464 19 WLS (default) 0.464
20 RMA (default) 0.457 20 RMA (default) 0.457
21 FMA (default) 0.342 21 FMA (default) 0.342
22 mean (default) 0.260 22 MAN (default) 0.273
23 MAN (default) 0.257 23 mean (default) 0.260
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Mean Rank Rank Method Mean Rank
1 FMA (default) 2.449 1 FMA (default) 2.437
2 WLS (default) 3.998 2 WLS (default) 3.950
3 WILS (default) 5.082 3 WILS (default) 5.129
4 mean (default) 7.581 4 mean (default) 7.660
5 PEESE (default) 7.679 5 PEESE (default) 7.744
6 WAAPWLS (default) 7.962 6 WAAPWLS (default) 7.995
7 RMA (default) 8.180 7 RMA (default) 8.153
8 trimfill (default) 8.202 8 trimfill (default) 8.179
9 AK (AK1) 9.412 9 AK (AK1) 9.603
10 PETPEESE (default) 10.139 10 PETPEESE (default) 10.333
11 RoBMA (PSMA) 11.693 11 MMPH (default) 11.225
12 MMPH (default) 11.804 12 RoBMA (PSMA) 12.011
13 puniform (default) 12.351 13 puniform (default) 12.503
14 SM (3PSM) 13.633 14 SM (3PSM) 13.993
15 PET (default) 14.775 15 MAN (default) 14.407
16 puniform (star) 14.897 16 PET (default) 15.100
17 MAN (default) 15.405 17 puniform (star) 15.406
18 EK (default) 16.064 18 EK (default) 16.434
19 AK (AK2) 16.952 19 AK (AK2) 16.562
20 SM (4PSM) 17.020 20 SM (4PSM) 17.261
21 MAIVE (default) 18.560 21 MAIVE (default) 18.953
22 MAIVE (WAIVE) 20.447 22 RTMA (relaxed) 19.613
23 RTMA (relaxed) 21.265 23 MAIVE (WAIVE) 20.898
24 pcurve (default) 23.974 24 pcurve (default) 23.974

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average CI width is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of CI width values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RoBMA (PSMA) 3.703 1 RoBMA (PSMA) 3.577
2 AK (AK2) 2.124 2 AK (AK2) 1.735
3 MMPH (default) 1.614 3 MAIVE (default) 1.579
4 MAIVE (default) 1.579 4 PETPEESE (default) 1.532
5 PETPEESE (default) 1.532 5 PET (default) 1.515
6 PET (default) 1.515 6 EK (default) 1.515
7 EK (default) 1.515 7 puniform (default) 1.501
8 puniform (default) 1.515 8 puniform (star) 1.325
9 puniform (star) 1.325 9 SM (3PSM) 1.321
10 SM (3PSM) 1.325 10 MMPH (default) 1.310
11 AK (AK1) 1.215 11 AK (AK1) 1.205
12 SM (4PSM) 1.156 12 SM (4PSM) 1.185
13 MAIVE (WAIVE) 1.111 13 RTMA (relaxed) 1.143
14 RMA (default) 0.998 14 MAIVE (WAIVE) 1.111
15 RTMA (relaxed) 0.948 15 RMA (default) 0.998
16 WAAPWLS (default) 0.945 16 WAAPWLS (default) 0.945
17 trimfill (default) 0.922 17 trimfill (default) 0.922
18 WILS (default) 0.871 18 WILS (default) 0.871
19 PEESE (default) 0.843 19 PEESE (default) 0.843
20 WLS (default) 0.790 20 WLS (default) 0.790
21 FMA (default) 0.503 21 FMA (default) 0.503
22 mean (default) 0.487 22 MAN (default) 0.500
23 MAN (default) 0.347 23 mean (default) 0.487
24 pcurve (default) NaN 24 pcurve (default) NaN

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 PETPEESE (default) -4.626 1 AK (AK2) -4.661
2 EK (default) -4.496 2 PETPEESE (default) -4.626
3 PET (default) -4.496 3 EK (default) -4.496
4 WAAPWLS (default) -4.042 4 PET (default) -4.496
5 MMPH (default) -3.951 5 MMPH (default) -4.144
6 PEESE (default) -3.890 6 WAAPWLS (default) -4.042
7 SM (3PSM) -3.560 7 PEESE (default) -3.890
8 WLS (default) -3.450 8 SM (3PSM) -3.593
9 trimfill (default) -3.445 9 WLS (default) -3.450
10 MAIVE (default) -3.394 10 trimfill (default) -3.446
11 puniform (default) -3.374 11 MAIVE (default) -3.394
12 RoBMA (PSMA) -3.331 12 puniform (default) -3.376
13 AK (AK1) -3.277 13 RoBMA (PSMA) -3.332
14 puniform (star) -3.208 14 AK (AK1) -3.281
15 AK (AK2) -3.158 15 puniform (star) -3.208
16 RMA (default) -3.121 16 RMA (default) -3.121
17 FMA (default) -3.058 17 FMA (default) -3.058
18 WILS (default) -3.037 18 WILS (default) -3.037
19 SM (4PSM) -2.636 19 SM (4PSM) -2.873
20 mean (default) -2.503 20 RTMA (relaxed) -2.740
21 MAIVE (WAIVE) -1.547 21 mean (default) -2.503
22 RTMA (relaxed) -1.499 22 MAIVE (WAIVE) -1.547
23 MAN (default) -0.717 23 MAN (default) -1.521
24 pcurve (default) NaN 24 pcurve (default) NaN

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.102 1 RoBMA (PSMA) 0.106
2 AK (AK2) 0.125 2 MAIVE (WAIVE) 0.176
3 MAIVE (WAIVE) 0.176 3 AK (AK2) 0.237
4 SM (4PSM) 0.245 4 SM (4PSM) 0.248
5 PET (default) 0.257 5 PET (default) 0.257
6 EK (default) 0.257 6 EK (default) 0.257
7 MAIVE (default) 0.264 7 MAIVE (default) 0.264
8 PETPEESE (default) 0.270 8 PETPEESE (default) 0.270
9 SM (3PSM) 0.277 9 SM (3PSM) 0.280
10 RTMA (relaxed) 0.281 10 puniform (star) 0.293
11 puniform (star) 0.293 11 WILS (default) 0.391
12 MMPH (default) 0.379 12 RTMA (relaxed) 0.412
13 WILS (default) 0.391 13 MMPH (default) 0.449
14 WAAPWLS (default) 0.523 14 WAAPWLS (default) 0.523
15 PEESE (default) 0.546 15 PEESE (default) 0.546
16 AK (AK1) 0.581 16 AK (AK1) 0.581
17 MAN (default) 0.583 17 trimfill (default) 0.586
18 trimfill (default) 0.586 18 puniform (default) 0.607
19 puniform (default) 0.608 19 MAN (default) 0.619
20 WLS (default) 0.621 20 WLS (default) 0.621
21 RMA (default) 0.622 21 RMA (default) 0.622
22 FMA (default) 0.772 22 FMA (default) 0.772
23 mean (default) 0.779 23 mean (default) 0.779
24 pcurve (default) NaN 24 pcurve (default) NaN

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 mean (default) 0.990 1 mean (default) 0.990
2 FMA (default) 0.989 2 FMA (default) 0.989
3 WLS (default) 0.976 3 WLS (default) 0.976
4 RMA (default) 0.974 4 RMA (default) 0.974
5 AK (AK1) 0.969 5 AK (AK1) 0.969
6 trimfill (default) 0.965 6 trimfill (default) 0.965
7 PEESE (default) 0.953 7 PEESE (default) 0.953
8 puniform (default) 0.939 8 puniform (default) 0.939
9 WAAPWLS (default) 0.934 9 WAAPWLS (default) 0.934
10 MMPH (default) 0.918 10 MMPH (default) 0.922
11 PETPEESE (default) 0.893 11 RTMA (relaxed) 0.913
12 EK (default) 0.873 12 PETPEESE (default) 0.893
13 PET (default) 0.873 13 AK (AK2) 0.885
14 WILS (default) 0.864 14 EK (default) 0.873
15 SM (3PSM) 0.828 15 PET (default) 0.873
16 AK (AK2) 0.812 16 WILS (default) 0.864
17 puniform (star) 0.808 17 MAN (default) 0.844
18 MAIVE (default) 0.779 18 SM (3PSM) 0.835
19 SM (4PSM) 0.754 19 puniform (star) 0.808
20 MAN (default) 0.747 20 MAIVE (default) 0.779
21 RoBMA (PSMA) 0.703 21 SM (4PSM) 0.772
22 RTMA (relaxed) 0.606 22 RoBMA (PSMA) 0.706
23 MAIVE (WAIVE) 0.498 23 MAIVE (WAIVE) 0.498
24 pcurve (default) NaN 24 pcurve (default) NaN

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Conditional on Method Convergence)

The results below are conditional on method convergence. Note that the methods might differ in convergence rate and are therefore not compared on the same data sets.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average RMSE is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of RMSE values on the corresponding outcome scale.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Methods are compared using condition-wise ranks. Direct comparison using the average bias is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Methods are compared using condition-wise ranks. Direct comparison using the empirical standard error is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the interval score is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average 95% CI width is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of 95% CI width values on the corresponding outcome scale.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Replacement in Case of Non-Convergence)

The results below incorporate method replacement to handle non-convergence. If a method fails to converge, its results are replaced with the results from a simpler method (e.g., random-effects meta-analysis without publication bias adjustment). This emulates what a data analyst may do in practice in case a method does not converge. However, note that these results do not correspond to “pure” method performance as they might combine multiple different methods. See Method Replacement Strategy for details of the method replacement specification.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average RMSE is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of RMSE values on the corresponding outcome scale.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Methods are compared using condition-wise ranks. Direct comparison using the average bias is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Methods are compared using condition-wise ranks. Direct comparison using the empirical standard error is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the interval score is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average 95% CI width is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of 95% CI width values on the corresponding outcome scale.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

Subset: Publication Bias Present

These results are based on Stanley (2017), Alinaghi (2018), Bom (2019), and Carter (2019) data-generating mechanisms with a total of 1143 conditions.

Average Performance

Method performance measures are aggregated across all simulated conditions to provide an overall impression of method performance. However, keep in mind that a method with a high overall ranking is not necessarily the “best” method for a particular application. To select a suitable method for your application, consider also non-aggregated performance measures in conditions most relevant to your application.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Mean Rank Rank Method Mean Rank
1 AK (AK1) 7.665 1 AK (AK1) 7.477
2 MMPH (default) 8.206 2 RoBMA (PSMA) 8.905
3 RoBMA (PSMA) 8.612 3 WAAPWLS (default) 8.950
4 WAAPWLS (default) 9.110 4 MMPH (default) 9.633
5 PEESE (default) 10.001 5 trimfill (default) 9.825
6 trimfill (default) 10.031 6 PEESE (default) 9.900
7 PETPEESE (default) 10.237 7 FMA (default) 10.070
8 FMA (default) 10.299 8 WLS (default) 10.074
9 WLS (default) 10.304 9 PETPEESE (default) 10.160
10 WILS (default) 10.614 10 WILS (default) 10.610
11 SM (3PSM) 11.044 11 SM (3PSM) 11.101
12 puniform (star) 11.565 12 puniform (star) 11.528
13 EK (default) 12.221 13 AK (AK2) 12.248
14 pcurve (default) 12.317 14 EK (default) 12.288
15 PET (default) 12.352 15 pcurve (default) 12.372
16 AK (AK2) 12.835 16 PET (default) 12.425
17 SM (4PSM) 13.798 17 SM (4PSM) 13.818
18 RMA (default) 14.125 18 RMA (default) 14.148
19 puniform (default) 14.749 19 puniform (default) 14.724
20 MAIVE (default) 14.857 20 MAIVE (default) 14.874
21 RTMA (relaxed) 17.411 21 RTMA (relaxed) 16.693
22 MAN (default) 18.187 22 MAN (default) 18.325
23 mean (default) 18.358 23 mean (default) 18.668
24 MAIVE (WAIVE) 18.822 24 MAIVE (WAIVE) 18.903

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average RMSE is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of RMSE values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Mean Rank Rank Method Mean Rank
1 PETPEESE (default) 8.761 1 PETPEESE (default) 8.945
2 WAAPWLS (default) 9.028 2 WAAPWLS (default) 9.088
3 RoBMA (PSMA) 9.325 3 RoBMA (PSMA) 9.454
4 AK (AK1) 9.719 4 AK (AK1) 9.823
5 PEESE (default) 9.829 5 PEESE (default) 9.958
6 SM (3PSM) 10.042 6 SM (3PSM) 10.189
7 WILS (default) 10.282 7 EK (default) 10.567
8 EK (default) 10.290 8 WILS (default) 10.591
9 MMPH (default) 10.306 9 PET (default) 10.703
10 PET (default) 10.421 10 puniform (star) 10.857
11 puniform (star) 10.662 11 trimfill (default) 11.226
12 trimfill (default) 11.104 12 MMPH (default) 11.444
13 SM (4PSM) 11.503 13 SM (4PSM) 11.718
14 FMA (default) 11.906 14 AK (AK2) 11.939
15 WLS (default) 11.913 15 FMA (default) 12.039
16 AK (AK2) 12.724 16 WLS (default) 12.046
17 MAIVE (default) 13.800 17 RTMA (relaxed) 12.731
18 pcurve (default) 13.815 18 pcurve (default) 14.041
19 puniform (default) 14.363 19 MAIVE (default) 14.086
20 MAIVE (WAIVE) 15.963 20 puniform (default) 14.539
21 RTMA (relaxed) 16.329 21 MAIVE (WAIVE) 16.127
22 RMA (default) 16.535 22 RMA (default) 17.108
23 MAN (default) 19.113 23 MAN (default) 17.979
24 mean (default) 20.061 24 mean (default) 20.597

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Methods are compared using condition-wise ranks. Direct comparison using the average bias is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Mean Rank Rank Method Mean Rank
1 RMA (default) 4.309 1 RMA (default) 4.038
2 AK (AK1) 5.176 2 AK (AK1) 5.245
3 WLS (default) 6.242 3 WLS (default) 5.976
4 FMA (default) 6.247 4 FMA (default) 5.981
5 trimfill (default) 7.320 5 trimfill (default) 7.178
6 pcurve (default) 7.877 6 pcurve (default) 7.766
7 MMPH (default) 7.972 7 mean (default) 8.155
8 mean (default) 8.346 8 MMPH (default) 9.458
9 WAAPWLS (default) 9.945 9 WAAPWLS (default) 9.688
10 RoBMA (PSMA) 10.946 10 PEESE (default) 11.710
11 PEESE (default) 11.904 11 RoBMA (PSMA) 11.965
12 puniform (default) 12.518 12 puniform (default) 12.392
13 SM (3PSM) 13.569 13 SM (3PSM) 13.255
14 WILS (default) 14.454 14 WILS (default) 14.140
15 MAN (default) 14.935 15 puniform (star) 14.623
16 puniform (star) 14.996 16 PETPEESE (default) 15.223
17 PETPEESE (default) 15.449 17 AK (AK2) 15.763
18 AK (AK2) 15.574 18 MAN (default) 16.318
19 SM (4PSM) 17.118 19 SM (4PSM) 16.801
20 EK (default) 17.180 20 EK (default) 17.013
21 PET (default) 17.243 21 PET (default) 17.083
22 MAIVE (default) 17.697 22 MAIVE (default) 17.605
23 RTMA (relaxed) 19.129 23 RTMA (relaxed) 18.811
24 MAIVE (WAIVE) 21.682 24 MAIVE (WAIVE) 21.640

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Methods are compared using condition-wise ranks. Direct comparison using the average empirical standard error is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of empirical standard error values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 6.944 1 RoBMA (PSMA) 7.144
2 MMPH (default) 7.611 2 AK (AK1) 7.685
3 AK (AK1) 7.900 3 puniform (star) 8.393
4 puniform (star) 8.370 4 SM (3PSM) 8.455
5 SM (3PSM) 8.379 5 MMPH (default) 9.257
6 WAAPWLS (default) 9.603 6 WAAPWLS (default) 9.495
7 trimfill (default) 10.276 7 trimfill (default) 10.085
8 SM (4PSM) 10.657 8 SM (4PSM) 10.368
9 PETPEESE (default) 11.037 9 PETPEESE (default) 10.932
10 PEESE (default) 11.073 10 PEESE (default) 10.989
11 EK (default) 11.346 11 AK (AK2) 11.089
12 AK (AK2) 11.462 12 EK (default) 11.431
13 WLS (default) 12.263 13 WLS (default) 12.124
14 PET (default) 12.297 14 PET (default) 12.384
15 MAIVE (default) 12.740 15 MAIVE (default) 12.740
16 WILS (default) 12.934 16 WILS (default) 12.909
17 puniform (default) 13.458 17 puniform (default) 13.384
18 RMA (default) 14.122 18 RMA (default) 14.160
19 FMA (default) 14.620 19 FMA (default) 14.572
20 MAIVE (WAIVE) 15.318 20 RTMA (relaxed) 14.932
21 RTMA (relaxed) 16.211 21 MAIVE (WAIVE) 15.297
22 MAN (default) 17.266 22 MAN (default) 17.705
23 mean (default) 19.703 23 mean (default) 20.058
24 pcurve (default) 23.970 24 pcurve (default) 23.970

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average Interval Score is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of empirical standard error values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RTMA (relaxed) 0.770 1 RoBMA (PSMA) 0.756
2 RoBMA (PSMA) 0.759 2 SM (4PSM) 0.717
3 AK (AK2) 0.754 3 AK (AK2) 0.716
4 SM (4PSM) 0.722 4 puniform (star) 0.688
5 SM (3PSM) 0.688 5 SM (3PSM) 0.682
6 puniform (star) 0.688 6 MAIVE (default) 0.632
7 MMPH (default) 0.636 7 EK (default) 0.611
8 MAIVE (default) 0.632 8 PETPEESE (default) 0.599
9 EK (default) 0.611 9 MAIVE (WAIVE) 0.588
10 PETPEESE (default) 0.599 10 PET (default) 0.588
11 MAIVE (WAIVE) 0.588 11 MMPH (default) 0.569
12 PET (default) 0.588 12 RTMA (relaxed) 0.530
13 AK (AK1) 0.526 13 AK (AK1) 0.526
14 WAAPWLS (default) 0.523 14 WAAPWLS (default) 0.523
15 trimfill (default) 0.485 15 trimfill (default) 0.484
16 WILS (default) 0.479 16 WILS (default) 0.479
17 puniform (default) 0.478 17 puniform (default) 0.478
18 PEESE (default) 0.467 18 PEESE (default) 0.467
19 WLS (default) 0.393 19 WLS (default) 0.393
20 RMA (default) 0.358 20 RMA (default) 0.358
21 FMA (default) 0.288 21 FMA (default) 0.288
22 MAN (default) 0.245 22 MAN (default) 0.257
23 mean (default) 0.148 23 mean (default) 0.148
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Mean Rank Rank Method Mean Rank
1 FMA (default) 2.411 1 FMA (default) 2.395
2 WLS (default) 4.081 2 WLS (default) 4.026
3 WILS (default) 5.843 3 WILS (default) 5.918
4 mean (default) 7.140 4 mean (default) 7.226
5 PEESE (default) 7.449 5 PEESE (default) 7.510
6 WAAPWLS (default) 8.002 6 WAAPWLS (default) 8.048
7 trimfill (default) 8.206 7 trimfill (default) 8.178
8 RMA (default) 8.273 8 RMA (default) 8.255
9 AK (AK1) 8.626 9 AK (AK1) 8.815
10 PETPEESE (default) 9.944 10 PETPEESE (default) 10.144
11 MMPH (default) 12.025 11 MMPH (default) 11.497
12 puniform (default) 12.029 12 puniform (default) 12.230
13 RoBMA (PSMA) 12.422 13 RoBMA (PSMA) 12.787
14 SM (3PSM) 13.989 14 SM (3PSM) 14.411
15 PET (default) 14.458 15 PET (default) 14.815
16 puniform (star) 14.898 16 MAN (default) 15.184
17 EK (default) 15.766 17 puniform (star) 15.448
18 MAN (default) 16.281 18 EK (default) 16.157
19 AK (AK2) 16.794 19 AK (AK2) 16.210
20 SM (4PSM) 17.297 20 SM (4PSM) 17.616
21 MAIVE (default) 18.130 21 MAIVE (default) 18.571
22 MAIVE (WAIVE) 20.186 22 RTMA (relaxed) 19.450
23 RTMA (relaxed) 21.341 23 MAIVE (WAIVE) 20.700
24 pcurve (default) 23.970 24 pcurve (default) 23.970

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average CI width is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of CI width values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RoBMA (PSMA) 3.072 1 RoBMA (PSMA) 2.925
2 AK (AK2) 1.862 2 puniform (default) 1.707
3 puniform (default) 1.712 3 PETPEESE (default) 1.411
4 MMPH (default) 1.494 4 PET (default) 1.402
5 PETPEESE (default) 1.411 5 EK (default) 1.401
6 PET (default) 1.402 6 AK (AK2) 1.346
7 EK (default) 1.401 7 MAIVE (default) 1.258
8 MAIVE (default) 1.258 8 puniform (star) 1.057
9 SM (3PSM) 1.058 9 SM (3PSM) 1.047
10 puniform (star) 1.057 10 MMPH (default) 1.035
11 SM (4PSM) 0.857 11 SM (4PSM) 0.891
12 AK (AK1) 0.816 12 RTMA (relaxed) 0.872
13 WILS (default) 0.780 13 AK (AK1) 0.814
14 RTMA (relaxed) 0.725 14 WILS (default) 0.780
15 trimfill (default) 0.710 15 trimfill (default) 0.710
16 WAAPWLS (default) 0.694 16 WAAPWLS (default) 0.694
17 MAIVE (WAIVE) 0.658 17 MAIVE (WAIVE) 0.658
18 RMA (default) 0.601 18 RMA (default) 0.601
19 PEESE (default) 0.592 19 PEESE (default) 0.592
20 WLS (default) 0.520 20 WLS (default) 0.520
21 FMA (default) 0.267 21 MAN (default) 0.378
22 MAN (default) 0.201 22 FMA (default) 0.267
23 mean (default) 0.143 23 mean (default) 0.143
24 pcurve (default) NaN 24 pcurve (default) NaN

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 PETPEESE (default) -4.416 1 PETPEESE (default) -4.416
2 PET (default) -4.295 2 PET (default) -4.295
3 EK (default) -4.295 3 EK (default) -4.295
4 MMPH (default) -3.810 4 AK (AK2) -4.176
5 WAAPWLS (default) -3.646 5 MMPH (default) -3.870
6 PEESE (default) -3.433 6 WAAPWLS (default) -3.646
7 puniform (default) -3.371 7 PEESE (default) -3.433
8 RoBMA (PSMA) -2.917 8 puniform (default) -3.370
9 MAIVE (default) -2.868 9 RoBMA (PSMA) -2.912
10 SM (3PSM) -2.841 10 SM (3PSM) -2.882
11 trimfill (default) -2.827 11 MAIVE (default) -2.868
12 WLS (default) -2.818 12 trimfill (default) -2.827
13 AK (AK2) -2.697 13 WLS (default) -2.818
14 AK (AK1) -2.628 14 AK (AK1) -2.629
15 puniform (star) -2.613 15 puniform (star) -2.613
16 WILS (default) -2.580 16 WILS (default) -2.580
17 RMA (default) -2.441 17 RMA (default) -2.441
18 FMA (default) -2.367 18 RTMA (relaxed) -2.381
19 SM (4PSM) -1.917 19 FMA (default) -2.367
20 mean (default) -1.692 20 SM (4PSM) -2.168
21 MAIVE (WAIVE) -1.366 21 mean (default) -1.692
22 RTMA (relaxed) -0.997 22 MAIVE (WAIVE) -1.366
23 MAN (default) -0.138 23 MAN (default) -1.055
24 pcurve (default) NaN 24 pcurve (default) NaN

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.124 1 RoBMA (PSMA) 0.129
2 AK (AK2) 0.150 2 MAIVE (WAIVE) 0.220
3 MAIVE (WAIVE) 0.220 3 SM (4PSM) 0.284
4 SM (4PSM) 0.279 4 PET (default) 0.293
5 PET (default) 0.293 5 EK (default) 0.294
6 EK (default) 0.294 6 AK (AK2) 0.307
7 RTMA (relaxed) 0.295 7 PETPEESE (default) 0.311
8 PETPEESE (default) 0.311 8 SM (3PSM) 0.323
9 SM (3PSM) 0.318 9 puniform (star) 0.327
10 puniform (star) 0.327 10 MAIVE (default) 0.328
11 MAIVE (default) 0.328 11 WILS (default) 0.404
12 WILS (default) 0.404 12 RTMA (relaxed) 0.477
13 MMPH (default) 0.451 13 MMPH (default) 0.550
14 MAN (default) 0.602 14 puniform (default) 0.618
15 puniform (default) 0.618 15 MAN (default) 0.646
16 WAAPWLS (default) 0.650 16 WAAPWLS (default) 0.650
17 PEESE (default) 0.664 17 PEESE (default) 0.664
18 trimfill (default) 0.706 18 trimfill (default) 0.706
19 AK (AK1) 0.731 19 AK (AK1) 0.731
20 WLS (default) 0.762 20 WLS (default) 0.762
21 RMA (default) 0.782 21 RMA (default) 0.782
22 FMA (default) 0.883 22 FMA (default) 0.883
23 mean (default) 0.931 23 mean (default) 0.931
24 pcurve (default) NaN 24 pcurve (default) NaN

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 mean (default) 0.996 1 mean (default) 0.996
2 FMA (default) 0.991 2 FMA (default) 0.991
3 RMA (default) 0.979 3 RMA (default) 0.979
4 WLS (default) 0.978 4 WLS (default) 0.978
5 AK (AK1) 0.974 5 AK (AK1) 0.974
6 trimfill (default) 0.971 6 trimfill (default) 0.971
7 PEESE (default) 0.953 7 PEESE (default) 0.953
8 WAAPWLS (default) 0.937 8 WAAPWLS (default) 0.937
9 puniform (default) 0.929 9 puniform (default) 0.929
10 MMPH (default) 0.912 10 MMPH (default) 0.918
11 PETPEESE (default) 0.886 11 RTMA (relaxed) 0.908
12 EK (default) 0.865 12 PETPEESE (default) 0.886
13 PET (default) 0.865 13 AK (AK2) 0.870
14 WILS (default) 0.839 14 EK (default) 0.865
15 SM (3PSM) 0.789 15 PET (default) 0.865
16 AK (AK2) 0.777 16 WILS (default) 0.839
17 puniform (star) 0.766 17 MAN (default) 0.831
18 MAIVE (default) 0.743 18 SM (3PSM) 0.798
19 MAN (default) 0.717 19 puniform (star) 0.766
20 SM (4PSM) 0.711 20 MAIVE (default) 0.743
21 RoBMA (PSMA) 0.665 21 SM (4PSM) 0.732
22 RTMA (relaxed) 0.536 22 RoBMA (PSMA) 0.669
23 MAIVE (WAIVE) 0.475 23 MAIVE (WAIVE) 0.475
24 pcurve (default) NaN 24 pcurve (default) NaN

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Conditional on Method Convergence)

The results below are conditional on method convergence. Note that the methods might differ in convergence rate and are therefore not compared on the same data sets.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average RMSE is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of RMSE values on the corresponding outcome scale.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Methods are compared using condition-wise ranks. Direct comparison using the average bias is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Methods are compared using condition-wise ranks. Direct comparison using the empirical standard error is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the interval score is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average 95% CI width is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of 95% CI width values on the corresponding outcome scale.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Replacement in Case of Non-Convergence)

The results below incorporate method replacement to handle non-convergence. If a method fails to converge, its results are replaced with the results from a simpler method (e.g., random-effects meta-analysis without publication bias adjustment). This emulates what a data analyst may do in practice in case a method does not converge. However, note that these results do not correspond to “pure” method performance as they might combine multiple different methods. See Method Replacement Strategy for details of the method replacement specification.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average RMSE is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of RMSE values on the corresponding outcome scale.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Methods are compared using condition-wise ranks. Direct comparison using the average bias is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Methods are compared using condition-wise ranks. Direct comparison using the empirical standard error is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the interval score is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average 95% CI width is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of 95% CI width values on the corresponding outcome scale.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

Subset: Publication Bias Absent

These results are based on Stanley (2017), Alinaghi (2018), Bom (2019), and Carter (2019) data-generating mechanisms with a total of 522 conditions.

Average Performance

Method performance measures are aggregated across all simulated conditions to provide an overall impression of method performance. However, keep in mind that a method with a high overall ranking is not necessarily the “best” method for a particular application. To select a suitable method for your application, consider also non-aggregated performance measures in conditions most relevant to your application.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Mean Rank Rank Method Mean Rank
1 RoBMA (PSMA) 6.328 1 RoBMA (PSMA) 6.416
2 AK (AK1) 6.908 2 AK (AK1) 7.084
3 MMPH (default) 7.594 3 FMA (default) 7.713
4 FMA (default) 7.751 4 WLS (default) 7.726
5 WLS (default) 7.764 5 MMPH (default) 8.054
6 RMA (default) 8.241 6 RMA (default) 8.220
7 WAAPWLS (default) 8.301 7 WAAPWLS (default) 8.282
8 SM (3PSM) 8.738 8 SM (3PSM) 8.868
9 trimfill (default) 9.996 9 trimfill (default) 10.008
10 puniform (star) 10.877 10 puniform (star) 10.962
11 PEESE (default) 11.282 11 PEESE (default) 11.328
12 WILS (default) 11.910 12 WILS (default) 12.029
13 PETPEESE (default) 12.649 13 PETPEESE (default) 12.755
14 SM (4PSM) 12.812 14 SM (4PSM) 12.916
15 mean (default) 13.207 15 mean (default) 13.412
16 AK (AK2) 14.490 16 AK (AK2) 13.709
17 EK (default) 15.270 17 EK (default) 15.444
18 PET (default) 15.418 18 PET (default) 15.590
19 MAIVE (default) 15.808 19 MAIVE (default) 15.992
20 pcurve (default) 17.385 20 pcurve (default) 17.414
21 puniform (default) 17.724 21 puniform (default) 17.738
22 MAN (default) 17.820 22 MAN (default) 17.787
23 RTMA (relaxed) 19.278 23 RTMA (relaxed) 17.939
24 MAIVE (WAIVE) 19.939 24 MAIVE (WAIVE) 20.107

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average RMSE is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of RMSE values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Mean Rank Rank Method Mean Rank
1 AK (AK1) 8.222 1 AK (AK1) 8.333
2 WAAPWLS (default) 8.352 2 WAAPWLS (default) 8.427
3 SM (3PSM) 8.910 3 SM (3PSM) 8.954
4 WLS (default) 8.931 4 WLS (default) 8.964
5 FMA (default) 8.960 5 FMA (default) 8.992
6 PEESE (default) 9.004 6 PEESE (default) 9.092
7 PETPEESE (default) 9.632 7 PETPEESE (default) 9.810
8 RMA (default) 9.849 8 RMA (default) 9.990
9 MAIVE (default) 10.088 9 MAIVE (default) 10.218
10 puniform (star) 10.623 10 puniform (star) 10.672
11 SM (4PSM) 10.701 11 SM (4PSM) 10.814
12 PET (default) 10.726 12 PET (default) 10.950
13 EK (default) 10.747 13 EK (default) 10.975
14 RoBMA (PSMA) 11.172 14 RoBMA (PSMA) 11.398
15 mean (default) 12.791 15 AK (AK2) 12.004
16 AK (AK2) 13.506 16 mean (default) 13.086
17 MMPH (default) 14.845 17 MMPH (default) 15.061
18 MAIVE (WAIVE) 14.912 18 MAIVE (WAIVE) 15.148
19 trimfill (default) 15.142 19 trimfill (default) 15.308
20 WILS (default) 15.322 20 WILS (default) 15.602
21 puniform (default) 17.241 21 puniform (default) 17.356
22 pcurve (default) 18.749 22 RTMA (relaxed) 18.579
23 RTMA (relaxed) 19.672 23 pcurve (default) 18.877
24 MAN (default) 19.814 24 MAN (default) 19.299

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Methods are compared using condition-wise ranks. Direct comparison using the average bias is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Mean Rank Rank Method Mean Rank
1 RMA (default) 4.383 1 RMA (default) 4.065
2 MMPH (default) 5.450 2 MMPH (default) 6.218
3 WLS (default) 6.728 3 WLS (default) 6.533
4 FMA (default) 6.734 4 FMA (default) 6.538
5 RoBMA (PSMA) 7.142 5 RoBMA (PSMA) 7.096
6 AK (AK1) 7.389 6 AK (AK1) 7.655
7 trimfill (default) 8.985 7 trimfill (default) 8.826
8 WAAPWLS (default) 10.057 8 WAAPWLS (default) 9.921
9 mean (default) 10.073 9 mean (default) 9.981
10 MAN (default) 10.692 10 SM (3PSM) 11.450
11 SM (3PSM) 11.462 11 MAN (default) 11.753
12 pcurve (default) 11.902 12 pcurve (default) 11.877
13 PEESE (default) 12.914 13 PEESE (default) 12.935
14 puniform (star) 13.496 14 puniform (star) 13.481
15 WILS (default) 13.575 15 WILS (default) 13.563
16 puniform (default) 13.854 16 puniform (default) 13.824
17 PETPEESE (default) 15.142 17 PETPEESE (default) 15.232
18 SM (4PSM) 15.573 18 SM (4PSM) 15.569
19 AK (AK2) 16.352 19 AK (AK2) 15.883
20 MAIVE (default) 17.939 20 RTMA (relaxed) 18.027
21 EK (default) 18.370 21 MAIVE (default) 18.048
22 PET (default) 18.485 22 EK (default) 18.400
23 RTMA (relaxed) 18.816 23 PET (default) 18.513
24 MAIVE (WAIVE) 22.021 24 MAIVE (WAIVE) 22.144

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Methods are compared using condition-wise ranks. Direct comparison using the average empirical standard error is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of empirical standard error values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 5.170 1 RoBMA (PSMA) 5.270
2 AK (AK1) 6.374 2 AK (AK1) 6.466
3 SM (3PSM) 7.042 3 SM (3PSM) 7.182
4 MMPH (default) 7.136 4 MMPH (default) 7.651
5 RMA (default) 8.372 5 RMA (default) 8.337
6 puniform (star) 8.552 6 puniform (star) 8.728
7 WAAPWLS (default) 8.709 7 WAAPWLS (default) 8.799
8 WLS (default) 9.485 8 WLS (default) 9.485
9 SM (4PSM) 9.900 9 SM (4PSM) 9.669
10 trimfill (default) 10.596 10 trimfill (default) 10.626
11 PEESE (default) 11.569 11 PEESE (default) 11.669
12 AK (AK2) 12.837 12 AK (AK2) 12.107
13 PETPEESE (default) 12.849 13 PETPEESE (default) 13.004
14 MAIVE (default) 13.475 14 MAIVE (default) 13.638
15 FMA (default) 13.573 15 FMA (default) 13.649
16 EK (default) 14.398 16 EK (default) 14.565
17 WILS (default) 14.847 17 WILS (default) 15.044
18 mean (default) 15.395 18 PET (default) 15.565
19 PET (default) 15.408 19 mean (default) 15.688
20 MAIVE (WAIVE) 16.510 20 MAIVE (WAIVE) 16.799
21 puniform (default) 16.927 21 puniform (default) 16.967
22 MAN (default) 17.322 22 RTMA (relaxed) 17.105
23 RTMA (relaxed) 19.013 23 MAN (default) 17.444
24 pcurve (default) 23.981 24 pcurve (default) 23.981

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average Interval Score is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of empirical standard error values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.888 1 RoBMA (PSMA) 0.888
2 AK (AK2) 0.879 2 AK (AK2) 0.879
3 SM (4PSM) 0.860 3 SM (4PSM) 0.856
4 MAIVE (default) 0.835 4 MAIVE (default) 0.835
5 puniform (star) 0.832 5 puniform (star) 0.832
6 SM (3PSM) 0.831 6 SM (3PSM) 0.830
7 AK (AK1) 0.791 7 AK (AK1) 0.790
8 MAIVE (WAIVE) 0.776 8 MAIVE (WAIVE) 0.776
9 MMPH (default) 0.766 9 MMPH (default) 0.733
10 RTMA (relaxed) 0.733 10 WAAPWLS (default) 0.711
11 WAAPWLS (default) 0.711 11 EK (default) 0.706
12 EK (default) 0.706 12 PETPEESE (default) 0.695
13 PETPEESE (default) 0.695 13 PET (default) 0.689
14 PET (default) 0.689 14 RMA (default) 0.675
15 RMA (default) 0.675 15 RTMA (relaxed) 0.673
16 trimfill (default) 0.673 16 trimfill (default) 0.673
17 PEESE (default) 0.656 17 PEESE (default) 0.656
18 WLS (default) 0.619 18 WLS (default) 0.619
19 WILS (default) 0.557 19 WILS (default) 0.557
20 mean (default) 0.505 20 mean (default) 0.505
21 puniform (default) 0.497 21 puniform (default) 0.499
22 FMA (default) 0.461 22 FMA (default) 0.461
23 MAN (default) 0.281 23 MAN (default) 0.304
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Mean Rank Rank Method Mean Rank
1 FMA (default) 2.533 1 FMA (default) 2.531
2 WILS (default) 3.414 2 WILS (default) 3.400
3 WLS (default) 3.814 3 WLS (default) 3.782
4 WAAPWLS (default) 7.875 4 WAAPWLS (default) 7.879
5 RMA (default) 7.975 5 RMA (default) 7.931
6 PEESE (default) 8.182 6 trimfill (default) 8.180
7 trimfill (default) 8.193 7 PEESE (default) 8.255
8 mean (default) 8.546 8 mean (default) 8.611
9 RoBMA (PSMA) 10.098 9 RoBMA (PSMA) 10.314
10 PETPEESE (default) 10.567 10 MMPH (default) 10.628
11 AK (AK1) 11.134 11 PETPEESE (default) 10.745
12 MMPH (default) 11.320 12 AK (AK1) 11.330
13 SM (3PSM) 12.854 13 MAN (default) 12.705
14 puniform (default) 13.056 14 SM (3PSM) 13.077
15 MAN (default) 13.487 15 puniform (default) 13.102
16 puniform (star) 14.895 16 puniform (star) 15.314
17 PET (default) 15.471 17 PET (default) 15.724
18 SM (4PSM) 16.414 18 SM (4PSM) 16.485
19 EK (default) 16.716 19 EK (default) 17.040
20 AK (AK2) 17.297 20 AK (AK2) 17.333
21 MAIVE (default) 19.500 21 MAIVE (default) 19.789
22 MAIVE (WAIVE) 21.017 22 RTMA (relaxed) 19.971
23 RTMA (relaxed) 21.100 23 MAIVE (WAIVE) 21.331
24 pcurve (default) 23.981 24 pcurve (default) 23.981

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average CI width is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of CI width values on the corresponding outcome scale.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RoBMA (PSMA) 5.048 1 RoBMA (PSMA) 4.967
2 AK (AK2) 2.632 2 AK (AK2) 2.489
3 MAIVE (default) 2.263 3 MAIVE (default) 2.263
4 MAIVE (WAIVE) 2.079 4 MAIVE (WAIVE) 2.079
5 AK (AK1) 2.066 5 AK (AK1) 2.038
6 puniform (star) 1.896 6 SM (3PSM) 1.903
7 SM (3PSM) 1.894 7 puniform (star) 1.896
8 MMPH (default) 1.872 8 MMPH (default) 1.880
9 RMA (default) 1.842 9 RMA (default) 1.842
10 SM (4PSM) 1.794 10 SM (4PSM) 1.814
11 PETPEESE (default) 1.790 11 PETPEESE (default) 1.790
12 EK (default) 1.759 12 EK (default) 1.759
13 PET (default) 1.758 13 PET (default) 1.758
14 WAAPWLS (default) 1.480 14 RTMA (relaxed) 1.718
15 RTMA (relaxed) 1.420 15 WAAPWLS (default) 1.480
16 PEESE (default) 1.380 16 PEESE (default) 1.380
17 trimfill (default) 1.375 17 trimfill (default) 1.375
18 WLS (default) 1.367 18 WLS (default) 1.367
19 mean (default) 1.222 19 mean (default) 1.222
20 puniform (default) 1.095 20 WILS (default) 1.065
21 WILS (default) 1.065 21 puniform (default) 1.063
22 FMA (default) 1.006 22 FMA (default) 1.006
23 MAN (default) 0.611 23 MAN (default) 0.720
24 pcurve (default) NaN 24 pcurve (default) NaN

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 SM (3PSM) -5.093 1 AK (AK2) -5.601
2 PETPEESE (default) -5.073 2 SM (3PSM) -5.111
3 EK (default) -4.924 3 PETPEESE (default) -5.073
4 PET (default) -4.924 4 EK (default) -4.924
5 WAAPWLS (default) -4.886 5 PET (default) -4.924
6 PEESE (default) -4.863 6 WAAPWLS (default) -4.886
7 WLS (default) -4.799 7 PEESE (default) -4.863
8 trimfill (default) -4.763 8 WLS (default) -4.799
9 AK (AK1) -4.662 9 trimfill (default) -4.764
10 RMA (default) -4.572 10 MMPH (default) -4.710
11 FMA (default) -4.531 11 AK (AK1) -4.670
12 MAIVE (default) -4.517 12 RMA (default) -4.572
13 puniform (star) -4.477 13 FMA (default) -4.531
14 MMPH (default) -4.251 14 MAIVE (default) -4.517
15 mean (default) -4.232 15 puniform (star) -4.477
16 RoBMA (PSMA) -4.216 16 SM (4PSM) -4.377
17 SM (4PSM) -4.170 17 mean (default) -4.232
18 AK (AK2) -4.051 18 RoBMA (PSMA) -4.227
19 WILS (default) -4.013 19 WILS (default) -4.013
20 puniform (default) -3.380 20 RTMA (relaxed) -3.500
21 RTMA (relaxed) -2.564 21 puniform (default) -3.388
22 MAIVE (WAIVE) -1.933 22 MAN (default) -2.364
23 MAN (default) -1.762 23 MAIVE (WAIVE) -1.933
24 pcurve (default) NaN 24 pcurve (default) NaN

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.053 1 RoBMA (PSMA) 0.054
2 AK (AK2) 0.071 2 MAIVE (WAIVE) 0.077
3 MAIVE (WAIVE) 0.077 3 AK (AK2) 0.090
4 MAIVE (default) 0.118 4 MAIVE (default) 0.118
5 SM (4PSM) 0.167 5 SM (4PSM) 0.167
6 PET (default) 0.172 6 PET (default) 0.172
7 EK (default) 0.172 7 EK (default) 0.172
8 PETPEESE (default) 0.176 8 PETPEESE (default) 0.176
9 SM (3PSM) 0.183 9 SM (3PSM) 0.183
10 puniform (star) 0.216 10 puniform (star) 0.216
11 MMPH (default) 0.219 11 MMPH (default) 0.224
12 WAAPWLS (default) 0.233 12 WAAPWLS (default) 0.233
13 AK (AK1) 0.236 13 AK (AK1) 0.236
14 RTMA (relaxed) 0.250 14 RMA (default) 0.255
15 RMA (default) 0.255 15 RTMA (relaxed) 0.263
16 PEESE (default) 0.276 16 PEESE (default) 0.276
17 WLS (default) 0.296 17 WLS (default) 0.296
18 trimfill (default) 0.310 18 trimfill (default) 0.310
19 WILS (default) 0.361 19 WILS (default) 0.361
20 mean (default) 0.430 20 mean (default) 0.430
21 FMA (default) 0.518 21 FMA (default) 0.518
22 MAN (default) 0.544 22 MAN (default) 0.564
23 puniform (default) 0.587 23 puniform (default) 0.582
24 pcurve (default) NaN 24 pcurve (default) NaN

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 FMA (default) 0.984 1 FMA (default) 0.984
2 mean (default) 0.978 2 mean (default) 0.978
3 WLS (default) 0.971 3 WLS (default) 0.971
4 RMA (default) 0.963 4 RMA (default) 0.963
5 AK (AK1) 0.960 5 puniform (default) 0.960
6 puniform (default) 0.959 6 AK (AK1) 0.960
7 trimfill (default) 0.953 7 trimfill (default) 0.953
8 PEESE (default) 0.952 8 PEESE (default) 0.952
9 MMPH (default) 0.931 9 MMPH (default) 0.932
10 WAAPWLS (default) 0.927 10 WAAPWLS (default) 0.927
11 WILS (default) 0.918 11 RTMA (relaxed) 0.925
12 SM (3PSM) 0.911 12 WILS (default) 0.918
13 PETPEESE (default) 0.910 13 AK (AK2) 0.915
14 puniform (star) 0.898 14 SM (3PSM) 0.914
15 EK (default) 0.889 15 PETPEESE (default) 0.910
16 PET (default) 0.889 16 puniform (star) 0.898
17 AK (AK2) 0.883 17 EK (default) 0.889
18 MAIVE (default) 0.857 18 PET (default) 0.889
19 SM (4PSM) 0.843 19 MAN (default) 0.866
20 MAN (default) 0.802 20 SM (4PSM) 0.858
21 RoBMA (PSMA) 0.784 21 MAIVE (default) 0.857
22 RTMA (relaxed) 0.756 22 RoBMA (PSMA) 0.785
23 MAIVE (WAIVE) 0.547 23 MAIVE (WAIVE) 0.547
24 pcurve (default) NaN 24 pcurve (default) NaN

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Conditional on Method Convergence)

The results below are conditional on method convergence. Note that the methods might differ in convergence rate and are therefore not compared on the same data sets.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average RMSE is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of RMSE values on the corresponding outcome scale.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Methods are compared using condition-wise ranks. Direct comparison using the average bias is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Methods are compared using condition-wise ranks. Direct comparison using the empirical standard error is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the interval score is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average 95% CI width is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of 95% CI width values on the corresponding outcome scale.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Replacement in Case of Non-Convergence)

The results below incorporate method replacement to handle non-convergence. If a method fails to converge, its results are replaced with the results from a simpler method (e.g., random-effects meta-analysis without publication bias adjustment). This emulates what a data analyst may do in practice in case a method does not converge. However, note that these results do not correspond to “pure” method performance as they might combine multiple different methods. See Method Replacement Strategy for details of the method replacement specification.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average RMSE is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of RMSE values on the corresponding outcome scale.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Methods are compared using condition-wise ranks. Direct comparison using the average bias is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Methods are compared using condition-wise ranks. Direct comparison using the empirical standard error is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the interval score is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of bias values on the corresponding outcome scale.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method. Methods are compared using condition-wise ranks. Direct comparison using the average 95% CI width is not possible because the data-generating mechanisms differ in the outcome scale. See the DGM-specific results (or subresults) to see the distribution of 95% CI width values on the corresponding outcome scale.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

Session Info

This report was compiled on Wed Sep 30 16:01:58 2026 (UTC) using the following computational environment

## R version 4.6.1 (2026-06-24)
## Platform: x86_64-pc-linux-gnu
## Running under: Ubuntu 24.04.5 LTS
## 
## Matrix products: default
## BLAS:   /usr/lib/x86_64-linux-gnu/openblas-pthread/libblas.so.3 
## LAPACK: /usr/lib/x86_64-linux-gnu/openblas-pthread/libopenblasp-r0.3.26.so;  LAPACK version 3.12.0
## 
## locale:
##  [1] LC_CTYPE=C.UTF-8       LC_NUMERIC=C           LC_TIME=C.UTF-8       
##  [4] LC_COLLATE=C.UTF-8     LC_MONETARY=C.UTF-8    LC_MESSAGES=C.UTF-8   
##  [7] LC_PAPER=C.UTF-8       LC_NAME=C              LC_ADDRESS=C          
## [10] LC_TELEPHONE=C         LC_MEASUREMENT=C.UTF-8 LC_IDENTIFICATION=C   
## 
## time zone: UTC
## tzcode source: system (glibc)
## 
## attached base packages:
## [1] stats     graphics  grDevices utils     datasets  methods   base     
## 
## other attached packages:
## [1] scales_1.4.0                   ggdist_3.3.3                  
## [3] ggplot2_4.0.3                  PublicationBiasBenchmark_0.3.0
## 
## loaded via a namespace (and not attached):
##  [1] gtable_0.3.6         xfun_0.61            bslib_0.12.0        
##  [4] htmlwidgets_1.6.4    lattice_0.22-9       vctrs_0.7.3         
##  [7] tools_4.6.1          Rdpack_2.6.6         generics_0.1.4      
## [10] curl_8.0.0           sandwich_3.1-3       tibble_3.3.1        
## [13] pkgconfig_2.0.3      RColorBrewer_1.1-3   S7_0.2.2            
## [16] desc_1.4.3           distributional_0.9.0 lifecycle_1.0.5     
## [19] compiler_4.6.1       farver_2.1.2         stringr_1.6.0       
## [22] textshaping_1.0.5    htmltools_0.5.9      sass_0.4.10         
## [25] clubSandwich_0.7.0   yaml_2.3.12          pillar_1.11.1       
## [28] pkgdown_2.2.1        jquerylib_0.1.4      cachem_1.1.0        
## [31] tidyselect_1.2.1     digest_0.6.39        stringi_1.8.9       
## [34] dplyr_1.2.1          purrr_1.2.2          labeling_0.4.3      
## [37] fastmap_1.2.0        grid_4.6.1           cli_3.6.6           
## [40] magrittr_2.0.5       triebeard_0.4.1      crul_1.6.0          
## [43] osfr_0.2.9           withr_3.0.3          rmarkdown_2.32      
## [46] httr_1.4.9           otel_0.2.0           ragg_1.5.2          
## [49] zoo_1.9-1            kableExtra_1.4.1     memoise_2.0.1       
## [52] evaluate_1.0.5       knitr_1.52           rbibutils_2.4.1     
## [55] viridisLite_0.4.3    rlang_1.3.0          urltools_1.7.3.1    
## [58] Rcpp_1.1.2           glue_1.8.1           httpcode_0.3.0      
## [61] xml2_1.6.0           svglite_2.2.2        rstudioapi_0.19.0   
## [64] jsonlite_2.0.0       R6_2.6.1             systemfonts_1.3.2   
## [67] fs_2.1.0