Skip to contents

Complete Results

These results are based on Alinaghi (2018) data-generating mechanism with a total of 81 conditions.

Average Performance

Method performance measures are aggregated across all simulated conditions to provide an overall impression of method performance. However, keep in mind that a method with a high overall ranking is not necessarily the “best” method for a particular application. To select a suitable method for your application, consider also non-aggregated performance measures in conditions most relevant to your application.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.216 1 RoBMA (PSMA) 0.216
2 AK (AK2) 0.229 2 trimfill (default) 0.236
3 trimfill (default) 0.236 3 AK (AK2) 0.245
4 AK (AK1) 0.255 4 AK (AK1) 0.255
5 SM (4PSM) 0.263 5 SM (4PSM) 0.263
6 MAIVE (WAIVE) 0.295 6 MAIVE (WAIVE) 0.295
7 SM (3PSM) 0.310 7 SM (3PSM) 0.310
8 puniform (star) 0.316 8 puniform (star) 0.316
9 MAIVE (default) 0.317 9 MAIVE (default) 0.317
10 RMA (default) 0.320 10 RMA (default) 0.320
11 MMPH (default) 0.330 11 MMPH (default) 0.330
12 FMA (default) 0.345 12 FMA (default) 0.345
12 WLS (default) 0.345 12 WLS (default) 0.345
14 PEESE (default) 0.359 14 PEESE (default) 0.359
15 PETPEESE (default) 0.363 15 PETPEESE (default) 0.363
16 WAAPWLS (default) 0.372 16 WAAPWLS (default) 0.372
17 EK (default) 0.437 17 EK (default) 0.437
18 PET (default) 0.438 18 PET (default) 0.438
19 mean (default) 0.496 19 mean (default) 0.496
20 WILS (default) 0.571 20 WILS (default) 0.571
21 puniform (default) 0.643 21 puniform (default) 0.643
22 RTMA (relaxed) 0.771 22 RTMA (relaxed) 0.755
23 MAN (default) 1.178 23 MAN (default) 1.180
24 pcurve (default) 1.376 24 pcurve (default) 1.376

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 SM (4PSM) 0.025 1 SM (4PSM) 0.025
2 PET (default) 0.078 2 PET (default) 0.078
3 EK (default) 0.080 3 EK (default) 0.080
4 AK (AK2) 0.086 4 trimfill (default) 0.096
5 trimfill (default) 0.096 5 SM (3PSM) 0.098
6 SM (3PSM) 0.098 6 RoBMA (PSMA) 0.099
7 RoBMA (PSMA) 0.099 7 AK (AK2) 0.108
8 puniform (star) 0.111 8 puniform (star) 0.111
9 PETPEESE (default) 0.112 9 PETPEESE (default) 0.112
10 WAAPWLS (default) 0.115 10 WAAPWLS (default) 0.115
11 PEESE (default) 0.116 11 PEESE (default) 0.116
12 FMA (default) 0.131 12 FMA (default) 0.131
12 WLS (default) 0.131 12 WLS (default) 0.131
14 MAIVE (WAIVE) 0.165 14 MAIVE (WAIVE) 0.165
15 AK (AK1) 0.183 15 AK (AK1) 0.182
16 WILS (default) -0.183 16 WILS (default) -0.183
17 MAIVE (default) 0.187 17 MAIVE (default) 0.187
18 MMPH (default) 0.213 18 MMPH (default) 0.217
19 RMA (default) 0.262 19 RMA (default) 0.262
20 RTMA (relaxed) 0.321 20 RTMA (relaxed) 0.339
21 mean (default) 0.429 21 mean (default) 0.429
22 puniform (default) 0.606 22 puniform (default) 0.606
23 MAN (default) -1.136 23 MAN (default) -1.121
24 pcurve (default) -1.219 24 pcurve (default) -1.219

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 pcurve (default) 0.056 1 pcurve (default) 0.056
2 RMA (default) 0.124 2 RMA (default) 0.124
3 AK (AK1) 0.132 3 AK (AK1) 0.132
4 RoBMA (PSMA) 0.138 4 RoBMA (PSMA) 0.138
5 mean (default) 0.149 5 mean (default) 0.149
6 puniform (default) 0.155 6 puniform (default) 0.155
7 trimfill (default) 0.158 7 trimfill (default) 0.158
8 puniform (star) 0.160 8 MMPH (default) 0.159
9 MMPH (default) 0.160 9 puniform (star) 0.160
10 SM (3PSM) 0.161 10 SM (3PSM) 0.161
11 MAN (default) 0.162 11 MAIVE (default) 0.187
12 MAIVE (default) 0.187 12 MAIVE (WAIVE) 0.188
13 MAIVE (WAIVE) 0.188 13 SM (4PSM) 0.191
14 SM (4PSM) 0.191 14 AK (AK2) 0.198
15 AK (AK2) 0.192 15 MAN (default) 0.210
16 FMA (default) 0.286 16 FMA (default) 0.286
17 WLS (default) 0.286 17 WLS (default) 0.286
18 PEESE (default) 0.307 18 PEESE (default) 0.307
19 PETPEESE (default) 0.312 19 PETPEESE (default) 0.312
20 WAAPWLS (default) 0.324 20 WAAPWLS (default) 0.324
21 EK (default) 0.395 21 EK (default) 0.395
22 PET (default) 0.395 22 PET (default) 0.395
23 WILS (default) 0.453 23 RTMA (relaxed) 0.439
24 RTMA (relaxed) 0.458 24 WILS (default) 0.453

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 2.220 1 RoBMA (PSMA) 2.220
2 AK (AK2) 3.037 2 MAIVE (WAIVE) 3.441
3 MAIVE (WAIVE) 3.441 3 AK (AK2) 3.550
4 FMA (default) 3.780 4 FMA (default) 3.780
5 SM (4PSM) 3.999 5 SM (4PSM) 3.999
6 MAIVE (default) 4.031 6 MAIVE (default) 4.031
7 trimfill (default) 4.897 7 trimfill (default) 4.897
8 AK (AK1) 5.772 8 AK (AK1) 5.764
9 RMA (default) 6.200 9 RMA (default) 6.200
10 SM (3PSM) 6.625 10 SM (3PSM) 6.625
11 puniform (star) 7.284 11 puniform (star) 7.284
12 MMPH (default) 7.645 12 MMPH (default) 7.615
13 WAAPWLS (default) 7.800 13 WAAPWLS (default) 7.800
14 WLS (default) 8.687 14 WLS (default) 8.687
15 PEESE (default) 8.983 15 PEESE (default) 8.983
16 PETPEESE (default) 9.060 16 PETPEESE (default) 9.060
17 EK (default) 10.612 17 EK (default) 10.612
18 PET (default) 10.651 18 PET (default) 10.651
19 RTMA (relaxed) 12.965 19 RTMA (relaxed) 12.999
20 mean (default) 14.940 20 mean (default) 14.940
21 WILS (default) 16.151 21 WILS (default) 16.151
22 puniform (default) 20.205 22 puniform (default) 20.205
23 MAN (default) 34.942 23 MAN (default) 34.875
24 pcurve (default) NaN 24 pcurve (default) NaN

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.868 1 RoBMA (PSMA) 0.868
2 AK (AK2) 0.802 2 AK (AK2) 0.759
3 SM (4PSM) 0.749 3 SM (4PSM) 0.749
4 MAIVE (WAIVE) 0.720 4 MAIVE (WAIVE) 0.720
5 MAIVE (default) 0.680 5 MAIVE (default) 0.680
6 AK (AK1) 0.652 6 AK (AK1) 0.651
7 SM (3PSM) 0.625 7 SM (3PSM) 0.625
8 trimfill (default) 0.614 8 trimfill (default) 0.614
9 RMA (default) 0.597 9 RMA (default) 0.597
10 puniform (star) 0.597 10 puniform (star) 0.597
11 FMA (default) 0.597 11 FMA (default) 0.597
12 RTMA (relaxed) 0.589 12 RTMA (relaxed) 0.563
13 WAAPWLS (default) 0.555 13 WAAPWLS (default) 0.555
14 MMPH (default) 0.490 14 MMPH (default) 0.481
15 PETPEESE (default) 0.467 15 PETPEESE (default) 0.467
16 PEESE (default) 0.455 16 PEESE (default) 0.455
17 WLS (default) 0.441 17 WLS (default) 0.441
18 EK (default) 0.412 18 EK (default) 0.412
19 PET (default) 0.361 19 PET (default) 0.361
20 puniform (default) 0.344 20 puniform (default) 0.344
21 WILS (default) 0.307 21 WILS (default) 0.307
22 mean (default) 0.299 22 mean (default) 0.299
23 MAN (default) 0.067 23 MAN (default) 0.067
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 WILS (default) 0.154 1 WILS (default) 0.154
2 WLS (default) 0.163 2 WLS (default) 0.163
3 PEESE (default) 0.168 3 PEESE (default) 0.168
4 PETPEESE (default) 0.170 4 PETPEESE (default) 0.170
5 EK (default) 0.208 5 EK (default) 0.208
6 PET (default) 0.208 6 PET (default) 0.208
7 trimfill (default) 0.229 7 trimfill (default) 0.229
8 AK (AK1) 0.246 8 AK (AK1) 0.246
9 mean (default) 0.247 9 mean (default) 0.247
10 WAAPWLS (default) 0.289 10 WAAPWLS (default) 0.289
11 MMPH (default) 0.290 11 puniform (star) 0.290
12 puniform (star) 0.290 12 MMPH (default) 0.291
13 SM (3PSM) 0.317 13 SM (3PSM) 0.317
14 puniform (default) 0.321 14 puniform (default) 0.321
15 SM (4PSM) 0.395 15 AK (AK2) 0.393
16 AK (AK2) 0.404 16 SM (4PSM) 0.395
17 RMA (default) 0.448 17 RMA (default) 0.448
18 RoBMA (PSMA) 0.494 18 RoBMA (PSMA) 0.494
19 MAN (default) 0.627 19 MAN (default) 0.626
20 MAIVE (WAIVE) 0.664 20 MAIVE (WAIVE) 0.664
21 MAIVE (default) 0.676 21 MAIVE (default) 0.676
22 FMA (default) 0.980 22 FMA (default) 0.980
23 RTMA (relaxed) 2.194 23 RTMA (relaxed) 1.890
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RoBMA (PSMA) 5.904 1 RoBMA (PSMA) 5.904
2 MAIVE (default) 2.518 2 MAIVE (default) 2.518
3 MAIVE (WAIVE) 2.467 3 MAIVE (WAIVE) 2.467
4 AK (AK2) 2.307 4 AK (AK2) 2.211
5 RMA (default) 2.161 5 RMA (default) 2.161
6 AK (AK1) 1.616 6 AK (AK1) 1.618
7 SM (4PSM) 1.521 7 SM (4PSM) 1.521
8 trimfill (default) 1.487 8 trimfill (default) 1.489
9 EK (default) 1.101 9 EK (default) 1.101
9 PET (default) 1.101 9 PET (default) 1.101
11 PETPEESE (default) 1.059 11 PETPEESE (default) 1.059
12 mean (default) 1.039 12 mean (default) 1.039
13 FMA (default) 1.010 13 FMA (default) 1.010
14 WAAPWLS (default) 0.905 14 WAAPWLS (default) 0.905
15 RTMA (relaxed) 0.848 15 RTMA (relaxed) 0.851
16 SM (3PSM) 0.819 16 SM (3PSM) 0.819
17 WLS (default) 0.811 17 WLS (default) 0.811
18 PEESE (default) 0.795 18 PEESE (default) 0.795
19 puniform (default) 0.749 19 MMPH (default) 0.761
20 puniform (star) 0.742 20 puniform (default) 0.749
21 MMPH (default) 0.635 21 puniform (star) 0.742
22 WILS (default) 0.446 22 WILS (default) 0.446
23 MAN (default) 0.072 23 MAN (default) 0.068
24 pcurve (default) NaN 24 pcurve (default) NaN

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RoBMA (PSMA) -6.322 1 AK (AK2) -6.345
2 MAIVE (default) -5.986 2 RoBMA (PSMA) -6.322
3 MAIVE (WAIVE) -5.825 3 MAIVE (default) -5.986
4 AK (AK2) -5.588 4 MAIVE (WAIVE) -5.825
5 SM (4PSM) -5.405 5 SM (4PSM) -5.405
6 WAAPWLS (default) -5.289 6 WAAPWLS (default) -5.289
7 PETPEESE (default) -5.199 7 PETPEESE (default) -5.199
8 EK (default) -5.148 8 EK (default) -5.148
8 PET (default) -5.148 8 PET (default) -5.148
10 RMA (default) -5.086 10 AK (AK1) -5.101
11 AK (AK1) -5.078 11 RMA (default) -5.086
12 trimfill (default) -4.937 12 trimfill (default) -4.936
13 mean (default) -4.588 13 MMPH (default) -4.864
14 WLS (default) -4.347 14 mean (default) -4.588
15 PEESE (default) -4.310 15 WLS (default) -4.347
16 MMPH (default) -4.148 16 PEESE (default) -4.310
17 SM (3PSM) -3.941 17 SM (3PSM) -3.941
18 FMA (default) -3.575 18 FMA (default) -3.575
19 puniform (star) -3.410 19 puniform (star) -3.410
20 WILS (default) -3.167 20 RTMA (relaxed) -3.191
21 RTMA (relaxed) -3.127 21 WILS (default) -3.167
22 puniform (default) -2.490 22 puniform (default) -2.490
23 MAN (default) -1.210 23 MAN (default) -1.201
24 pcurve (default) NaN 24 pcurve (default) NaN

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.092 1 RoBMA (PSMA) 0.092
2 AK (AK2) 0.196 2 AK (AK2) 0.206
3 MAIVE (default) 0.229 3 MAIVE (default) 0.229
4 MAIVE (WAIVE) 0.230 4 MAIVE (WAIVE) 0.230
5 RMA (default) 0.359 5 RMA (default) 0.359
6 SM (4PSM) 0.372 6 SM (4PSM) 0.372
7 AK (AK1) 0.452 7 AK (AK1) 0.452
8 trimfill (default) 0.484 8 trimfill (default) 0.484
9 FMA (default) 0.535 9 FMA (default) 0.535
10 PETPEESE (default) 0.536 10 PETPEESE (default) 0.536
11 EK (default) 0.545 11 EK (default) 0.545
11 PET (default) 0.545 11 PET (default) 0.545
13 mean (default) 0.552 13 mean (default) 0.552
14 WAAPWLS (default) 0.574 14 WAAPWLS (default) 0.574
15 RTMA (relaxed) 0.621 15 MMPH (default) 0.583
16 SM (3PSM) 0.637 16 RTMA (relaxed) 0.620
17 WLS (default) 0.645 17 SM (3PSM) 0.637
18 PEESE (default) 0.647 18 WLS (default) 0.645
19 MMPH (default) 0.654 19 PEESE (default) 0.647
20 puniform (star) 0.697 20 puniform (star) 0.697
21 puniform (default) 0.707 21 puniform (default) 0.707
22 WILS (default) 0.775 22 WILS (default) 0.775
23 MAN (default) 0.813 23 MAN (default) 0.815
24 pcurve (default) NaN 24 pcurve (default) NaN

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 puniform (default) 1.000 1 puniform (default) 1.000
2 mean (default) 0.998 2 mean (default) 0.998
3 AK (AK1) 0.997 3 AK (AK1) 0.997
4 MMPH (default) 0.994 4 MMPH (default) 0.994
5 WLS (default) 0.993 5 WLS (default) 0.993
6 PEESE (default) 0.993 6 PEESE (default) 0.993
7 trimfill (default) 0.992 7 trimfill (default) 0.992
8 EK (default) 0.988 8 EK (default) 0.988
8 PET (default) 0.988 8 PET (default) 0.988
10 PETPEESE (default) 0.986 10 PETPEESE (default) 0.986
11 RMA (default) 0.985 11 RMA (default) 0.985
12 puniform (star) 0.983 12 puniform (star) 0.983
13 WILS (default) 0.982 13 WILS (default) 0.982
14 WAAPWLS (default) 0.976 14 WAAPWLS (default) 0.976
15 AK (AK2) 0.974 15 AK (AK2) 0.974
16 SM (3PSM) 0.971 16 SM (3PSM) 0.971
17 MAIVE (default) 0.960 17 MAIVE (default) 0.960
18 MAIVE (WAIVE) 0.959 18 MAIVE (WAIVE) 0.959
19 SM (4PSM) 0.957 19 RTMA (relaxed) 0.957
20 RTMA (relaxed) 0.955 20 SM (4PSM) 0.957
21 RoBMA (PSMA) 0.945 21 RoBMA (PSMA) 0.945
22 FMA (default) 0.928 22 FMA (default) 0.928
23 MAN (default) 0.868 23 MAN (default) 0.869
24 pcurve (default) NaN 24 pcurve (default) NaN

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Conditional on Method Convergence)

The results below are conditional on method convergence. Note that the methods might differ in convergence rate and are therefore not compared on the same data sets.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Replacement in Case of Non-Convergence)

The results below incorporate method replacement to handle non-convergence. If a method fails to converge, its results are replaced with the results from a simpler method (e.g., random-effects meta-analysis without publication bias adjustment). This emulates what a data analyst may do in practice in case a method does not converge. However, note that these results do not correspond to “pure” method performance as they might combine multiple different methods. See Method Replacement Strategy for details of the method replacement specification.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

Subset: Fixed Effects

These results are based on Alinaghi (2018) data-generating mechanism with a total of 27 conditions.

Average Performance

Method performance measures are aggregated across all simulated conditions to provide an overall impression of method performance. However, keep in mind that a method with a high overall ranking is not necessarily the “best” method for a particular application. To select a suitable method for your application, consider also non-aggregated performance measures in conditions most relevant to your application.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.008 1 RoBMA (PSMA) 0.008
2 AK (AK2) 0.009 2 PETPEESE (default) 0.010
3 PETPEESE (default) 0.010 3 PEESE (default) 0.012
4 PEESE (default) 0.012 4 WAAPWLS (default) 0.012
5 WAAPWLS (default) 0.012 5 WLS (default) 0.014
6 WLS (default) 0.014 6 FMA (default) 0.014
7 FMA (default) 0.014 7 EK (default) 0.015
8 trimfill (default) 0.015 8 trimfill (default) 0.015
9 EK (default) 0.015 9 WILS (default) 0.016
10 WILS (default) 0.016 10 SM (4PSM) 0.017
11 SM (4PSM) 0.017 11 AK (AK2) 0.017
12 PET (default) 0.020 12 PET (default) 0.020
13 RMA (default) 0.022 13 RMA (default) 0.022
14 AK (AK1) 0.037 14 AK (AK1) 0.036
15 SM (3PSM) 0.041 15 MMPH (default) 0.038
16 MMPH (default) 0.043 16 SM (3PSM) 0.041
17 puniform (star) 0.047 17 puniform (star) 0.047
18 puniform (default) 0.080 18 puniform (default) 0.080
19 MAIVE (WAIVE) 0.085 19 MAIVE (WAIVE) 0.085
20 MAIVE (default) 0.114 20 MAIVE (default) 0.114
21 mean (default) 0.348 21 mean (default) 0.348
22 RTMA (relaxed) 0.523 22 RTMA (relaxed) 0.510
23 MAN (default) 0.769 23 MAN (default) 0.769
24 pcurve (default) 1.340 24 pcurve (default) 1.340

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 AK (AK2) 0.000 1 RoBMA (PSMA) 0.000
2 RoBMA (PSMA) 0.000 2 PETPEESE (default) -0.001
3 PETPEESE (default) -0.001 3 PEESE (default) 0.001
4 PEESE (default) 0.001 4 MMPH (default) -0.002
5 trimfill (default) 0.002 5 trimfill (default) 0.002
6 WILS (default) -0.003 6 WILS (default) -0.003
7 WAAPWLS (default) 0.003 7 WAAPWLS (default) 0.003
8 SM (4PSM) -0.006 8 AK (AK2) 0.003
9 FMA (default) 0.007 9 SM (4PSM) -0.006
10 WLS (default) 0.007 10 FMA (default) 0.007
11 EK (default) -0.008 11 WLS (default) 0.007
12 RMA (default) 0.012 12 EK (default) -0.008
13 PET (default) -0.013 13 RMA (default) 0.012
14 puniform (default) 0.018 14 PET (default) -0.013
15 SM (3PSM) 0.019 15 puniform (default) 0.018
16 MAIVE (WAIVE) 0.023 16 SM (3PSM) 0.019
17 puniform (star) 0.025 17 MAIVE (WAIVE) 0.023
18 AK (AK1) 0.028 18 puniform (star) 0.025
19 MMPH (default) -0.031 19 AK (AK1) 0.027
20 MAIVE (default) 0.038 20 MAIVE (default) 0.038
21 mean (default) 0.318 21 mean (default) 0.318
22 RTMA (relaxed) 0.369 22 RTMA (relaxed) 0.359
23 MAN (default) -0.756 23 MAN (default) -0.756
24 pcurve (default) -1.305 24 pcurve (default) -1.305

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.008 1 RoBMA (PSMA) 0.008
2 AK (AK2) 0.009 2 PEESE (default) 0.010
3 PEESE (default) 0.010 3 WLS (default) 0.010
4 WLS (default) 0.010 4 FMA (default) 0.010
5 FMA (default) 0.010 5 PETPEESE (default) 0.010
6 PETPEESE (default) 0.010 6 WAAPWLS (default) 0.010
7 WAAPWLS (default) 0.010 7 EK (default) 0.011
8 EK (default) 0.011 8 WILS (default) 0.011
9 WILS (default) 0.011 9 PET (default) 0.011
10 PET (default) 0.011 10 trimfill (default) 0.013
11 trimfill (default) 0.013 11 SM (4PSM) 0.014
12 SM (4PSM) 0.014 12 RMA (default) 0.015
13 RMA (default) 0.015 13 AK (AK2) 0.017
14 puniform (star) 0.017 14 puniform (star) 0.017
15 AK (AK1) 0.018 15 AK (AK1) 0.019
16 SM (3PSM) 0.021 16 SM (3PSM) 0.021
17 MMPH (default) 0.024 17 MMPH (default) 0.030
18 pcurve (default) 0.036 18 pcurve (default) 0.036
19 MAIVE (WAIVE) 0.063 19 MAIVE (WAIVE) 0.063
20 mean (default) 0.068 20 mean (default) 0.068
21 MAIVE (default) 0.070 21 MAIVE (default) 0.070
22 puniform (default) 0.075 22 puniform (default) 0.075
23 MAN (default) 0.090 23 MAN (default) 0.090
24 RTMA (relaxed) 0.247 24 RTMA (relaxed) 0.238

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.035 1 RoBMA (PSMA) 0.035
2 AK (AK2) 0.043 2 PETPEESE (default) 0.053
3 PETPEESE (default) 0.053 3 PEESE (default) 0.096
4 PEESE (default) 0.096 4 AK (AK2) 0.100
5 SM (4PSM) 0.101 5 SM (4PSM) 0.101
6 WAAPWLS (default) 0.108 6 WAAPWLS (default) 0.108
7 trimfill (default) 0.126 7 trimfill (default) 0.127
8 WLS (default) 0.131 8 WLS (default) 0.131
9 FMA (default) 0.136 9 FMA (default) 0.136
10 EK (default) 0.148 10 EK (default) 0.148
11 WILS (default) 0.172 11 WILS (default) 0.172
12 PET (default) 0.242 12 PET (default) 0.242
13 RMA (default) 0.289 13 RMA (default) 0.289
14 puniform (default) 0.397 14 MMPH (default) 0.290
15 MMPH (default) 0.513 15 puniform (default) 0.397
16 MAIVE (WAIVE) 0.681 16 AK (AK1) 0.668
17 AK (AK1) 0.692 17 MAIVE (WAIVE) 0.681
18 SM (3PSM) 0.728 18 SM (3PSM) 0.728
19 puniform (star) 1.047 19 puniform (star) 1.047
20 MAIVE (default) 1.207 20 MAIVE (default) 1.207
21 RTMA (relaxed) 9.741 21 RTMA (relaxed) 9.454
22 mean (default) 10.014 22 mean (default) 10.014
23 MAN (default) 23.894 23 MAN (default) 23.895
24 pcurve (default) NaN 24 pcurve (default) NaN

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.953 1 RoBMA (PSMA) 0.953
2 AK (AK2) 0.952 2 AK (AK2) 0.942
3 SM (4PSM) 0.939 3 SM (4PSM) 0.939
4 puniform (default) 0.927 4 puniform (default) 0.927
5 PETPEESE (default) 0.920 5 PETPEESE (default) 0.920
6 WAAPWLS (default) 0.909 6 WAAPWLS (default) 0.909
7 trimfill (default) 0.905 7 trimfill (default) 0.904
8 AK (AK1) 0.903 8 AK (AK1) 0.901
9 PEESE (default) 0.886 9 PEESE (default) 0.886
10 SM (3PSM) 0.885 10 SM (3PSM) 0.885
11 puniform (star) 0.879 11 puniform (star) 0.879
12 MMPH (default) 0.875 12 WLS (default) 0.849
13 WLS (default) 0.849 13 RMA (default) 0.839
14 RMA (default) 0.839 14 FMA (default) 0.829
15 FMA (default) 0.829 15 MMPH (default) 0.802
16 MAIVE (WAIVE) 0.789 16 MAIVE (WAIVE) 0.789
17 EK (default) 0.757 17 EK (default) 0.757
18 WILS (default) 0.748 18 WILS (default) 0.748
19 MAIVE (default) 0.727 19 MAIVE (default) 0.727
20 RTMA (relaxed) 0.607 20 RTMA (relaxed) 0.605
21 PET (default) 0.603 21 PET (default) 0.603
22 mean (default) 0.378 22 mean (default) 0.378
23 MAN (default) 0.079 23 MAN (default) 0.079
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.029 1 RoBMA (PSMA) 0.029
2 WILS (default) 0.033 2 WILS (default) 0.033
3 FMA (default) 0.034 3 FMA (default) 0.034
4 PEESE (default) 0.035 4 PEESE (default) 0.035
5 WAAPWLS (default) 0.036 5 WAAPWLS (default) 0.036
6 PETPEESE (default) 0.036 6 PETPEESE (default) 0.036
7 AK (AK2) 0.036 7 WLS (default) 0.036
8 WLS (default) 0.036 8 EK (default) 0.038
9 EK (default) 0.038 9 AK (AK2) 0.038
10 PET (default) 0.040 10 PET (default) 0.040
11 SM (4PSM) 0.045 11 SM (4PSM) 0.045
12 trimfill (default) 0.053 12 trimfill (default) 0.053
13 AK (AK1) 0.054 13 AK (AK1) 0.053
14 RMA (default) 0.056 14 RMA (default) 0.056
15 SM (3PSM) 0.058 15 SM (3PSM) 0.058
16 puniform (star) 0.060 16 puniform (star) 0.060
17 MMPH (default) 0.088 17 MMPH (default) 0.095
18 MAIVE (WAIVE) 0.220 18 MAIVE (WAIVE) 0.220
19 mean (default) 0.247 19 mean (default) 0.247
20 MAIVE (default) 0.267 20 MAIVE (default) 0.267
21 puniform (default) 0.326 21 puniform (default) 0.326
22 MAN (default) 0.349 22 MAN (default) 0.349
23 RTMA (relaxed) 1.753 23 RTMA (relaxed) 1.499
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RoBMA (PSMA) 7.601 1 RoBMA (PSMA) 7.601
2 MAIVE (WAIVE) 3.338 2 MAIVE (WAIVE) 3.338
3 MAIVE (default) 3.323 3 MAIVE (default) 3.323
4 AK (AK2) 3.114 4 AK (AK2) 2.832
5 EK (default) 2.803 5 EK (default) 2.803
5 PET (default) 2.803 5 PET (default) 2.803
7 PETPEESE (default) 2.631 7 PETPEESE (default) 2.631
8 RMA (default) 2.357 8 RMA (default) 2.357
9 puniform (default) 2.246 9 puniform (default) 2.246
10 AK (AK1) 2.233 10 AK (AK1) 2.236
11 trimfill (default) 2.164 11 trimfill (default) 2.169
12 SM (4PSM) 2.158 12 SM (4PSM) 2.158
13 WAAPWLS (default) 1.936 13 WAAPWLS (default) 1.936
14 WLS (default) 1.900 14 WLS (default) 1.900
15 PEESE (default) 1.860 15 PEESE (default) 1.860
16 mean (default) 1.671 16 mean (default) 1.671
17 FMA (default) 1.398 17 FMA (default) 1.398
18 RTMA (relaxed) 1.286 18 RTMA (relaxed) 1.287
19 WILS (default) 1.255 19 WILS (default) 1.255
20 puniform (star) 1.106 20 MMPH (default) 1.160
21 SM (3PSM) 1.100 21 puniform (star) 1.106
22 MAN (default) 0.680 22 SM (3PSM) 1.100
23 MMPH (default) 0.201 23 MAN (default) 0.679
24 pcurve (default) NaN 24 pcurve (default) NaN

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RoBMA (PSMA) -7.601 1 RoBMA (PSMA) -7.601
2 EK (default) -7.539 2 EK (default) -7.539
2 PET (default) -7.539 2 PET (default) -7.539
4 MAIVE (default) -7.533 4 MAIVE (default) -7.533
5 PETPEESE (default) -7.523 5 AK (AK2) -7.531
6 puniform (default) -7.469 6 PETPEESE (default) -7.523
7 SM (4PSM) -7.359 7 puniform (default) -7.469
8 MAIVE (WAIVE) -7.314 8 SM (4PSM) -7.359
9 WAAPWLS (default) -6.809 9 MAIVE (WAIVE) -7.314
10 AK (AK2) -6.178 10 MMPH (default) -7.226
11 AK (AK1) -6.073 11 WAAPWLS (default) -6.809
12 SM (3PSM) -5.936 12 AK (AK1) -6.141
13 RMA (default) -5.045 13 SM (3PSM) -5.936
14 trimfill (default) -5.043 14 RMA (default) -5.045
15 WLS (default) -5.028 15 trimfill (default) -5.040
16 PEESE (default) -5.026 16 WLS (default) -5.028
17 MAN (default) -5.026 17 PEESE (default) -5.026
18 mean (default) -4.983 18 MAN (default) -5.013
19 FMA (default) -4.956 19 mean (default) -4.983
20 WILS (default) -4.923 20 FMA (default) -4.956
21 puniform (star) -4.716 21 WILS (default) -4.923
22 RTMA (relaxed) -4.419 22 puniform (star) -4.716
23 MMPH (default) -1.668 23 RTMA (relaxed) -4.463
24 pcurve (default) NaN 24 pcurve (default) NaN

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.000 1 RoBMA (PSMA) 0.000
2 MAIVE (WAIVE) 0.038 2 MAIVE (WAIVE) 0.038
3 MAIVE (default) 0.039 3 MAIVE (default) 0.039
4 AK (AK2) 0.044 4 EK (default) 0.060
5 EK (default) 0.060 4 PET (default) 0.060
5 PET (default) 0.060 6 AK (AK2) 0.067
7 PETPEESE (default) 0.075 7 PETPEESE (default) 0.075
8 puniform (default) 0.121 8 puniform (default) 0.121
9 SM (4PSM) 0.187 9 SM (4PSM) 0.187
10 WAAPWLS (default) 0.337 10 MMPH (default) 0.313
11 AK (AK1) 0.353 11 WAAPWLS (default) 0.337
12 RMA (default) 0.356 12 AK (AK1) 0.353
13 trimfill (default) 0.360 13 RMA (default) 0.356
14 WLS (default) 0.372 14 trimfill (default) 0.360
15 PEESE (default) 0.374 15 WLS (default) 0.372
16 mean (default) 0.410 16 PEESE (default) 0.374
17 FMA (default) 0.433 17 mean (default) 0.410
18 WILS (default) 0.459 18 FMA (default) 0.433
19 RTMA (relaxed) 0.493 19 WILS (default) 0.459
20 SM (3PSM) 0.557 20 RTMA (relaxed) 0.492
21 puniform (star) 0.563 21 SM (3PSM) 0.557
22 MAN (default) 0.670 22 puniform (star) 0.563
23 MMPH (default) 0.797 23 MAN (default) 0.670
24 pcurve (default) NaN 24 pcurve (default) NaN

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 AK (AK1) 1.000 1 AK (AK1) 1.000
1 AK (AK2) 1.000 1 AK (AK2) 1.000
1 EK (default) 1.000 1 EK (default) 1.000
1 FMA (default) 1.000 1 FMA (default) 1.000
1 mean (default) 1.000 1 mean (default) 1.000
1 MMPH (default) 1.000 1 MMPH (default) 1.000
1 PEESE (default) 1.000 1 PEESE (default) 1.000
1 PET (default) 1.000 1 PET (default) 1.000
1 PETPEESE (default) 1.000 1 PETPEESE (default) 1.000
1 puniform (default) 1.000 1 puniform (default) 1.000
1 puniform (star) 1.000 1 puniform (star) 1.000
1 RMA (default) 1.000 1 RMA (default) 1.000
1 RoBMA (PSMA) 1.000 1 RoBMA (PSMA) 1.000
1 SM (3PSM) 1.000 1 SM (3PSM) 1.000
1 SM (4PSM) 1.000 1 SM (4PSM) 1.000
1 trimfill (default) 1.000 1 trimfill (default) 1.000
1 WAAPWLS (default) 1.000 1 WAAPWLS (default) 1.000
1 WILS (default) 1.000 1 WILS (default) 1.000
1 WLS (default) 1.000 1 WLS (default) 1.000
20 MAIVE (default) 1.000 20 MAIVE (default) 1.000
21 MAIVE (WAIVE) 0.999 21 MAIVE (WAIVE) 0.999
22 RTMA (relaxed) 0.986 22 RTMA (relaxed) 0.987
23 MAN (default) 0.969 23 MAN (default) 0.969
24 pcurve (default) NaN 24 pcurve (default) NaN

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Conditional on Method Convergence)

The results below are conditional on method convergence. Note that the methods might differ in convergence rate and are therefore not compared on the same data sets.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Replacement in Case of Non-Convergence)

The results below incorporate method replacement to handle non-convergence. If a method fails to converge, its results are replaced with the results from a simpler method (e.g., random-effects meta-analysis without publication bias adjustment). This emulates what a data analyst may do in practice in case a method does not converge. However, note that these results do not correspond to “pure” method performance as they might combine multiple different methods. See Method Replacement Strategy for details of the method replacement specification.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

Subset: Random Effects

These results are based on Alinaghi (2018) data-generating mechanism with a total of 27 conditions.

Average Performance

Method performance measures are aggregated across all simulated conditions to provide an overall impression of method performance. However, keep in mind that a method with a high overall ranking is not necessarily the “best” method for a particular application. To select a suitable method for your application, consider also non-aggregated performance measures in conditions most relevant to your application.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.098 1 RoBMA (PSMA) 0.098
2 AK (AK2) 0.101 2 AK (AK2) 0.120
3 trimfill (default) 0.151 3 trimfill (default) 0.151
4 AK (AK1) 0.173 4 AK (AK1) 0.173
5 MMPH (default) 0.183 5 MMPH (default) 0.183
6 SM (4PSM) 0.186 6 SM (4PSM) 0.186
7 MAIVE (WAIVE) 0.191 7 MAIVE (WAIVE) 0.191
8 PEESE (default) 0.198 8 PEESE (default) 0.198
9 PETPEESE (default) 0.199 9 PETPEESE (default) 0.199
10 FMA (default) 0.199 10 FMA (default) 0.199
10 WLS (default) 0.199 10 WLS (default) 0.199
12 WAAPWLS (default) 0.207 12 WAAPWLS (default) 0.207
13 MAIVE (default) 0.216 13 MAIVE (default) 0.216
14 EK (default) 0.223 14 EK (default) 0.223
14 PET (default) 0.223 14 PET (default) 0.223
16 RMA (default) 0.272 16 RMA (default) 0.272
17 SM (3PSM) 0.280 17 SM (3PSM) 0.280
18 puniform (star) 0.287 18 puniform (star) 0.287
19 puniform (default) 0.448 19 puniform (default) 0.448
20 mean (default) 0.461 20 mean (default) 0.461
21 WILS (default) 0.475 21 WILS (default) 0.475
22 RTMA (relaxed) 0.715 22 RTMA (relaxed) 0.705
23 MAN (default) 1.256 23 MAN (default) 1.254
24 pcurve (default) 1.405 24 pcurve (default) 1.405

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 AK (AK2) 0.003 1 SM (3PSM) 0.020
2 SM (3PSM) 0.020 2 AK (AK2) 0.028
3 EK (default) 0.033 3 EK (default) 0.033
3 PET (default) 0.033 3 PET (default) 0.033
5 puniform (star) 0.036 5 puniform (star) 0.036
6 RoBMA (PSMA) -0.040 6 RoBMA (PSMA) -0.040
7 PETPEESE (default) 0.070 7 PETPEESE (default) 0.070
8 PEESE (default) 0.070 8 PEESE (default) 0.070
9 WAAPWLS (default) 0.072 9 WAAPWLS (default) 0.072
10 FMA (default) 0.086 10 FMA (default) 0.086
11 WLS (default) 0.086 11 WLS (default) 0.086
12 trimfill (default) 0.093 12 trimfill (default) 0.093
13 SM (4PSM) -0.107 13 SM (4PSM) -0.107
14 MMPH (default) 0.115 14 MMPH (default) 0.115
15 AK (AK1) 0.131 15 AK (AK1) 0.131
16 MAIVE (WAIVE) 0.135 16 MAIVE (WAIVE) 0.135
17 MAIVE (default) 0.158 17 MAIVE (default) 0.158
18 RMA (default) 0.247 18 RMA (default) 0.247
19 RTMA (relaxed) 0.272 19 WILS (default) -0.283
20 WILS (default) -0.283 20 RTMA (relaxed) 0.286
21 mean (default) 0.430 21 mean (default) 0.430
22 puniform (default) 0.438 22 puniform (default) 0.438
23 MAN (default) -1.209 23 MAN (default) -1.204
24 pcurve (default) -1.236 24 pcurve (default) -1.236

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 pcurve (default) 0.034 1 pcurve (default) 0.034
2 RMA (default) 0.058 2 RMA (default) 0.058
3 MMPH (default) 0.060 3 MMPH (default) 0.060
4 AK (AK1) 0.064 4 AK (AK1) 0.064
5 trimfill (default) 0.069 5 trimfill (default) 0.069
6 mean (default) 0.076 6 mean (default) 0.076
7 puniform (default) 0.080 7 puniform (default) 0.080
8 SM (3PSM) 0.082 8 SM (3PSM) 0.082
9 MAIVE (WAIVE) 0.084 9 MAIVE (WAIVE) 0.084
10 MAIVE (default) 0.084 10 MAIVE (default) 0.084
11 puniform (star) 0.085 11 puniform (star) 0.085
12 RoBMA (PSMA) 0.085 12 RoBMA (PSMA) 0.085
13 AK (AK2) 0.101 13 SM (4PSM) 0.104
14 SM (4PSM) 0.104 14 AK (AK2) 0.110
15 FMA (default) 0.148 15 FMA (default) 0.148
16 WLS (default) 0.148 16 WLS (default) 0.148
17 MAN (default) 0.149 17 PEESE (default) 0.155
18 PEESE (default) 0.155 18 PETPEESE (default) 0.156
19 PETPEESE (default) 0.156 19 WAAPWLS (default) 0.166
20 WAAPWLS (default) 0.166 20 MAN (default) 0.169
21 EK (default) 0.190 21 EK (default) 0.190
21 PET (default) 0.190 21 PET (default) 0.190
23 WILS (default) 0.270 23 WILS (default) 0.270
24 RTMA (relaxed) 0.324 24 RTMA (relaxed) 0.314

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.426 1 RoBMA (PSMA) 0.426
2 AK (AK2) 0.471 2 AK (AK2) 0.887
3 SM (4PSM) 2.479 3 SM (4PSM) 2.479
4 trimfill (default) 2.707 4 trimfill (default) 2.707
5 WAAPWLS (default) 2.898 5 WAAPWLS (default) 2.898
6 MAIVE (WAIVE) 3.371 6 MAIVE (WAIVE) 3.371
7 MMPH (default) 3.798 7 MMPH (default) 3.797
8 AK (AK1) 3.950 8 AK (AK1) 3.950
9 MAIVE (default) 4.196 9 MAIVE (default) 4.196
10 PEESE (default) 4.234 10 PEESE (default) 4.234
11 PETPEESE (default) 4.247 11 PETPEESE (default) 4.247
12 WLS (default) 4.339 12 WLS (default) 4.339
13 EK (default) 4.512 13 EK (default) 4.512
14 PET (default) 4.516 14 PET (default) 4.516
15 FMA (default) 5.792 15 FMA (default) 5.792
16 SM (3PSM) 6.604 16 SM (3PSM) 6.604
17 puniform (star) 6.865 17 puniform (star) 6.865
18 RMA (default) 7.368 18 RMA (default) 7.368
19 puniform (default) 12.917 19 puniform (default) 12.917
20 RTMA (relaxed) 13.906 20 RTMA (relaxed) 13.966
21 WILS (default) 14.064 21 WILS (default) 14.064
22 mean (default) 14.386 22 mean (default) 14.386
23 MAN (default) 40.077 23 MAN (default) 39.973
24 pcurve (default) NaN 24 pcurve (default) NaN

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 AK (AK2) 0.951 1 RoBMA (PSMA) 0.941
2 RoBMA (PSMA) 0.941 2 AK (AK2) 0.866
3 SM (4PSM) 0.835 3 SM (4PSM) 0.835
4 AK (AK1) 0.719 4 AK (AK1) 0.719
5 MAIVE (WAIVE) 0.646 5 MAIVE (WAIVE) 0.646
6 trimfill (default) 0.626 6 trimfill (default) 0.626
7 MAIVE (default) 0.615 7 MAIVE (default) 0.615
8 SM (3PSM) 0.598 8 SM (3PSM) 0.598
9 RTMA (relaxed) 0.591 9 puniform (star) 0.586
10 puniform (star) 0.586 10 MMPH (default) 0.575
11 MMPH (default) 0.575 11 RTMA (relaxed) 0.556
12 WAAPWLS (default) 0.529 12 WAAPWLS (default) 0.529
13 RMA (default) 0.422 13 RMA (default) 0.422
14 mean (default) 0.342 14 mean (default) 0.342
15 PETPEESE (default) 0.335 15 PETPEESE (default) 0.335
16 EK (default) 0.335 16 EK (default) 0.335
17 PEESE (default) 0.335 17 PEESE (default) 0.335
18 PET (default) 0.335 18 PET (default) 0.335
19 WLS (default) 0.330 19 WLS (default) 0.330
20 FMA (default) 0.113 20 FMA (default) 0.113
21 puniform (default) 0.098 21 puniform (default) 0.098
22 WILS (default) 0.091 22 WILS (default) 0.091
23 MAN (default) 0.023 23 MAN (default) 0.024
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 FMA (default) 0.051 1 FMA (default) 0.051
2 WILS (default) 0.147 2 WILS (default) 0.147
3 WLS (default) 0.153 3 WLS (default) 0.153
4 PEESE (default) 0.156 4 PEESE (default) 0.156
5 PETPEESE (default) 0.157 5 PETPEESE (default) 0.157
6 PET (default) 0.185 6 PET (default) 0.185
7 EK (default) 0.185 7 EK (default) 0.185
8 RMA (default) 0.228 8 RMA (default) 0.228
9 trimfill (default) 0.234 9 trimfill (default) 0.234
10 MMPH (default) 0.239 10 MMPH (default) 0.239
11 mean (default) 0.244 11 mean (default) 0.244
12 AK (AK1) 0.248 12 AK (AK1) 0.248
13 puniform (default) 0.254 13 puniform (default) 0.254
14 WAAPWLS (default) 0.304 14 WAAPWLS (default) 0.304
15 RoBMA (PSMA) 0.310 15 RoBMA (PSMA) 0.310
16 SM (3PSM) 0.314 16 SM (3PSM) 0.314
17 MAIVE (WAIVE) 0.315 17 MAIVE (WAIVE) 0.315
18 puniform (star) 0.324 18 puniform (star) 0.324
19 MAIVE (default) 0.328 19 MAIVE (default) 0.328
20 AK (AK2) 0.392 20 AK (AK2) 0.383
21 SM (4PSM) 0.408 21 SM (4PSM) 0.408
22 MAN (default) 0.515 22 MAN (default) 0.514
23 RTMA (relaxed) 2.147 23 RTMA (relaxed) 1.862
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RoBMA (PSMA) 6.663 1 RoBMA (PSMA) 6.663
2 AK (AK2) 3.115 2 AK (AK2) 3.117
3 MAIVE (default) 2.279 3 MAIVE (default) 2.279
4 MAIVE (WAIVE) 2.224 4 MAIVE (WAIVE) 2.224
5 RMA (default) 2.150 5 RMA (default) 2.150
6 AK (AK1) 2.110 6 AK (AK1) 2.110
7 trimfill (default) 1.957 7 trimfill (default) 1.957
8 SM (4PSM) 1.835 8 SM (4PSM) 1.835
9 mean (default) 1.179 9 mean (default) 1.179
10 SM (3PSM) 0.935 10 SM (3PSM) 0.935
11 puniform (star) 0.930 11 puniform (star) 0.930
12 MMPH (default) 0.896 12 MMPH (default) 0.902
13 RTMA (relaxed) 0.851 13 RTMA (relaxed) 0.852
14 WAAPWLS (default) 0.598 14 WAAPWLS (default) 0.598
15 PETPEESE (default) 0.403 15 PETPEESE (default) 0.403
16 WLS (default) 0.399 16 WLS (default) 0.399
17 PEESE (default) 0.395 17 PEESE (default) 0.395
18 EK (default) 0.384 18 EK (default) 0.384
18 PET (default) 0.384 18 PET (default) 0.384
20 FMA (default) 0.098 20 FMA (default) 0.098
21 WILS (default) 0.044 21 WILS (default) 0.044
22 puniform (default) 0.000 22 puniform (default) 0.000
23 MAN (default) -0.206 23 MAN (default) -0.206
24 pcurve (default) NaN 24 pcurve (default) NaN

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RoBMA (PSMA) -6.364 1 AK (AK2) -6.830
2 AK (AK2) -6.064 2 RoBMA (PSMA) -6.364
3 WAAPWLS (default) -5.828 3 WAAPWLS (default) -5.828
4 MAIVE (default) -5.648 4 MAIVE (default) -5.648
5 MAIVE (WAIVE) -5.474 5 MAIVE (WAIVE) -5.474
6 RMA (default) -5.039 6 RMA (default) -5.039
7 AK (AK1) -5.039 7 AK (AK1) -5.039
8 trimfill (default) -5.023 8 trimfill (default) -5.023
9 mean (default) -4.918 9 mean (default) -4.918
10 PETPEESE (default) -4.860 10 PETPEESE (default) -4.860
11 EK (default) -4.812 11 EK (default) -4.812
11 PET (default) -4.812 11 PET (default) -4.812
13 MMPH (default) -4.440 13 MMPH (default) -4.460
14 WLS (default) -4.362 14 WLS (default) -4.362
15 PEESE (default) -4.345 15 PEESE (default) -4.345
16 SM (4PSM) -4.038 16 SM (4PSM) -4.038
17 FMA (default) -3.715 17 FMA (default) -3.715
18 WILS (default) -2.818 18 WILS (default) -2.818
19 SM (3PSM) -1.905 19 SM (3PSM) -1.905
20 puniform (star) -1.874 20 puniform (star) -1.874
21 RTMA (relaxed) -1.622 21 RTMA (relaxed) -1.669
22 puniform (default) 0.000 22 puniform (default) 0.000
23 MAN (default) 2.515 23 MAN (default) 2.512
24 pcurve (default) NaN 24 pcurve (default) NaN

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RoBMA (PSMA) 0.002 1 RoBMA (PSMA) 0.002
2 AK (AK2) 0.043 2 AK (AK2) 0.043
3 MAIVE (default) 0.353 3 MAIVE (default) 0.353
4 MAIVE (WAIVE) 0.356 4 MAIVE (WAIVE) 0.356
5 RMA (default) 0.361 5 RMA (default) 0.361
6 AK (AK1) 0.361 6 AK (AK1) 0.361
7 SM (4PSM) 0.369 7 SM (4PSM) 0.369
8 trimfill (default) 0.376 8 trimfill (default) 0.376
9 mean (default) 0.464 9 mean (default) 0.464
10 WAAPWLS (default) 0.594 10 WAAPWLS (default) 0.594
11 MMPH (default) 0.618 11 MMPH (default) 0.614
12 puniform (star) 0.684 12 puniform (star) 0.684
12 SM (3PSM) 0.684 12 SM (3PSM) 0.684
14 RTMA (relaxed) 0.686 14 RTMA (relaxed) 0.686
15 PETPEESE (default) 0.702 15 PETPEESE (default) 0.702
16 WLS (default) 0.704 16 WLS (default) 0.704
17 PEESE (default) 0.707 17 PEESE (default) 0.707
18 EK (default) 0.712 18 EK (default) 0.712
18 PET (default) 0.712 18 PET (default) 0.712
20 FMA (default) 0.909 20 FMA (default) 0.909
21 WILS (default) 0.937 21 WILS (default) 0.937
22 MAN (default) 0.996 22 MAN (default) 0.996
23 puniform (default) 1.000 23 puniform (default) 1.000
24 pcurve (default) NaN 24 pcurve (default) NaN

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 AK (AK1) 1.000 1 AK (AK1) 1.000
1 mean (default) 1.000 1 mean (default) 1.000
1 puniform (default) 1.000 1 puniform (default) 1.000
1 RMA (default) 1.000 1 RMA (default) 1.000
1 trimfill (default) 1.000 1 trimfill (default) 1.000
6 FMA (default) 1.000 6 FMA (default) 1.000
7 MMPH (default) 1.000 7 MMPH (default) 1.000
8 WLS (default) 1.000 8 WLS (default) 1.000
9 PEESE (default) 1.000 9 PEESE (default) 1.000
10 MAIVE (default) 0.999 10 MAIVE (default) 0.999
11 MAIVE (WAIVE) 0.998 11 MAIVE (WAIVE) 0.998
12 PETPEESE (default) 0.997 12 PETPEESE (default) 0.997
13 EK (default) 0.997 13 EK (default) 0.997
13 PET (default) 0.997 13 PET (default) 0.997
15 WAAPWLS (default) 0.981 15 WAAPWLS (default) 0.981
16 AK (AK2) 0.978 16 AK (AK2) 0.978
17 WILS (default) 0.978 17 WILS (default) 0.978
18 SM (3PSM) 0.962 18 SM (3PSM) 0.962
19 puniform (star) 0.958 19 puniform (star) 0.958
20 RoBMA (PSMA) 0.944 20 RoBMA (PSMA) 0.944
21 RTMA (relaxed) 0.944 21 RTMA (relaxed) 0.944
22 SM (4PSM) 0.935 22 SM (4PSM) 0.935
23 MAN (default) 0.896 23 MAN (default) 0.896
24 pcurve (default) NaN 24 pcurve (default) NaN

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Conditional on Method Convergence)

The results below are conditional on method convergence. Note that the methods might differ in convergence rate and are therefore not compared on the same data sets.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Replacement in Case of Non-Convergence)

The results below incorporate method replacement to handle non-convergence. If a method fails to converge, its results are replaced with the results from a simpler method (e.g., random-effects meta-analysis without publication bias adjustment). This emulates what a data analyst may do in practice in case a method does not converge. However, note that these results do not correspond to “pure” method performance as they might combine multiple different methods. See Method Replacement Strategy for details of the method replacement specification.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

Subset: Panel Random Effects

These results are based on Alinaghi (2018) data-generating mechanism with a total of 27 conditions.

Average Performance

Method performance measures are aggregated across all simulated conditions to provide an overall impression of method performance. However, keep in mind that a method with a high overall ranking is not necessarily the “best” method for a particular application. To select a suitable method for your application, consider also non-aggregated performance measures in conditions most relevant to your application.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 trimfill (default) 0.542 1 trimfill (default) 0.542
2 RoBMA (PSMA) 0.543 2 RoBMA (PSMA) 0.543
3 AK (AK1) 0.555 3 AK (AK1) 0.555
4 MMPH (default) 0.563 4 MMPH (default) 0.563
5 AK (AK2) 0.575 5 SM (4PSM) 0.587
6 SM (4PSM) 0.587 6 AK (AK2) 0.598
7 SM (3PSM) 0.608 7 SM (3PSM) 0.608
8 MAIVE (WAIVE) 0.610 8 MAIVE (WAIVE) 0.610
9 puniform (star) 0.615 9 puniform (star) 0.615
10 MAIVE (default) 0.622 10 MAIVE (default) 0.622
11 RMA (default) 0.667 11 RMA (default) 0.667
12 mean (default) 0.678 12 mean (default) 0.678
13 FMA (default) 0.824 13 FMA (default) 0.824
13 WLS (default) 0.824 13 WLS (default) 0.824
15 PEESE (default) 0.868 15 PEESE (default) 0.868
16 PETPEESE (default) 0.879 16 PETPEESE (default) 0.879
17 WAAPWLS (default) 0.897 17 WAAPWLS (default) 0.897
18 PET (default) 1.071 18 RTMA (relaxed) 1.048
19 EK (default) 1.071 19 PET (default) 1.071
20 RTMA (relaxed) 1.074 20 EK (default) 1.071
21 WILS (default) 1.222 21 WILS (default) 1.222
22 pcurve (default) 1.381 22 pcurve (default) 1.381
23 puniform (default) 1.400 23 puniform (default) 1.400
24 MAN (default) 1.508 24 MAN (default) 1.517

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 SM (4PSM) 0.190 1 SM (4PSM) 0.190
2 trimfill (default) 0.194 2 trimfill (default) 0.194
3 PET (default) 0.214 3 PET (default) 0.214
4 EK (default) 0.214 4 EK (default) 0.214
5 SM (3PSM) 0.256 5 SM (3PSM) 0.256
6 AK (AK2) 0.256 6 WILS (default) -0.263
7 WILS (default) -0.263 7 PETPEESE (default) 0.266
8 PETPEESE (default) 0.266 8 WAAPWLS (default) 0.270
9 WAAPWLS (default) 0.270 9 puniform (star) 0.272
10 puniform (star) 0.272 10 PEESE (default) 0.276
11 PEESE (default) 0.276 11 AK (AK2) 0.293
12 WLS (default) 0.301 12 WLS (default) 0.301
13 FMA (default) 0.301 13 FMA (default) 0.301
14 RTMA (relaxed) 0.323 14 RoBMA (PSMA) 0.336
15 RoBMA (PSMA) 0.336 15 MAIVE (WAIVE) 0.338
16 MAIVE (WAIVE) 0.338 16 MAIVE (default) 0.365
17 MAIVE (default) 0.365 17 RTMA (relaxed) 0.371
18 MMPH (default) 0.383 18 MMPH (default) 0.383
19 AK (AK1) 0.389 19 AK (AK1) 0.389
20 RMA (default) 0.528 20 RMA (default) 0.528
21 mean (default) 0.538 21 mean (default) 0.538
22 pcurve (default) -1.115 22 pcurve (default) -1.115
23 puniform (default) 1.362 23 puniform (default) 1.362
24 MAN (default) -1.443 24 MAN (default) -1.403

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 pcurve (default) 0.098 1 pcurve (default) 0.098
2 MAN (default) 0.246 2 MMPH (default) 0.295
3 MMPH (default) 0.295 3 RMA (default) 0.299
4 RMA (default) 0.299 4 mean (default) 0.302
5 mean (default) 0.302 5 puniform (default) 0.309
6 puniform (default) 0.309 6 AK (AK1) 0.313
7 AK (AK1) 0.313 7 RoBMA (PSMA) 0.321
8 RoBMA (PSMA) 0.321 8 MAN (default) 0.371
9 puniform (star) 0.378 9 puniform (star) 0.378
10 SM (3PSM) 0.380 10 SM (3PSM) 0.380
11 trimfill (default) 0.393 11 trimfill (default) 0.393
12 MAIVE (default) 0.408 12 MAIVE (default) 0.408
13 MAIVE (WAIVE) 0.417 13 MAIVE (WAIVE) 0.417
14 SM (4PSM) 0.454 14 SM (4PSM) 0.454
15 AK (AK2) 0.466 15 AK (AK2) 0.467
16 FMA (default) 0.699 16 FMA (default) 0.699
17 WLS (default) 0.699 17 WLS (default) 0.699
18 PEESE (default) 0.758 18 PEESE (default) 0.758
19 PETPEESE (default) 0.771 19 RTMA (relaxed) 0.765
20 WAAPWLS (default) 0.795 20 PETPEESE (default) 0.771
21 RTMA (relaxed) 0.803 21 WAAPWLS (default) 0.795
22 PET (default) 0.985 22 PET (default) 0.985
23 EK (default) 0.985 23 EK (default) 0.985
24 WILS (default) 1.079 24 WILS (default) 1.079

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 FMA (default) 5.413 1 FMA (default) 5.413
2 RoBMA (PSMA) 6.200 2 RoBMA (PSMA) 6.200
3 MAIVE (WAIVE) 6.273 3 MAIVE (WAIVE) 6.273
4 MAIVE (default) 6.690 4 MAIVE (default) 6.690
5 AK (AK2) 8.596 5 SM (4PSM) 9.418
6 SM (4PSM) 9.418 6 AK (AK2) 9.662
7 RMA (default) 10.943 7 RMA (default) 10.943
8 trimfill (default) 11.857 8 trimfill (default) 11.857
9 SM (3PSM) 12.542 9 SM (3PSM) 12.542
10 AK (AK1) 12.674 10 AK (AK1) 12.674
11 MMPH (default) 13.605 11 MMPH (default) 13.604
12 puniform (star) 13.940 12 puniform (star) 13.940
13 RTMA (relaxed) 15.249 13 RTMA (relaxed) 15.577
14 WAAPWLS (default) 20.395 14 WAAPWLS (default) 20.395
15 mean (default) 20.420 15 mean (default) 20.420
16 WLS (default) 21.589 16 WLS (default) 21.589
17 PEESE (default) 22.620 17 PEESE (default) 22.620
18 PETPEESE (default) 22.879 18 PETPEESE (default) 22.879
19 EK (default) 27.177 19 EK (default) 27.177
20 PET (default) 27.193 20 PET (default) 27.193
21 WILS (default) 34.217 21 WILS (default) 34.217
22 MAN (default) 40.856 22 MAN (default) 40.758
23 puniform (default) 47.302 23 puniform (default) 47.302
24 pcurve (default) NaN 24 pcurve (default) NaN

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 FMA (default) 0.848 1 FMA (default) 0.848
2 MAIVE (WAIVE) 0.724 2 MAIVE (WAIVE) 0.724
3 RoBMA (PSMA) 0.710 3 RoBMA (PSMA) 0.710
4 MAIVE (default) 0.697 4 MAIVE (default) 0.697
5 RTMA (relaxed) 0.569 5 RMA (default) 0.531
6 RMA (default) 0.531 6 RTMA (relaxed) 0.528
7 AK (AK2) 0.502 7 SM (4PSM) 0.474
8 SM (4PSM) 0.474 8 AK (AK2) 0.467
9 SM (3PSM) 0.392 9 SM (3PSM) 0.392
10 AK (AK1) 0.334 10 AK (AK1) 0.334
11 puniform (star) 0.326 11 puniform (star) 0.326
12 trimfill (default) 0.313 12 trimfill (default) 0.313
13 MMPH (default) 0.291 13 MMPH (default) 0.291
14 WAAPWLS (default) 0.226 14 WAAPWLS (default) 0.226
15 mean (default) 0.175 15 mean (default) 0.175
16 WLS (default) 0.145 16 WLS (default) 0.145
17 PETPEESE (default) 0.145 17 PETPEESE (default) 0.145
18 EK (default) 0.145 18 EK (default) 0.145
19 PET (default) 0.144 19 PET (default) 0.144
20 PEESE (default) 0.144 20 PEESE (default) 0.144
21 MAN (default) 0.099 21 MAN (default) 0.098
22 WILS (default) 0.081 22 WILS (default) 0.081
23 puniform (default) 0.006 23 puniform (default) 0.006
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 mean (default) 0.251 1 mean (default) 0.251
2 WILS (default) 0.283 2 WILS (default) 0.283
3 WLS (default) 0.299 3 WLS (default) 0.299
4 PEESE (default) 0.312 4 PEESE (default) 0.312
5 PETPEESE (default) 0.318 5 PETPEESE (default) 0.318
6 puniform (default) 0.383 6 puniform (default) 0.383
7 trimfill (default) 0.401 7 trimfill (default) 0.401
8 PET (default) 0.401 8 PET (default) 0.401
9 MMPH (default) 0.401 9 MMPH (default) 0.402
10 EK (default) 0.402 10 EK (default) 0.402
11 AK (AK1) 0.437 11 AK (AK1) 0.437
12 puniform (star) 0.487 12 puniform (star) 0.487
13 WAAPWLS (default) 0.527 13 WAAPWLS (default) 0.527
14 SM (3PSM) 0.579 14 SM (3PSM) 0.579
15 SM (4PSM) 0.733 15 SM (4PSM) 0.733
16 AK (AK2) 0.784 16 AK (AK2) 0.759
17 MAN (default) 1.016 17 MAN (default) 1.014
18 RMA (default) 1.060 18 RMA (default) 1.060
19 RoBMA (PSMA) 1.144 19 RoBMA (PSMA) 1.144
20 MAIVE (default) 1.433 20 MAIVE (default) 1.433
21 MAIVE (WAIVE) 1.456 21 MAIVE (WAIVE) 1.456
22 RTMA (relaxed) 2.681 22 RTMA (relaxed) 2.309
23 FMA (default) 2.855 23 FMA (default) 2.855
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RoBMA (PSMA) 3.448 1 RoBMA (PSMA) 3.448
2 RMA (default) 1.974 2 RMA (default) 1.974
3 MAIVE (default) 1.952 3 MAIVE (default) 1.952
4 MAIVE (WAIVE) 1.839 4 MAIVE (WAIVE) 1.839
5 FMA (default) 1.533 5 FMA (default) 1.533
6 AK (AK2) 0.691 6 AK (AK2) 0.686
7 SM (4PSM) 0.570 7 SM (4PSM) 0.570
8 AK (AK1) 0.506 8 AK (AK1) 0.506
9 MMPH (default) 0.501 9 MMPH (default) 0.503
10 SM (3PSM) 0.421 10 SM (3PSM) 0.421
11 RTMA (relaxed) 0.405 11 RTMA (relaxed) 0.415
12 trimfill (default) 0.341 12 trimfill (default) 0.341
13 mean (default) 0.265 13 mean (default) 0.265
14 puniform (star) 0.189 14 puniform (star) 0.189
15 WAAPWLS (default) 0.182 15 WAAPWLS (default) 0.182
16 PETPEESE (default) 0.144 16 PETPEESE (default) 0.144
17 WLS (default) 0.134 17 WLS (default) 0.134
18 PEESE (default) 0.132 18 PEESE (default) 0.132
19 EK (default) 0.116 19 EK (default) 0.116
19 PET (default) 0.116 19 PET (default) 0.116
21 WILS (default) 0.038 21 WILS (default) 0.038
22 puniform (default) 0.000 22 puniform (default) 0.000
23 MAN (default) -0.257 23 MAN (default) -0.267
24 pcurve (default) NaN 24 pcurve (default) NaN

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RMA (default) -5.174 1 RMA (default) -5.174
2 RoBMA (PSMA) -5.000 2 RoBMA (PSMA) -5.000
3 SM (4PSM) -4.819 3 SM (4PSM) -4.819
4 MAIVE (default) -4.778 4 MAIVE (default) -4.778
5 trimfill (default) -4.744 5 trimfill (default) -4.744
6 MAIVE (WAIVE) -4.687 6 MAIVE (WAIVE) -4.687
7 MMPH (default) -4.578 7 AK (AK2) -4.674
8 AK (AK2) -4.522 8 MMPH (default) -4.580
9 AK (AK1) -4.122 9 AK (AK1) -4.122
10 SM (3PSM) -3.982 10 SM (3PSM) -3.982
11 mean (default) -3.862 11 mean (default) -3.862
12 WLS (default) -3.651 12 WLS (default) -3.651
13 puniform (star) -3.640 13 puniform (star) -3.640
14 PEESE (default) -3.558 14 PEESE (default) -3.558
15 RTMA (relaxed) -3.340 15 RTMA (relaxed) -3.440
16 WAAPWLS (default) -3.230 16 WAAPWLS (default) -3.230
17 PETPEESE (default) -3.213 17 PETPEESE (default) -3.213
18 EK (default) -3.094 18 EK (default) -3.094
18 PET (default) -3.094 18 PET (default) -3.094
20 FMA (default) -2.055 20 FMA (default) -2.055
21 WILS (default) -1.760 21 WILS (default) -1.760
22 MAN (default) -1.119 22 MAN (default) -1.102
23 puniform (default) 0.000 23 puniform (default) 0.000
24 pcurve (default) NaN 24 pcurve (default) NaN

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 FMA (default) 0.263 1 FMA (default) 0.263
2 RoBMA (PSMA) 0.273 2 RoBMA (PSMA) 0.273
3 MAIVE (default) 0.294 3 MAIVE (default) 0.294
4 MAIVE (WAIVE) 0.297 4 MAIVE (WAIVE) 0.297
5 RMA (default) 0.361 5 RMA (default) 0.361
6 AK (AK2) 0.501 6 AK (AK2) 0.506
7 SM (4PSM) 0.559 7 SM (4PSM) 0.559
8 MMPH (default) 0.642 8 MMPH (default) 0.641
9 AK (AK1) 0.642 9 AK (AK1) 0.642
10 SM (3PSM) 0.671 10 SM (3PSM) 0.671
11 RTMA (relaxed) 0.683 11 RTMA (relaxed) 0.682
12 trimfill (default) 0.715 12 trimfill (default) 0.715
13 MAN (default) 0.773 13 MAN (default) 0.780
14 mean (default) 0.783 14 mean (default) 0.783
15 WAAPWLS (default) 0.791 15 WAAPWLS (default) 0.791
16 PETPEESE (default) 0.832 16 PETPEESE (default) 0.832
17 puniform (star) 0.844 17 puniform (star) 0.844
18 WLS (default) 0.860 18 WLS (default) 0.860
19 PEESE (default) 0.861 19 PEESE (default) 0.861
20 EK (default) 0.864 20 EK (default) 0.864
20 PET (default) 0.864 20 PET (default) 0.864
22 WILS (default) 0.931 22 WILS (default) 0.931
23 puniform (default) 1.000 23 puniform (default) 1.000
24 pcurve (default) NaN 24 pcurve (default) NaN

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 puniform (default) 1.000 1 puniform (default) 1.000
2 mean (default) 0.995 2 mean (default) 0.995
3 puniform (star) 0.992 3 puniform (star) 0.992
4 AK (AK1) 0.991 4 AK (AK1) 0.991
5 MMPH (default) 0.988 5 MMPH (default) 0.988
6 WLS (default) 0.980 6 WLS (default) 0.980
7 PEESE (default) 0.979 7 PEESE (default) 0.979
8 trimfill (default) 0.977 8 trimfill (default) 0.977
9 EK (default) 0.969 9 EK (default) 0.969
9 PET (default) 0.969 9 PET (default) 0.969
11 WILS (default) 0.968 11 WILS (default) 0.968
12 PETPEESE (default) 0.959 12 PETPEESE (default) 0.959
13 RMA (default) 0.955 13 RMA (default) 0.955
14 SM (3PSM) 0.950 14 SM (3PSM) 0.950
15 WAAPWLS (default) 0.948 15 WAAPWLS (default) 0.948
16 AK (AK2) 0.943 16 AK (AK2) 0.943
17 RTMA (relaxed) 0.935 17 RTMA (relaxed) 0.942
18 SM (4PSM) 0.935 18 SM (4PSM) 0.935
19 RoBMA (PSMA) 0.890 19 RoBMA (PSMA) 0.890
20 MAIVE (default) 0.881 20 MAIVE (default) 0.881
21 MAIVE (WAIVE) 0.880 21 MAIVE (WAIVE) 0.880
22 FMA (default) 0.785 22 FMA (default) 0.785
23 MAN (default) 0.741 23 MAN (default) 0.744
24 pcurve (default) NaN 24 pcurve (default) NaN

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Conditional on Method Convergence)

The results below are conditional on method convergence. Note that the methods might differ in convergence rate and are therefore not compared on the same data sets.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Replacement in Case of Non-Convergence)

The results below incorporate method replacement to handle non-convergence. If a method fails to converge, its results are replaced with the results from a simpler method (e.g., random-effects meta-analysis without publication bias adjustment). This emulates what a data analyst may do in practice in case a method does not converge. However, note that these results do not correspond to “pure” method performance as they might combine multiple different methods. See Method Replacement Strategy for details of the method replacement specification.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

Session Info

This report was compiled on Wed Sep 30 15:38:44 2026 (UTC) using the following computational environment

## R version 4.6.1 (2026-06-24)
## Platform: x86_64-pc-linux-gnu
## Running under: Ubuntu 24.04.5 LTS
## 
## Matrix products: default
## BLAS:   /usr/lib/x86_64-linux-gnu/openblas-pthread/libblas.so.3 
## LAPACK: /usr/lib/x86_64-linux-gnu/openblas-pthread/libopenblasp-r0.3.26.so;  LAPACK version 3.12.0
## 
## locale:
##  [1] LC_CTYPE=C.UTF-8       LC_NUMERIC=C           LC_TIME=C.UTF-8       
##  [4] LC_COLLATE=C.UTF-8     LC_MONETARY=C.UTF-8    LC_MESSAGES=C.UTF-8   
##  [7] LC_PAPER=C.UTF-8       LC_NAME=C              LC_ADDRESS=C          
## [10] LC_TELEPHONE=C         LC_MEASUREMENT=C.UTF-8 LC_IDENTIFICATION=C   
## 
## time zone: UTC
## tzcode source: system (glibc)
## 
## attached base packages:
## [1] stats     graphics  grDevices utils     datasets  methods   base     
## 
## other attached packages:
## [1] scales_1.4.0                   ggdist_3.3.3                  
## [3] ggplot2_4.0.3                  PublicationBiasBenchmark_0.3.0
## 
## loaded via a namespace (and not attached):
##  [1] gtable_0.3.6         xfun_0.61            bslib_0.12.0        
##  [4] htmlwidgets_1.6.4    lattice_0.22-9       vctrs_0.7.3         
##  [7] tools_4.6.1          Rdpack_2.6.6         generics_0.1.4      
## [10] curl_8.0.0           sandwich_3.1-3       tibble_3.3.1        
## [13] pkgconfig_2.0.3      RColorBrewer_1.1-3   S7_0.2.2            
## [16] desc_1.4.3           distributional_0.9.0 lifecycle_1.0.5     
## [19] compiler_4.6.1       farver_2.1.2         stringr_1.6.0       
## [22] textshaping_1.0.5    htmltools_0.5.9      sass_0.4.10         
## [25] clubSandwich_0.7.0   yaml_2.3.12          pillar_1.11.1       
## [28] pkgdown_2.2.1        jquerylib_0.1.4      cachem_1.1.0        
## [31] tidyselect_1.2.1     digest_0.6.39        stringi_1.8.9       
## [34] dplyr_1.2.1          purrr_1.2.2          labeling_0.4.3      
## [37] fastmap_1.2.0        grid_4.6.1           cli_3.6.6           
## [40] magrittr_2.0.5       triebeard_0.4.1      crul_1.6.0          
## [43] osfr_0.2.9           withr_3.0.3          rmarkdown_2.32      
## [46] httr_1.4.9           otel_0.2.0           ragg_1.5.2          
## [49] zoo_1.9-1            kableExtra_1.4.1     memoise_2.0.1       
## [52] evaluate_1.0.5       knitr_1.52           rbibutils_2.4.1     
## [55] viridisLite_0.4.3    rlang_1.3.0          urltools_1.7.3.1    
## [58] Rcpp_1.1.2           glue_1.8.1           httpcode_0.3.0      
## [61] xml2_1.6.0           svglite_2.2.2        rstudioapi_0.19.0   
## [64] jsonlite_2.0.0       R6_2.6.1             systemfonts_1.3.2   
## [67] fs_2.1.0