mirror of
https://github.com/MCCTeam/Minecraft-Console-Client
synced 2026-08-15 13:04:36 +00:00
Compare commits
1466 commits
20220903-3
...
master
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
d50e90d860 | ||
|
|
196dc54222 | ||
|
|
a0710fd82f | ||
|
|
cfb77088b6 | ||
|
|
f1d844afbe | ||
|
|
1fd4b0244d | ||
|
|
844fa7af2e | ||
|
|
9523275eb6 | ||
|
|
8e11304bf9 | ||
|
|
34e9fd46fd | ||
|
|
15dd55d10a | ||
|
|
8f8e03b1c9 | ||
|
|
c9225aac58 | ||
|
|
7b100a3149 | ||
|
|
dbce402842 | ||
|
|
503652a760 | ||
|
|
3ae1d746fa | ||
|
|
4a68205f02 | ||
|
|
c377e52306 | ||
|
|
7fee87c98f | ||
|
|
07aa032214 | ||
|
|
6b0f1a8f77 | ||
|
|
a780312c19 | ||
|
|
19f02dfc25 | ||
|
|
dc472c3df6 | ||
|
|
90fda17365 | ||
|
|
5655e3fc89 | ||
|
|
456a548cbc | ||
|
|
92212d2b95 | ||
|
|
eea631f69e | ||
|
|
6e4bca5073 | ||
|
|
630a5fabef | ||
|
|
47bc47786a | ||
|
|
9b90b535d8 | ||
|
|
1686d00f30 | ||
|
|
23b83a45f5 | ||
|
|
97631acc55 | ||
|
|
03beb51fb7 | ||
|
|
9d1c7ac4bd | ||
|
|
668cbf7e6a | ||
|
|
bf99df7cde | ||
|
|
4c6b2b3190 | ||
|
|
fd0c99b2d3 | ||
|
|
ba134476af | ||
|
|
1dc1d06f51 | ||
|
|
78da20b4e4 | ||
|
|
a951b5b6c6 | ||
|
|
4a2837d862 | ||
|
|
85672cc0a2 | ||
|
|
b1200af588 | ||
|
|
64c1dc55e0 | ||
|
|
c19fdd6634 | ||
|
|
fceed9b4d7 | ||
|
|
25bbe35718 | ||
|
|
419c140dc4 | ||
|
|
720183955d | ||
|
|
142fb107db | ||
|
|
99f5744715 | ||
|
|
d1df44aae5 | ||
|
|
e4161af158 | ||
|
|
22685a720c | ||
|
|
f34b8ba989 | ||
|
|
95031338bf | ||
|
|
d3c87fcd36 | ||
|
|
3c452dfcc5 | ||
|
|
aee35c3f58 | ||
|
|
c23a71c489 | ||
|
|
800f757804 | ||
|
|
e0ba34f7fe | ||
|
|
9edf53250a | ||
|
|
3501ea142f | ||
|
|
e8afeea3d6 | ||
|
|
aa1b6f968e | ||
|
|
25a2f5a4a3 | ||
|
|
85fe6b451b | ||
|
|
72200fc1ef | ||
|
|
c05c1f3c02 | ||
|
|
73377685b4 | ||
|
|
e6b461e7b2 | ||
|
|
327d3a1c37 | ||
|
|
ef1c8fe740 | ||
|
|
31876d485a | ||
|
|
3f71f7afef | ||
|
|
aaa4908f7d | ||
|
|
7f2a71f9fe | ||
|
|
c0d983272b | ||
|
|
2aa85cbb17 | ||
|
|
bd29c7fc5c | ||
|
|
35db711ea4 | ||
|
|
3d988c932f | ||
|
|
a1b542dad6 | ||
|
|
7d7b880a7c | ||
|
|
37cf19c175 | ||
|
|
5e5fa539bb | ||
|
|
981a996344 | ||
|
|
98dfde6cc2 | ||
|
|
28e845d1ca | ||
|
|
de96462659 | ||
|
|
02b34a8d13 | ||
|
|
20756b351e | ||
|
|
16637f7eba | ||
|
|
d38afe8e1e | ||
|
|
883ed12ccd | ||
|
|
98d4781d2c | ||
|
|
e744ff206b | ||
|
|
61c2cbf7c3 | ||
|
|
2e383403c3 | ||
|
|
80c07634e2 | ||
|
|
eeb89d676b | ||
|
|
9980e08bac | ||
|
|
861084db9b | ||
|
|
71f935eaeb | ||
|
|
a5423211ad | ||
|
|
fc74cc066c | ||
|
|
5b13e740ef | ||
|
|
b1d02edd99 | ||
|
|
6bfcd9e9ca | ||
|
|
3c59bfe13a | ||
|
|
654d9d4dd7 | ||
|
|
0a514d1351 | ||
|
|
913c0f0f72 | ||
|
|
4bbc751cc1 | ||
|
|
5f0ca8837f | ||
|
|
f25fe53881 | ||
|
|
82555bd823 | ||
|
|
29b0087a3a | ||
|
|
b10bd81bee | ||
|
|
5545aabb60 | ||
|
|
c98c2a8f83 | ||
|
|
b698f122cf | ||
|
|
5228670131 | ||
|
|
c4df73ab6a | ||
|
|
8757533fb9 | ||
|
|
2b05b7420e | ||
|
|
701441d424 | ||
|
|
546345e816 | ||
|
|
d835da76d2 | ||
|
|
6db8c817e0 | ||
|
|
c6ca3dc13d | ||
|
|
a4ff7ae438 | ||
|
|
45734ba294 | ||
|
|
6670685f51 | ||
|
|
de7af06d11 | ||
|
|
59e98867b3 | ||
|
|
a7dd9fff09 | ||
|
|
1796e7f80c | ||
|
|
e331eae03c | ||
|
|
4cc5fc00c3 | ||
|
|
03e2e0d3fd | ||
|
|
47de24cf1d | ||
|
|
04a975d011 | ||
|
|
2d31199f97 | ||
|
|
f39fe2d067 | ||
|
|
8e26e55650 | ||
|
|
75f8f29d1a | ||
|
|
b74364d465 | ||
|
|
6456149d39 | ||
|
|
858f88d9ec | ||
|
|
a2dabe1aa8 | ||
|
|
67bd944179 | ||
|
|
b13511f201 | ||
|
|
f91319cdfe | ||
|
|
44e3685454 | ||
|
|
5c1361169c | ||
|
|
b282476607 | ||
|
|
ea66dd4c13 | ||
|
|
3f64a9158c | ||
|
|
6331478e30 | ||
|
|
d55fd75520 | ||
|
|
6673b398b1 | ||
|
|
93d9939378 | ||
|
|
a57fb7f61f | ||
|
|
ba04636d06 | ||
|
|
21a52b3bcc | ||
|
|
7ce4f82871 | ||
|
|
67066a145d | ||
|
|
d256517e21 | ||
|
|
90a3777dac | ||
|
|
effc05aec3 | ||
|
|
b543825200 | ||
|
|
b87b5abbe2 | ||
|
|
9f974da00d | ||
|
|
34f45543c3 | ||
|
|
9c47456455 | ||
|
|
92ab2f7378 | ||
|
|
e515921df2 | ||
|
|
1a1caec30a | ||
|
|
7bf342baab | ||
|
|
9c994de2ee | ||
|
|
03d3d6f950 | ||
|
|
98a1551f6a | ||
|
|
acabe95801 | ||
|
|
12ab759f3b | ||
|
|
41b33b6599 | ||
|
|
c282f2783c | ||
|
|
72001554a8 | ||
|
|
8e31a37ea9 | ||
|
|
c94be261fd | ||
|
|
b5af631a40 | ||
|
|
05741da9f1 | ||
|
|
5ad0d55705 | ||
|
|
5279609561 | ||
|
|
9549743c6d | ||
|
|
8ba95c6140 | ||
|
|
0852d54991 | ||
|
|
2bccb7f8e9 | ||
|
|
e564da3e09 | ||
|
|
4dbce45676 | ||
|
|
5ec1bcd9db | ||
|
|
c1635a999f | ||
|
|
4a28430ef5 | ||
|
|
c7053ec015 | ||
|
|
abffb58e5f | ||
|
|
e511d33452 | ||
|
|
4bd0929ebf | ||
|
|
443f4874df | ||
|
|
b445d2bdec | ||
|
|
58ad896021 | ||
|
|
8dccdaa7aa | ||
|
|
5890cdef50 | ||
|
|
4013d39068 | ||
|
|
ab29180e39 | ||
|
|
248a7a2f6f | ||
|
|
8e0db1402f | ||
|
|
ca5c9f73e3 | ||
|
|
af285866dc | ||
|
|
b665ebfeac | ||
|
|
e4d71d970b | ||
|
|
0a041b37c5 | ||
|
|
e0752d8d98 | ||
|
|
0519eac4a8 | ||
|
|
bf292ea1be | ||
|
|
afee31e918 | ||
|
|
e03b7de959 | ||
|
|
1be001bdf8 | ||
|
|
4c6d5b566b | ||
|
|
d9c6826443 | ||
|
|
fdb6e5db4c | ||
|
|
3b65085a66 | ||
|
|
6b986dcddd | ||
|
|
158a4b1b13 | ||
|
|
f3e8f65fe5 | ||
|
|
8c8962f1ea | ||
|
|
fbacc550e2 | ||
|
|
d96396f74d | ||
|
|
e0b21dd1d2 | ||
|
|
229c9eba0a | ||
|
|
2ac7e84b37 | ||
|
|
5aa1816621 | ||
|
|
b10e332e31 | ||
|
|
906d66a713 | ||
|
|
b69708c637 | ||
|
|
5363ef7bc1 | ||
|
|
740de13d98 | ||
|
|
edadd91016 | ||
|
|
ff00d01fd7 | ||
|
|
64c42bb4b5 | ||
|
|
698a562156 | ||
|
|
7ec9d64b5b | ||
|
|
756bc1e11f | ||
|
|
5a5165e89f | ||
|
|
0e15806722 | ||
|
|
37271c5988 | ||
|
|
35fd34b19d | ||
|
|
ad065225ba | ||
|
|
0445d95c23 | ||
|
|
37bda82828 | ||
|
|
5dde778b53 | ||
|
|
47f13a815c | ||
|
|
f432637710 | ||
|
|
e81700bf0e | ||
|
|
b6bfecc622 | ||
|
|
44e7f205c8 | ||
|
|
3073137006 | ||
|
|
5c323772b3 | ||
|
|
0f32771075 | ||
|
|
e6ea29c914 | ||
|
|
e1f3dd5fdb | ||
|
|
7fd70ccf38 | ||
|
|
914c56fc6e | ||
|
|
7f3b2c26da | ||
|
|
49b7b98ba8 | ||
|
|
7405aad68f | ||
|
|
b083336bf6 | ||
|
|
a40188cb09 | ||
|
|
c7166f1cbc | ||
|
|
c19ce62e98 | ||
|
|
98ad174999 | ||
|
|
296070f7e4 | ||
|
|
0dde017ecc | ||
|
|
871bc7651f | ||
|
|
a6c832b744 | ||
|
|
33144b16d4 | ||
|
|
81be654601 | ||
|
|
0cc34e6e22 | ||
|
|
32d5ad1108 | ||
|
|
16e3a69211 | ||
|
|
6e86edc632 | ||
|
|
76c3403700 | ||
|
|
c8286f60c1 | ||
|
|
23262de6c2 | ||
|
|
a12ba29c45 | ||
|
|
1ca023be36 | ||
|
|
0fda37b017 | ||
|
|
75df2e0c71 | ||
|
|
d6b3dec8ad | ||
|
|
4b86e93764 | ||
|
|
8906f69839 | ||
|
|
f4c160979c | ||
|
|
691057fe13 | ||
|
|
dfef24ad10 | ||
|
|
a16af89dac | ||
|
|
744dcd73cd | ||
|
|
bca4532328 | ||
|
|
3557239c43 | ||
|
|
22e987070a | ||
|
|
e987b525be | ||
|
|
512445cb1d | ||
|
|
2003786608 | ||
|
|
2441c67178 | ||
|
|
a88ca3cb71 | ||
|
|
fc05d2c98c | ||
|
|
ee6eb84bd8 | ||
|
|
82ebdd00ee | ||
|
|
737a94475a | ||
|
|
1a86655dfb | ||
|
|
65dec59687 | ||
|
|
5f9549586a | ||
|
|
4c54d2bbb7 | ||
|
|
fdbffcc5bb | ||
|
|
cf65c6f241 | ||
|
|
ee48c24f69 | ||
|
|
c9b0913c1a | ||
|
|
c3c57c058a | ||
|
|
968800b95a | ||
|
|
d5358d3d1a | ||
|
|
5de72de234 | ||
|
|
d158bfdf21 | ||
|
|
b9b7160e19 | ||
|
|
35dd3c4f06 | ||
|
|
eae96a8fbc | ||
|
|
be20b97478 | ||
|
|
1e35a4c624 | ||
|
|
354a498f21 | ||
|
|
435887cc04 | ||
|
|
f7bc817408 | ||
|
|
eaf4704473 | ||
|
|
76e5cab248 | ||
|
|
6b5435629f | ||
|
|
62740ee94e | ||
|
|
0881cbaa1c | ||
|
|
22455905a7 | ||
|
|
962c8b1ab2 | ||
|
|
c55d32bb70 | ||
|
|
c95fe131e3 | ||
|
|
0f3289dfdf | ||
|
|
5705df43bd | ||
|
|
7b3e5ee492 | ||
|
|
65ef3dde6b | ||
|
|
6dc42d9bd1 | ||
|
|
1165f32bb7 | ||
|
|
534e337f10 | ||
|
|
031dbb16d4 | ||
|
|
ec53f53f42 | ||
|
|
f10556a162 | ||
|
|
46fdd687c7 | ||
|
|
53afc252ea | ||
|
|
8c4acb0ad5 | ||
|
|
d08cf803b8 | ||
|
|
301ea6b9db | ||
|
|
ef106c80c2 | ||
|
|
7b415d5388 | ||
|
|
d97861888d | ||
|
|
90ea05b17d | ||
|
|
b05c8cfe0d | ||
|
|
b25579a105 | ||
|
|
dfc1164839 | ||
|
|
893be203e5 | ||
|
|
d5308ba8c6 | ||
|
|
3a28634592 | ||
|
|
871305bd72 | ||
|
|
2f03f66f5c | ||
|
|
d05e3148f1 | ||
|
|
c6b9121a59 | ||
|
|
d427a6e160 | ||
|
|
0566f2518b | ||
|
|
a1516e9680 | ||
|
|
6631180f8a | ||
|
|
c2475ea9e5 | ||
|
|
157d90273b | ||
|
|
17d4b78022 | ||
|
|
b757215dbb | ||
|
|
83f14a64f8 | ||
|
|
0d34d5e33d | ||
|
|
9cada9d19d | ||
|
|
cf382122e9 | ||
|
|
f0fda8ce9f | ||
|
|
2d677e0f0d | ||
|
|
7f9023e7bb | ||
|
|
79b091cab6 | ||
|
|
5270aed315 | ||
|
|
0194380fbc | ||
|
|
ef133f3d6d | ||
|
|
22f46d14a9 | ||
|
|
448e1ef6f5 | ||
|
|
19a60fd85c | ||
|
|
ef8e41196b | ||
|
|
da330a158e | ||
|
|
c42133797e | ||
|
|
ce7bf0f400 | ||
|
|
0a9948a92a | ||
|
|
0c9ff137b2 | ||
|
|
57cf5c75c0 | ||
|
|
df05eb6dc4 | ||
|
|
f6446b198d | ||
|
|
4b9e0e93aa | ||
|
|
4993498d52 | ||
|
|
20f536188a | ||
|
|
a7a3566756 | ||
|
|
cac9e5c625 | ||
|
|
7a75328f0a | ||
|
|
747662ea8c | ||
|
|
ac1d5c37a6 | ||
|
|
fe7ab9f373 | ||
|
|
b33ee7e4a1 | ||
|
|
ca5e7520cb | ||
|
|
0ceaaead10 | ||
|
|
129f5bda9f | ||
|
|
59332be062 | ||
|
|
5df03b0abd | ||
|
|
a2f7511ef7 | ||
|
|
7da47d36dc | ||
|
|
05afeaf5bc | ||
|
|
966b787a4c | ||
|
|
6beb00bae4 | ||
|
|
e5b81fe428 | ||
|
|
dcfff3f1ba | ||
|
|
8456e363f5 | ||
|
|
d9f79bf512 | ||
|
|
a79373c713 | ||
|
|
71c01c739a | ||
|
|
d4faea9ecc | ||
|
|
a17b7d4bd3 | ||
|
|
a4ade66b8e | ||
|
|
69070f0157 | ||
|
|
d61a44ba34 | ||
|
|
d3c3e5a554 | ||
|
|
2c993247d2 | ||
|
|
00c3a8cb51 | ||
|
|
3931d15c7d | ||
|
|
cb4eb15784 | ||
|
|
a242f97c14 | ||
|
|
bd55fc3ed3 | ||
|
|
5fecd667a3 | ||
|
|
a384d29b64 | ||
|
|
eeb4383099 | ||
|
|
ed0a69185f | ||
|
|
d5a9ae3deb | ||
|
|
6d6d1105eb | ||
|
|
c8c9f9ada5 | ||
|
|
7fca2acd24 | ||
|
|
ba45f71f6e | ||
|
|
154778a488 | ||
|
|
41a1a1db9b | ||
|
|
d40a3d9a0b | ||
|
|
f405548529 | ||
|
|
87593b17b9 | ||
|
|
fb41b1ffe1 | ||
|
|
52eda4ccdd | ||
|
|
2f40cc1538 | ||
|
|
a1ab366787 | ||
|
|
5c9319c10d | ||
|
|
33d050583a | ||
|
|
fec46ccd76 | ||
|
|
7e9be430db | ||
|
|
c85dd4b498 | ||
|
|
c43b157b55 | ||
|
|
17808bb652 | ||
|
|
6c70685d52 | ||
|
|
6f30f9b79a | ||
|
|
699b2cad6b | ||
|
|
3fb7625b39 | ||
|
|
8ab9f8505f | ||
|
|
7f8ee0e2d6 | ||
|
|
461be9f368 | ||
|
|
c9f1f88e62 | ||
|
|
7b534ca547 | ||
|
|
0168cf8aa1 | ||
|
|
802cc696d6 | ||
|
|
4c1b2197d7 | ||
|
|
16a3ad2b69 | ||
|
|
a2ff094369 | ||
|
|
79008cb0e5 | ||
|
|
be084f6bea | ||
|
|
3745834f33 | ||
|
|
4d851b6339 | ||
|
|
db5faa22b1 | ||
|
|
f15c000868 | ||
|
|
dacf09b553 | ||
|
|
9eba4d026c | ||
|
|
da3b2de2ce | ||
|
|
54e07cd233 | ||
|
|
687d1746ed | ||
|
|
34bc6c5400 | ||
|
|
84858a55ee | ||
|
|
a98b15fb3c | ||
|
|
54d7b81e7b | ||
|
|
bce1cd9290 | ||
|
|
d6aa9d101f | ||
|
|
7aee32d079 | ||
|
|
8f99f1b2be | ||
|
|
f64757edea | ||
|
|
47e943562e | ||
|
|
bd6aae0060 | ||
|
|
0a409ad394 | ||
|
|
4808b61b9a | ||
|
|
eca27be1fa | ||
|
|
aeec731ae5 | ||
|
|
3a5283285a | ||
|
|
654a16907b | ||
|
|
81b756292e | ||
|
|
640a4e39b7 | ||
|
|
94bf42710a | ||
|
|
c7bc25aa17 | ||
|
|
446c6e7739 | ||
|
|
2ae31e269f | ||
|
|
c5df6a49c6 | ||
|
|
902e944cbb | ||
|
|
cf34db526b | ||
|
|
74def7c512 | ||
|
|
e09b997cdf | ||
|
|
99e73b1df4 | ||
|
|
23ac2f9639 | ||
|
|
d4c6f4d708 | ||
|
|
102a175881 | ||
|
|
a9930b0a67 | ||
|
|
94c6425a2b | ||
|
|
5147fb42df | ||
|
|
daeec49789 | ||
|
|
8f4e6b08b3 | ||
|
|
b2eeb80659 | ||
|
|
445c236914 | ||
|
|
932d01fb30 | ||
|
|
9efb7ef767 | ||
|
|
484133f07a | ||
|
|
a18b5d2415 | ||
|
|
a0a69caeea | ||
|
|
a16152410e | ||
|
|
7e2e2f3f53 | ||
|
|
a0171af007 | ||
|
|
5f7e302213 | ||
|
|
2d8412ce29 | ||
|
|
1d4ff226e0 | ||
|
|
8bedcece52 | ||
|
|
19a4f8132a | ||
|
|
a5896d47b5 | ||
|
|
15a50dc844 | ||
|
|
9db4ad86b4 | ||
|
|
221b82643d | ||
|
|
6c1449439c | ||
|
|
5a9786cb0b | ||
|
|
02686de79c | ||
|
|
8743129e72 | ||
|
|
a2d032b549 | ||
|
|
5d21b4575f | ||
|
|
0a463f2380 | ||
|
|
10593749dd | ||
|
|
c82b6c844b | ||
|
|
3ee0177d62 | ||
|
|
ff9aa62702 | ||
|
|
72ec3122ae | ||
|
|
6c9dcfd216 | ||
|
|
d44fc22c8a | ||
|
|
02873957a8 | ||
|
|
07189509b0 | ||
|
|
af405e5632 | ||
|
|
5c39d3b83d | ||
|
|
5e99e3a42d | ||
|
|
757cbe1a68 | ||
|
|
d2c1cbf2a5 | ||
|
|
64ccdcb39b | ||
|
|
37fe626325 | ||
|
|
15aabd9423 | ||
|
|
f31ed82ce1 | ||
|
|
611951668e | ||
|
|
7893bd7fe4 | ||
|
|
d20b139215 | ||
|
|
dc50df3b94 | ||
|
|
9657ed91d2 | ||
|
|
95fc0fd0f8 | ||
|
|
4ac3601aa7 | ||
|
|
454ce331b2 | ||
|
|
1e2847a48e | ||
|
|
4d51282040 | ||
|
|
7162691b57 | ||
|
|
4fb4c018e4 | ||
|
|
bdb561ed4a | ||
|
|
5c2c176ba1 | ||
|
|
535b95f4b9 | ||
|
|
d95b6bb9eb | ||
|
|
aa59c4d75a | ||
|
|
e4aeb51f71 | ||
|
|
9613e0df51 | ||
|
|
cffddb398e | ||
|
|
d1837d104b | ||
|
|
d4014f7c87 | ||
|
|
e3d3a86ac2 | ||
|
|
052993a2c1 | ||
|
|
33fccf85b5 | ||
|
|
d25f2f40ec | ||
|
|
397ab07d1f | ||
|
|
3afbae4f89 | ||
|
|
38e86eb209 | ||
|
|
e3a2593781 | ||
|
|
a3c918e946 | ||
|
|
cca4134e8b | ||
|
|
9e3bfaa868 | ||
|
|
d3151925fe | ||
|
|
e2746ae0d1 | ||
|
|
c5b22eb1b2 | ||
|
|
b4e69da81f | ||
|
|
9f4abfdd39 | ||
|
|
c50b9a5b63 | ||
|
|
bd6fb7332f | ||
|
|
6833e60cb5 | ||
|
|
c38162584d | ||
|
|
abe6b9f01b | ||
|
|
1ac9df86a4 | ||
|
|
5b9cc4aa0e | ||
|
|
fb896c45a7 | ||
|
|
28834e7a86 | ||
|
|
5488140261 | ||
|
|
3f431f1104 | ||
|
|
78dc74a6f0 | ||
|
|
bf04141167 | ||
|
|
c0c4c078c0 | ||
|
|
cf5c4f00cb | ||
|
|
bc60e99a63 | ||
|
|
6c36dc341a | ||
|
|
50b4b3c8fe | ||
|
|
a1c1dbc182 | ||
|
|
918f1a560d | ||
|
|
c5537c0444 | ||
|
|
34671fdab2 | ||
|
|
60c2b9ea7e | ||
|
|
aebfcd93e4 | ||
|
|
6d3c58b29f | ||
|
|
88e9c671da | ||
|
|
67ad5fcf97 | ||
|
|
2af0409d00 | ||
|
|
7e26542055 | ||
|
|
5e8d715358 | ||
|
|
067395ab2a | ||
|
|
1b0e27fde7 | ||
|
|
96a7e4a659 | ||
|
|
26bee79a5d | ||
|
|
4efe75174e | ||
|
|
779de6c996 | ||
|
|
668f27a221 | ||
|
|
ec094399df | ||
|
|
a277cde85a | ||
|
|
e2f6494933 | ||
|
|
23a1629113 | ||
|
|
066a1d1238 | ||
|
|
5d00fa6ea6 | ||
|
|
26e8a2f2e6 | ||
|
|
9af6c643e3 | ||
|
|
8df39ba74a | ||
|
|
9c1502600a | ||
|
|
c31739ddd1 | ||
|
|
067bd8faff | ||
|
|
f73148e5ec | ||
|
|
5dec0ac471 | ||
|
|
3acfc0aaa4 | ||
|
|
e5de803613 | ||
|
|
a92bf5b071 | ||
|
|
57a0dedb33 | ||
|
|
df833ae2a8 | ||
|
|
896263acc8 | ||
|
|
7609832976 | ||
|
|
e1cb18c6f8 | ||
|
|
ee02974abe | ||
|
|
c23c229eb2 | ||
|
|
2fb09342a3 | ||
|
|
cc10f4effa | ||
|
|
5e416d6b72 | ||
|
|
c689343371 | ||
|
|
4944497f5a | ||
|
|
a36ba23ba6 | ||
|
|
79a0dff8cd | ||
|
|
1e2b853b14 | ||
|
|
56f2426c1f | ||
|
|
b692b13bbc | ||
|
|
99ac3d028a | ||
|
|
967f67190c | ||
|
|
8eac21b4a4 | ||
|
|
bb18399523 | ||
|
|
41a701b6b2 | ||
|
|
a7a95d991c | ||
|
|
ff1c570a78 | ||
|
|
e2ca8e082b | ||
|
|
df79e10f26 | ||
|
|
494be0930b | ||
|
|
fdd77b562e | ||
|
|
57483646b1 | ||
|
|
f785f509f2 | ||
|
|
8b20973b02 | ||
|
|
ef7aa04e47 | ||
|
|
c0c176a226 | ||
|
|
831a86cb4f | ||
|
|
19781985f7 | ||
|
|
1cec3f7397 | ||
|
|
8726a73b5f | ||
|
|
2409de2a2f | ||
|
|
304c8f04f2 | ||
|
|
a5ab30f3da | ||
|
|
a1acd559d6 | ||
|
|
7152072598 | ||
|
|
c38f92a9ec | ||
|
|
17d43958e1 | ||
|
|
48c69ac5a6 | ||
|
|
98a3b70e2c | ||
|
|
814cd69380 | ||
|
|
f83e7c5707 | ||
|
|
d0c9695a79 | ||
|
|
7bd213a154 | ||
|
|
43c6620475 | ||
|
|
0da4a718cb | ||
|
|
4dea688ca2 | ||
|
|
49319fe781 | ||
|
|
76e873ed54 | ||
|
|
27e66433cd | ||
|
|
63b027d84a | ||
|
|
f54c5aab44 | ||
|
|
c5dc517c43 | ||
|
|
c69cdaddbf | ||
|
|
8bdb20a22c | ||
|
|
efe23eb1f9 | ||
|
|
ca966a464c | ||
|
|
c50b360eae | ||
|
|
2f9cf7bc8c | ||
|
|
729381fb87 | ||
|
|
95e7f08319 | ||
|
|
5a6fd577e5 | ||
|
|
58a5260b5b | ||
|
|
67e36a92d2 | ||
|
|
08551097c6 | ||
|
|
8756ff5b3c | ||
|
|
4919db8820 | ||
|
|
7cd8e3500c | ||
|
|
08c5c15557 | ||
|
|
8270a2d9a3 | ||
|
|
fc2373b6d5 | ||
|
|
8037794601 | ||
|
|
4146897e3c | ||
|
|
d0caf4c9ee | ||
|
|
d6d0800215 | ||
|
|
403284cc53 | ||
|
|
2fe376e7ca | ||
|
|
3ccdbc2a4f | ||
|
|
5f3923973f | ||
|
|
bf54def51f | ||
|
|
552f563eb0 | ||
|
|
f5c7a2801f | ||
|
|
dc71332dd3 | ||
|
|
706a41ee26 | ||
|
|
5044ec965b | ||
|
|
60ba4bd6c2 | ||
|
|
2ce0311949 | ||
|
|
691f1a136e | ||
|
|
c9c16818a4 | ||
|
|
221d5948e2 | ||
|
|
a19a91e37f | ||
|
|
6891d446a5 | ||
|
|
ac9a70c159 | ||
|
|
4bb25c377e | ||
|
|
79910b50f7 | ||
|
|
a72a8cf833 | ||
|
|
c78245c056 | ||
|
|
873bd79fd6 | ||
|
|
5f4227ad11 | ||
|
|
3b213296ac | ||
|
|
e2b6dc27c8 | ||
|
|
86fb4feffc | ||
|
|
df9443381b | ||
|
|
91ef890bb6 | ||
|
|
fde50c1728 | ||
|
|
438311787d | ||
|
|
744de0dbd4 | ||
|
|
bc0781cee9 | ||
|
|
620e8bf274 | ||
|
|
1db0792d7b | ||
|
|
a749cd7fc4 | ||
|
|
ecc88fac06 | ||
|
|
798a24e8be | ||
|
|
3522a16b0d | ||
|
|
13de67b6f8 | ||
|
|
8e1822b0d2 | ||
|
|
8be66daab1 | ||
|
|
092854532a | ||
|
|
6949276779 | ||
|
|
970ba19172 | ||
|
|
f749840d89 | ||
|
|
576575ff65 | ||
|
|
e569ffe0cc | ||
|
|
b2ef5cb23b | ||
|
|
7edf458135 | ||
|
|
3fab7eb78f | ||
|
|
9a147b57e5 | ||
|
|
d7e898c6d4 | ||
|
|
317d8cca49 | ||
|
|
35cfd4a7db | ||
|
|
6d016332fb | ||
|
|
4546e6946e | ||
|
|
a9f1ad4433 | ||
|
|
975aab88e3 | ||
|
|
790e0bfe55 | ||
|
|
1e60b611e9 | ||
|
|
350c1cdd51 | ||
|
|
1479c646d1 | ||
|
|
4f89e4fe36 | ||
|
|
db1fade2c2 | ||
|
|
ad684fb5b6 | ||
|
|
e13ba93f47 | ||
|
|
7aabe8ba28 | ||
|
|
f325dd7475 | ||
|
|
3bfe5aa855 | ||
|
|
88bca839f7 | ||
|
|
6faddae16e | ||
|
|
ae5e016f5f | ||
|
|
a8643d85fc | ||
|
|
8eee50044f | ||
|
|
79fa297e2b | ||
|
|
2c8b15b02e | ||
|
|
e7d519e4aa | ||
|
|
22cf7a046b | ||
|
|
725510d3ef | ||
|
|
9ae3bc4d0d | ||
|
|
644014e42f | ||
|
|
269a890a23 | ||
|
|
734de2a9ac | ||
|
|
f6797cb4b5 | ||
|
|
ba2402ee11 | ||
|
|
480f0d85f0 | ||
|
|
f2e1c57b23 | ||
|
|
e19de8eb0b | ||
|
|
ceff78a821 | ||
|
|
1c17da2665 | ||
|
|
6714d9a9d0 | ||
|
|
a41b6719ce | ||
|
|
2fb5c163d5 | ||
|
|
0ad892ef50 | ||
|
|
eb7bfff8ab | ||
|
|
549f39fab1 | ||
|
|
c6da4e2ac3 | ||
|
|
a08bfca4e5 | ||
|
|
f07b1e964c | ||
|
|
8dfcf9c5d5 | ||
|
|
f50bfbb857 | ||
|
|
782481816d | ||
|
|
cf6db27088 | ||
|
|
21e2f41f25 | ||
|
|
d49a597671 | ||
|
|
1f4db70f8a | ||
|
|
fd1009b43f | ||
|
|
0a149647b6 | ||
|
|
04c1f941b6 | ||
|
|
93112d2c02 | ||
|
|
4fc1aacca5 | ||
|
|
49dec51588 | ||
|
|
a97096cddf | ||
|
|
4f957cee7e | ||
|
|
3c97193b70 | ||
|
|
eb8ccc43d7 | ||
|
|
4626ccbc67 | ||
|
|
b8af534438 | ||
|
|
1aea8d3a4e | ||
|
|
c3fa413b4e | ||
|
|
542ff78ddf | ||
|
|
911908bfaf | ||
|
|
e1b018c333 | ||
|
|
67662c5df7 | ||
|
|
968f864f34 | ||
|
|
37bcad37e0 | ||
|
|
bdad4f302d | ||
|
|
a8200b6e14 | ||
|
|
ac65482296 | ||
|
|
c5a0409edc | ||
|
|
fe5f07306d | ||
|
|
8891b65eb8 | ||
|
|
dbba63342d | ||
|
|
469e6667bf | ||
|
|
272900d52e | ||
|
|
ac1d2b7142 | ||
|
|
497a1174de | ||
|
|
85210464a5 | ||
|
|
7c7b58e941 | ||
|
|
3f3f614640 | ||
|
|
b4829eaca4 | ||
|
|
8f337cc0f3 | ||
|
|
0beaae13f7 | ||
|
|
431ed0466d | ||
|
|
89d918a0c9 | ||
|
|
b631fcb487 | ||
|
|
ae7ce35cc8 | ||
|
|
b21f40593e | ||
|
|
d417335325 | ||
|
|
353771e307 | ||
|
|
a4a058aab2 | ||
|
|
a113dc8430 | ||
|
|
2f30528e4f | ||
|
|
025093a11b | ||
|
|
5de84d7e59 | ||
|
|
f77d58402a | ||
|
|
a357d2c87a | ||
|
|
081cebcf76 | ||
|
|
051ee70805 | ||
|
|
fce12db33f | ||
|
|
95f6c5768d | ||
|
|
1efa55206f | ||
|
|
9855e2e0f1 | ||
|
|
48b9e8b4fb | ||
|
|
61357d23c7 | ||
|
|
bee1efa1e0 | ||
|
|
e1313dad4e | ||
|
|
9e267113a5 | ||
|
|
9c78f7958f | ||
|
|
df24f28c97 | ||
|
|
8f1e5cd48d | ||
|
|
1bba41c395 | ||
|
|
cfc2baa4e7 | ||
|
|
3a0b32ec68 | ||
|
|
413cb1860d | ||
|
|
a4fb677566 | ||
|
|
64eb48f46d | ||
|
|
852be6e90d | ||
|
|
97af063d79 | ||
|
|
c7597e8822 | ||
|
|
f72355d974 | ||
|
|
adc6fc0029 | ||
|
|
4ff7712f20 | ||
|
|
599c6aac09 | ||
|
|
7107023d37 | ||
|
|
f94c1ed60d | ||
|
|
df6eeb9b4a | ||
|
|
3570ef605e | ||
|
|
42abf928fb | ||
|
|
949c1f6b67 | ||
|
|
09b3ec1a81 | ||
|
|
0ff64910f7 | ||
|
|
03b06ea3ac | ||
|
|
d90635f40b | ||
|
|
f855839bb3 | ||
|
|
78f5d246f9 | ||
|
|
0a57c927d7 | ||
|
|
beabe14c92 | ||
|
|
f215921e60 | ||
|
|
db6422d154 | ||
|
|
bbb3b1e38a | ||
|
|
e8b3e6b52e | ||
|
|
2a82c35b36 | ||
|
|
3f9fd47e5a | ||
|
|
978dc4b896 | ||
|
|
cd39c1ec12 | ||
|
|
5d4977af6f | ||
|
|
1901f4b62d | ||
|
|
74d29321b4 | ||
|
|
0b98628572 | ||
|
|
4ec3d0c13a | ||
|
|
21cd24e056 | ||
|
|
2f1da9e8c9 | ||
|
|
0f77828ac5 | ||
|
|
f6850920a3 | ||
|
|
033c615a13 | ||
|
|
56bfd21ee2 | ||
|
|
f467f4d6e4 | ||
|
|
930fecde23 | ||
|
|
7c95acfd0e | ||
|
|
b3d7942aca | ||
|
|
29be211946 | ||
|
|
8da8f6044f | ||
|
|
4c33c5fc27 | ||
|
|
b65b3afc3c | ||
|
|
750295b1e3 | ||
|
|
c36dec435d | ||
|
|
f4ad24746c | ||
|
|
055def372b | ||
|
|
1a22002bde | ||
|
|
8b27386a5b | ||
|
|
e0361d2183 | ||
|
|
425cff529c | ||
|
|
2e55a6bc85 | ||
|
|
046cb15c75 | ||
|
|
5d4ea515a7 | ||
|
|
e44ade8688 | ||
|
|
0aa117d023 | ||
|
|
057d6f515a | ||
|
|
91791e99c7 | ||
|
|
4f608687ee | ||
|
|
3735cab9dd | ||
|
|
8cdd32a52a | ||
|
|
db6bc3d2e6 | ||
|
|
d25480b88c | ||
|
|
70d1f70cca | ||
|
|
ee90be3005 | ||
|
|
b79dd1d379 | ||
|
|
cc92cd66d4 | ||
|
|
cb57db8328 | ||
|
|
e52fe6cb8b | ||
|
|
1f54a7c247 | ||
|
|
c2edd6f619 | ||
|
|
92a911ce99 | ||
|
|
730cdf92e7 | ||
|
|
c9c4d8dd62 | ||
|
|
d65f69fad4 | ||
|
|
7011927c71 | ||
|
|
30e95f2d23 | ||
|
|
950d9bcfdc | ||
|
|
50dd5a3ba3 | ||
|
|
338f534239 | ||
|
|
8db0467f69 | ||
|
|
957054eb12 | ||
|
|
b4d7d64cdd | ||
|
|
1298654693 | ||
|
|
ced65122f1 | ||
|
|
d4b3c42d8c | ||
|
|
ba0d9ba3fc | ||
|
|
de9b47b2d9 | ||
|
|
e0294f1beb | ||
|
|
052060b23c | ||
|
|
0ce9690778 | ||
|
|
fe0b268878 | ||
|
|
e447fea16b | ||
|
|
4be7a05006 | ||
|
|
6f87710232 | ||
|
|
aa722b519e | ||
|
|
08da49756f | ||
|
|
657fc6117b | ||
|
|
96245193ee | ||
|
|
640abd2d78 | ||
|
|
8758e8c8c4 | ||
|
|
9c8afb7d3c | ||
|
|
f3b50738f1 | ||
|
|
7d3512ef87 | ||
|
|
7fc69efd0a | ||
|
|
dd2f102eec | ||
|
|
86eeb76a9e | ||
|
|
20273af554 | ||
|
|
bb160d6d84 | ||
|
|
d81a67762e | ||
|
|
7a9bc7bd1d | ||
|
|
5677c4187c | ||
|
|
1a47eddb8f | ||
|
|
7ee08092d4 | ||
|
|
7900108763 | ||
|
|
94a3c92b36 | ||
|
|
127978615c | ||
|
|
5e11ed3896 | ||
|
|
1d0066f1d8 | ||
|
|
a7a774ecba | ||
|
|
892999ac98 | ||
|
|
84cf749344 | ||
|
|
597af24edb | ||
|
|
2ad6d02d59 | ||
|
|
28827b720a | ||
|
|
3713fa2dbe | ||
|
|
759eb888a6 | ||
|
|
046f5ddbc6 | ||
|
|
c77b3e705f | ||
|
|
60bec055a5 | ||
|
|
a75e22e792 | ||
|
|
2331590588 | ||
|
|
6c3bfb82ee | ||
|
|
5b471b518f | ||
|
|
2b9c58de56 | ||
|
|
ef39e8329c | ||
|
|
a4e55e8a93 | ||
|
|
851ecf50c7 | ||
|
|
f6a11dffdf | ||
|
|
30861a92ee | ||
|
|
3d2b9b22a3 | ||
|
|
c67a2955cc | ||
|
|
8bb81c796d | ||
|
|
bee568192c | ||
|
|
bde5f76d62 | ||
|
|
9b60f24c93 | ||
|
|
75b159653a | ||
|
|
7b017336a5 | ||
|
|
5d2589b10f | ||
|
|
c1ccdc07a1 | ||
|
|
5a540f646c | ||
|
|
2e7c024f45 | ||
|
|
8990c974cb | ||
|
|
09d4e71554 | ||
|
|
0b5a562f7f | ||
|
|
c0266685a8 | ||
|
|
fc30e38e78 | ||
|
|
0dbc6086d1 | ||
|
|
b46fef1b34 | ||
|
|
ae23ead4c3 | ||
|
|
d56dd4a74a | ||
|
|
8fef126444 | ||
|
|
02e344724d | ||
|
|
55dda7bc43 | ||
|
|
c88d32e83c | ||
|
|
8ee288496f | ||
|
|
50cd64d4b9 | ||
|
|
86c933dc82 | ||
|
|
86338f8a92 | ||
|
|
b3ca0a1b33 | ||
|
|
a23747e236 | ||
|
|
edeea6cae6 | ||
|
|
237d14577f | ||
|
|
f61c73eb68 | ||
|
|
5624e77125 | ||
|
|
27c40f27d2 | ||
|
|
e6849ae91a | ||
|
|
52b6372be6 | ||
|
|
e7cb7970bd | ||
|
|
bbdf5a3699 | ||
|
|
262c84123f | ||
|
|
c03b4470a0 | ||
|
|
524dbea8fc | ||
|
|
4341304c05 | ||
|
|
cb89eb10d0 | ||
|
|
115a5c4a10 | ||
|
|
06d519add4 | ||
|
|
2c52b150e8 | ||
|
|
648f42f058 | ||
|
|
61b5d96c2b | ||
|
|
17e1f6b805 | ||
|
|
be9a8b9a66 | ||
|
|
e5529eead9 | ||
|
|
de3e21dd64 | ||
|
|
e954801a0d | ||
|
|
d7471e3741 | ||
|
|
e5ed0dc04c | ||
|
|
f2f88ac009 | ||
|
|
7b09230e10 | ||
|
|
95e03476f8 | ||
|
|
aed47732e6 | ||
|
|
bccf7120cc | ||
|
|
077e3a5e9f | ||
|
|
a27491c1b6 | ||
|
|
0468bde434 | ||
|
|
52b80088cd | ||
|
|
80e227c3a7 | ||
|
|
01eee22921 | ||
|
|
2a32fbadb5 | ||
|
|
f8aefaf129 | ||
|
|
a1259edcae | ||
|
|
4925689496 | ||
|
|
404689836d | ||
|
|
b8d6914615 | ||
|
|
07b1f59285 | ||
|
|
7b19a1137f | ||
|
|
46e8818b5b | ||
|
|
477da50fe0 | ||
|
|
4338bef440 | ||
|
|
cea9d710c5 | ||
|
|
e038ff1bcf | ||
|
|
3c6de23d61 | ||
|
|
1272ffda0b | ||
|
|
1a739eeab5 | ||
|
|
9f5b7d50df | ||
|
|
0355482da3 | ||
|
|
b89aeda198 | ||
|
|
2ca470de57 | ||
|
|
6851b8a96c | ||
|
|
c49544e0e5 | ||
|
|
cbfc592e27 | ||
|
|
8071cd1b78 | ||
|
|
d8c9e1587b | ||
|
|
3945d5631f | ||
|
|
614e98d424 | ||
|
|
8fb74c63a6 | ||
|
|
d4f8dd0bfb | ||
|
|
a2eb3606ce | ||
|
|
12b317644d | ||
|
|
148cb34878 | ||
|
|
ec27ec53d7 | ||
|
|
315029b0e8 | ||
|
|
ed910aa9c7 | ||
|
|
e006943535 | ||
|
|
78f9c35800 | ||
|
|
0e7423d1d9 | ||
|
|
e6c66fbc44 | ||
|
|
730990cee5 | ||
|
|
616591ef35 | ||
|
|
ee57d101c0 | ||
|
|
eb51f72fb8 | ||
|
|
51742c5606 | ||
|
|
428c5eb71e | ||
|
|
a4175d2981 | ||
|
|
a31b4a792b | ||
|
|
735dc49d92 | ||
|
|
cbc3f7883e | ||
|
|
836a1b5801 | ||
|
|
db568b320c | ||
|
|
117a38b5f9 | ||
|
|
7993d347d8 | ||
|
|
bce2bc8b7f | ||
|
|
b949db57cf | ||
|
|
e9f227ca5b | ||
|
|
d3cf0f4adf | ||
|
|
aa2036b792 | ||
|
|
db6d6c80bf | ||
|
|
8d15acaef3 | ||
|
|
df410a9c8e | ||
|
|
d4129e04bb | ||
|
|
10f21dc01c | ||
|
|
12e2c6b4bb | ||
|
|
5495f71db7 | ||
|
|
e44973e900 | ||
|
|
6365e12735 | ||
|
|
c47add39a4 | ||
|
|
0a43c871ba | ||
|
|
01802dfcff | ||
|
|
d30bda4777 | ||
|
|
a7e69fd9fd | ||
|
|
10372fe5b3 | ||
|
|
023cc2e2d4 | ||
|
|
0bd7ee0f8e | ||
|
|
89b7110839 | ||
|
|
4dc1b420f5 | ||
|
|
aadede5ae2 | ||
|
|
fcaea59a7a | ||
|
|
4f83e43a68 | ||
|
|
6524fe1734 | ||
|
|
d222d5f683 | ||
|
|
77500deef2 | ||
|
|
dee085686f | ||
|
|
7fba4fb9ab | ||
|
|
b2514673bd | ||
|
|
625612909e | ||
|
|
f754f6ab0f | ||
|
|
9c6102ccd1 | ||
|
|
f567cadc47 | ||
|
|
2f5914dd6f | ||
|
|
c57ac183d5 | ||
|
|
4cb95731bf | ||
|
|
c1a04fe5bf | ||
|
|
76dcf08734 | ||
|
|
075b2b54dd | ||
|
|
568e1f2e53 | ||
|
|
8fafd99e3e | ||
|
|
d9db5a48a3 | ||
|
|
82193806d1 | ||
|
|
77d858dc35 | ||
|
|
2e18317f3f | ||
|
|
f538b9e948 | ||
|
|
066c932778 | ||
|
|
453914a740 | ||
|
|
a98438131f | ||
|
|
e5c713192e | ||
|
|
e44192ab7b | ||
|
|
bdcc22b465 | ||
|
|
e92f449749 | ||
|
|
a6bb0cb1ac | ||
|
|
6f456cb1d7 | ||
|
|
d05170a1e2 | ||
|
|
510b67906b | ||
|
|
7f2ede8ad2 | ||
|
|
48fcdce4ad | ||
|
|
642a85661f | ||
|
|
25dfcd8856 | ||
|
|
a118dc96e9 | ||
|
|
e0678ea7a5 | ||
|
|
e4952dbee0 | ||
|
|
16c1d1fd77 | ||
|
|
f16b1c118b | ||
|
|
1611f3c12c | ||
|
|
edaa309559 | ||
|
|
b97b226166 | ||
|
|
53898f3446 | ||
|
|
ccb4ce51cc | ||
|
|
f2cbe74a35 | ||
|
|
81a9955081 | ||
|
|
1d52d1eadd | ||
|
|
4aa6c1c99f | ||
|
|
ba6a954f45 | ||
|
|
5eb1ffde48 | ||
|
|
ff7021e798 | ||
|
|
993771dc5d | ||
|
|
f4ca38f143 | ||
|
|
d66e009979 | ||
|
|
3656dfdbcd | ||
|
|
a52da095b4 | ||
|
|
f4f1e7e8fa | ||
|
|
cfdc035617 | ||
|
|
e01eab28a2 | ||
|
|
36d2a2c731 | ||
|
|
cd3adfa14c | ||
|
|
6a266c68c3 | ||
|
|
3a9c9f3c8e | ||
|
|
a6ad163aad | ||
|
|
6c07415e24 | ||
|
|
5d1eabd74c | ||
|
|
9903ad7535 | ||
|
|
5571b99a03 | ||
|
|
387319b353 | ||
|
|
83c4f5ad66 | ||
|
|
9d68670dbd | ||
|
|
0a2f777b34 | ||
|
|
f7263873db | ||
|
|
656aed4705 | ||
|
|
36a97f4955 | ||
|
|
932d25f125 | ||
|
|
58b171cec0 | ||
|
|
aec38d83c7 | ||
|
|
3366b15937 | ||
|
|
4d276ced71 | ||
|
|
58924c6d6d | ||
|
|
1147b3e15c | ||
|
|
94a7ef2c2c | ||
|
|
8dc3e83f6d | ||
|
|
cf06c6f8e1 | ||
|
|
697025040f | ||
|
|
8809ad41d8 | ||
|
|
9b407dbdad | ||
|
|
ef79ca1fe8 | ||
|
|
24044303f5 | ||
|
|
5fec15186c | ||
|
|
23fadc0c33 | ||
|
|
6cb7a25a16 | ||
|
|
5b8d5e8e4a | ||
|
|
55057b3157 | ||
|
|
a03bab277b | ||
|
|
12c8a60ad7 | ||
|
|
c945a33f8f | ||
|
|
8fbd1e4801 | ||
|
|
5e8b2a2cbc | ||
|
|
0d58887b6a | ||
|
|
0445454c72 | ||
|
|
b5cf315ca1 | ||
|
|
85f6d807bf | ||
|
|
2a2e47a955 | ||
|
|
1b352d40ad | ||
|
|
1e07fa576a | ||
|
|
f47c240920 | ||
|
|
59e02c2da9 | ||
|
|
c00468c103 | ||
|
|
7f49ee3120 | ||
|
|
26bc6f16c0 | ||
|
|
206a5f1e72 | ||
|
|
a2b2aeb748 | ||
|
|
a0f0c634ff | ||
|
|
0907958ded | ||
|
|
97248aad16 | ||
|
|
fc444fcbd6 | ||
|
|
dac60200e0 | ||
|
|
99ea0d8997 | ||
|
|
34277e3fbd | ||
|
|
cb22387008 | ||
|
|
949126c9cb | ||
|
|
77da611411 | ||
|
|
bf3acb9cad | ||
|
|
c64efe6775 | ||
|
|
ccb8610020 | ||
|
|
effb3050b4 | ||
|
|
eb517a8e78 | ||
|
|
efd7c6b75b | ||
|
|
59cc4cea7c | ||
|
|
c7ba5e5fa3 | ||
|
|
223c13561c | ||
|
|
531f3408a0 | ||
|
|
5181395bbd | ||
|
|
ac3f346f14 | ||
|
|
0d1f930c65 | ||
|
|
bfd01a5f78 | ||
|
|
65bcd83330 | ||
|
|
b976550828 | ||
|
|
3c884a3d37 | ||
|
|
09b1d06bd1 | ||
|
|
c6e7dfedab | ||
|
|
696fa65c16 | ||
|
|
5321509380 | ||
|
|
c3d1b4e2ea | ||
|
|
67fd01138f | ||
|
|
8dd3efaa54 | ||
|
|
c5d5287938 | ||
|
|
317f2e78a9 | ||
|
|
d900824d6a | ||
|
|
4d4940a3b9 | ||
|
|
8ce5c40b28 | ||
|
|
7e71fbf241 | ||
|
|
5cb97ee00b | ||
|
|
ecb5d1e2ce | ||
|
|
81579a7e40 | ||
|
|
dcf1442b4f | ||
|
|
e69305f4fc | ||
|
|
3dac1f41d1 | ||
|
|
c50477a712 | ||
|
|
6430f13d3e | ||
|
|
7a33e65c61 | ||
|
|
e5c3b914dd | ||
|
|
0eb8d9998c | ||
|
|
8f6b962607 | ||
|
|
e56139a56d | ||
|
|
015d28ec94 | ||
|
|
dfc310b3f2 | ||
|
|
db17babe58 | ||
|
|
bcded40476 | ||
|
|
afdf2f9e2c | ||
|
|
58ca80908f | ||
|
|
3fd222b2c6 | ||
|
|
00d972c417 | ||
|
|
a383ead27f | ||
|
|
ca3cb39f12 | ||
|
|
5d05fc1d53 | ||
|
|
7b68c0c45a | ||
|
|
2c5b444ffd | ||
|
|
3b95cbcce0 | ||
|
|
6cb0c35ab8 | ||
|
|
11fe93a128 | ||
|
|
0382e07d50 | ||
|
|
70d354b016 | ||
|
|
f4ab4997c8 | ||
|
|
71ecebedce | ||
|
|
8ed2bc9d07 | ||
|
|
4538095d74 | ||
|
|
290ec9cc79 | ||
|
|
b26949e483 | ||
|
|
a13af47b3e | ||
|
|
98dd645fb5 | ||
|
|
db64515b78 | ||
|
|
c0be6a61c8 | ||
|
|
0a689e407e | ||
|
|
9089bb4cdb | ||
|
|
c90ea0e92b | ||
|
|
3d13eb51e6 | ||
|
|
aceccaf5b5 | ||
|
|
e4c77b0fef | ||
|
|
a6b98de43f | ||
|
|
1a90b6d942 | ||
|
|
c941d086d9 | ||
|
|
e09016cea5 | ||
|
|
cd45c64300 | ||
|
|
42de4378e1 | ||
|
|
68b9c81c0b | ||
|
|
da02f8004f | ||
|
|
684dcb3fc6 | ||
|
|
c8cecabc5c | ||
|
|
bc5298bf5f | ||
|
|
9e2f6d3b57 | ||
|
|
9e4184a98d | ||
|
|
003c4c3ab8 | ||
|
|
75e7b0e37d | ||
|
|
dff3f23b03 | ||
|
|
d10ad138f1 | ||
|
|
4757c4be53 | ||
|
|
842c968220 | ||
|
|
13d1a9856a | ||
|
|
7ceb4807f3 | ||
|
|
93c89f879d | ||
|
|
c34dd46067 | ||
|
|
a3971f9097 | ||
|
|
ed8e97fd2d | ||
|
|
5f520e2cf4 | ||
|
|
64915c87cf | ||
|
|
b125b0f10a | ||
|
|
58eafdfd5c | ||
|
|
01ef9a89ca | ||
|
|
af1485c753 | ||
|
|
e6164cd179 | ||
|
|
ec0f94a870 |
955 changed files with 269380 additions and 67802 deletions
1
.claude/skills
Symbolic link
1
.claude/skills
Symbolic link
|
|
@ -0,0 +1 @@
|
|||
../.skills
|
||||
16
.codex/hooks.json
Normal file
16
.codex/hooks.json
Normal file
|
|
@ -0,0 +1,16 @@
|
|||
{
|
||||
"hooks": {
|
||||
"PreToolUse": [
|
||||
{
|
||||
"matcher": "Bash",
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "/usr/bin/python3 \"$(git rev-parse --show-toplevel)/.codex/hooks/pre_tool_use_mcc_build_guard.py\"",
|
||||
"statusMessage": "Checking MCC build command policy"
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
67
.codex/hooks/pre_tool_use_mcc_build_guard.py
Normal file
67
.codex/hooks/pre_tool_use_mcc_build_guard.py
Normal file
|
|
@ -0,0 +1,67 @@
|
|||
#!/usr/bin/env python3
|
||||
import json
|
||||
import re
|
||||
import sys
|
||||
|
||||
|
||||
RAW_DOTNET_BUILD_RE = re.compile(r"(^|[\s;&|()])dotnet\s+build(\s|$)")
|
||||
ABSOLUTE_DOTNET_BUILD_RE = re.compile(r"(^|[\s;&|()])/\S*dotnet\s+build(\s|$)")
|
||||
RAW_DOTNET_PUBLISH_RE = re.compile(r"(^|[\s;&|()])dotnet\s+publish(\s|$)")
|
||||
ABSOLUTE_DOTNET_PUBLISH_RE = re.compile(r"(^|[\s;&|()])/\S*dotnet\s+publish(\s|$)")
|
||||
|
||||
|
||||
def main() -> int:
|
||||
try:
|
||||
payload = json.load(sys.stdin)
|
||||
except json.JSONDecodeError:
|
||||
return 0
|
||||
|
||||
command = payload.get("tool_input", {}).get("command", "")
|
||||
if not isinstance(command, str) or not command:
|
||||
return 0
|
||||
|
||||
if ABSOLUTE_DOTNET_BUILD_RE.search(command) or ABSOLUTE_DOTNET_PUBLISH_RE.search(command):
|
||||
return 0
|
||||
|
||||
if RAW_DOTNET_BUILD_RE.search(command):
|
||||
response = {
|
||||
"hookSpecificOutput": {
|
||||
"hookEventName": "PreToolUse",
|
||||
"permissionDecision": "deny",
|
||||
"permissionDecisionReason": (
|
||||
"Raw 'dotnet build' is blocked in this repository. "
|
||||
"Use 'source tools/mcc-env.sh && mcc-build' instead so MCC temp-build routing stays active. "
|
||||
"If you intentionally need the raw .NET CLI, call it by absolute path such as '/usr/bin/dotnet build ...' to bypass this guard."
|
||||
),
|
||||
},
|
||||
"systemMessage": (
|
||||
"Blocked raw 'dotnet build'. Use 'source tools/mcc-env.sh && mcc-build'. "
|
||||
"If you intentionally need raw .NET CLI behavior, call '/usr/bin/dotnet build ...' explicitly."
|
||||
),
|
||||
}
|
||||
json.dump(response, sys.stdout)
|
||||
sys.stdout.write("\n")
|
||||
elif RAW_DOTNET_PUBLISH_RE.search(command):
|
||||
response = {
|
||||
"hookSpecificOutput": {
|
||||
"hookEventName": "PreToolUse",
|
||||
"permissionDecision": "deny",
|
||||
"permissionDecisionReason": (
|
||||
"Raw 'dotnet publish' is blocked in this repository. "
|
||||
"Use 'source tools/mcc-env.sh && mcc-publish --rid <RID>' instead so MCC publish defaults stay aligned with the repo workflow. "
|
||||
"If you intentionally need the raw .NET CLI, call it by absolute path such as '/usr/bin/dotnet publish ...' to bypass this guard."
|
||||
),
|
||||
},
|
||||
"systemMessage": (
|
||||
"Blocked raw 'dotnet publish'. Use 'source tools/mcc-env.sh && mcc-publish --rid <RID>'. "
|
||||
"If you intentionally need raw .NET CLI behavior, call '/usr/bin/dotnet publish ...' explicitly."
|
||||
),
|
||||
}
|
||||
json.dump(response, sys.stdout)
|
||||
sys.stdout.write("\n")
|
||||
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
1
.codex/skills
Symbolic link
1
.codex/skills
Symbolic link
|
|
@ -0,0 +1 @@
|
|||
../.skills
|
||||
1
.cursor/skills
Symbolic link
1
.cursor/skills
Symbolic link
|
|
@ -0,0 +1 @@
|
|||
../.skills
|
||||
3
.cursorindexingignore
Normal file
3
.cursorindexingignore
Normal file
|
|
@ -0,0 +1,3 @@
|
|||
|
||||
# Don't index SpecStory auto-save files, but allow explicit context inclusion via @ references
|
||||
.specstory/**
|
||||
2
.gitattributes
vendored
Normal file
2
.gitattributes
vendored
Normal file
|
|
@ -0,0 +1,2 @@
|
|||
# Set the default line endings to LF
|
||||
* text=auto
|
||||
4
.github/ISSUE_TEMPLATE/bug_report_form.yaml
vendored
4
.github/ISSUE_TEMPLATE/bug_report_form.yaml
vendored
|
|
@ -9,7 +9,7 @@ body:
|
|||
attributes:
|
||||
label: Prerequisites
|
||||
options:
|
||||
- label: I made sure I am running the latest [development build](https://ci.appveyor.com/project/ORelio/minecraft-console-client/build/artifacts)
|
||||
- label: I made sure I am running the latest [development build](https://github.com/MCCTeam/Minecraft-Console-Client/releases/latest)
|
||||
required: true
|
||||
- label: I tried to [look for similar issues](https://github.com/MCCTeam/Minecraft-Console-Client/issues?q=is%3Aissue) before opening a new one
|
||||
required: true
|
||||
|
|
@ -95,4 +95,4 @@ body:
|
|||
- type: markdown
|
||||
id: credit
|
||||
attributes:
|
||||
value: Thank you for filling the bug report. Feel free to submit the report to us.
|
||||
value: Thank you for filling the bug report. Feel free to submit the report to us.
|
||||
|
|
|
|||
|
|
@ -10,9 +10,9 @@ body:
|
|||
attributes:
|
||||
label: Prerequisites
|
||||
options:
|
||||
- label: I have read and understood the [user manual](https://github.com/MCCTeam/Minecraft-Console-Client/tree/master/MinecraftClient/config)
|
||||
- label: I have read and understood the [user manual](https://mccteam.github.io/guide/)
|
||||
required: true
|
||||
- label: I made sure I am running the latest [development build](https://ci.appveyor.com/project/ORelio/Minecraft-Console-Client/build/artifacts)
|
||||
- label: I made sure I am running the latest [development build](https://github.com/MCCTeam/Minecraft-Console-Client/releases/latest)
|
||||
required: true
|
||||
- label: I tried to [look for similar feature requests](https://github.com/MCCTeam/Minecraft-Console-Client/issues?q=is%3Aissue) before opening a new one
|
||||
required: true
|
||||
|
|
@ -69,4 +69,4 @@ body:
|
|||
- type: markdown
|
||||
id: credit
|
||||
attributes:
|
||||
value: Thank you for filling the request form. Feel free to submit the request to us.
|
||||
value: Thank you for filling the request form. Feel free to submit the request to us.
|
||||
|
|
|
|||
307
.github/workflows/build-and-release.yml
vendored
307
.github/workflows/build-and-release.yml
vendored
|
|
@ -1,119 +1,246 @@
|
|||
name: Build
|
||||
name: Build MCC and Documents
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [ master ]
|
||||
branches:
|
||||
- master
|
||||
workflow_dispatch:
|
||||
|
||||
env:
|
||||
PROJECT: "MinecraftClient"
|
||||
target-version: "net6.0"
|
||||
compile-flags: "--no-self-contained -c Release -p:UseAppHost=true -p:IncludeNativeLibrariesForSelfExtract=true -p:DebugType=None"
|
||||
target-version: "net10.0"
|
||||
dotnet-version: "10.0.x"
|
||||
compile-flags: "--self-contained=true -c Release -p:UseAppHost=true -p:IncludeNativeLibrariesForSelfExtract=true -p:EnableCompressionInSingleFile=true -p:DebugType=Embedded -p:PublishSingleFile=true"
|
||||
|
||||
jobs:
|
||||
build:
|
||||
determine-build:
|
||||
runs-on: ubuntu-slim
|
||||
outputs:
|
||||
skip: ${{ steps.check-skip.outputs.skip }}
|
||||
steps:
|
||||
- name: Check skip CI
|
||||
id: check-skip
|
||||
run: |
|
||||
LOWER=$(echo "$COMMIT_MSG" | tr '[:upper:]' '[:lower:]')
|
||||
if echo "$LOWER" | grep -qE 'skip.?ci|ci.?skip'; then
|
||||
echo "skip=true" >> $GITHUB_OUTPUT
|
||||
else
|
||||
echo "skip=false" >> $GITHUB_OUTPUT
|
||||
fi
|
||||
env:
|
||||
COMMIT_MSG: ${{ github.event.head_commit.message }}
|
||||
|
||||
fetch-translations:
|
||||
strategy:
|
||||
fail-fast: true
|
||||
runs-on: ubuntu-latest
|
||||
needs: determine-build
|
||||
|
||||
if: ${{ needs.determine-build.outputs.skip != 'true' }}
|
||||
|
||||
timeout-minutes: 15
|
||||
|
||||
steps:
|
||||
- name: Setup Project Path
|
||||
run: |
|
||||
echo project-path=${{ github.workspace }}/${{ env.PROJECT }} >> $GITHUB_ENV
|
||||
|
||||
- name: Setup Output Paths
|
||||
run: |
|
||||
echo win-out-path=${{ env.project-path }}/bin/Release/${{ env.target-version }}/win-x64/publish/ >> $GITHUB_ENV
|
||||
echo linux-out-path=${{ env.project-path }}/bin/Release/${{ env.target-version }}/linux-x64/publish/ >> $GITHUB_ENV
|
||||
echo osx-out-path=${{ env.project-path }}/bin/Release/${{ env.target-version }}/osx-x64/publish/ >> $GITHUB_ENV
|
||||
echo linux-arm64-out-path=${{ env.project-path }}/bin/Release/${{ env.target-version }}/linux-arm64/publish/ >> $GITHUB_ENV
|
||||
|
||||
- name: Setup .NET SDK
|
||||
uses: actions/setup-dotnet@v2.1.0
|
||||
- name: Check cache
|
||||
uses: actions/cache/restore@v3
|
||||
id: cache-check
|
||||
with:
|
||||
path: ${{ github.workspace }}/*
|
||||
key: "translation-${{ github.sha }}"
|
||||
lookup-only: true
|
||||
restore-keys: "translation-"
|
||||
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v2
|
||||
if: steps.cache-check.outputs.cache-hit != 'true'
|
||||
uses: actions/checkout@v3
|
||||
with:
|
||||
fetch-depth: 0
|
||||
submodules: 'true'
|
||||
|
||||
- name: Get Version DateTime
|
||||
id: date-version
|
||||
uses: nanzm/get-time-action@v1.0
|
||||
with:
|
||||
timeZone: 0
|
||||
format: 'YYYY-MM-DD'
|
||||
|
||||
- name: VersionInfo
|
||||
- name: Check Crowdin secrets
|
||||
id: crowdin-check
|
||||
run: |
|
||||
COMMIT=$(echo ${{ github.sha }} | cut -c 1-7)
|
||||
echo '' >> ${{ env.project-path }}\Properties\AssemblyInfo.cs
|
||||
echo "[assembly: AssemblyConfiguration(\"GitHub build ${{ github.run_number }}, built on ${{ steps.date-version.outputs.time }} from commit $COMMIT\")]" >> ${{ env.project-path }}\Properties\AssemblyInfo.cs
|
||||
|
||||
- name: Build for Windows
|
||||
run: dotnet publish ${{ env.project-path }}.sln -f ${{ env.target-version }} -r win-x64 ${{ env.compile-flags }}
|
||||
if [ -z "$CROWDIN_PROJECT_ID" ] || [ -z "$CROWDIN_PERSONAL_TOKEN" ]; then
|
||||
echo "available=false" >> $GITHUB_OUTPUT
|
||||
else
|
||||
echo "available=true" >> $GITHUB_OUTPUT
|
||||
fi
|
||||
env:
|
||||
CROWDIN_PROJECT_ID: ${{ secrets.CROWDIN_PROJECT_ID }}
|
||||
CROWDIN_PERSONAL_TOKEN: ${{ secrets.CROWDIN_TOKEN }}
|
||||
|
||||
- name: Zip Windows Build
|
||||
run: zip -qq -r windows.zip *
|
||||
working-directory: ${{ env.win-out-path }}
|
||||
- name: Download translations from crowdin
|
||||
uses: crowdin/github-action@v2.4.0
|
||||
if: steps.cache-check.outputs.cache-hit != 'true' && steps.crowdin-check.outputs.available == 'true'
|
||||
with:
|
||||
upload_sources: ${{ github.repository == 'MCCTeam/Minecraft-Console-Client' }}
|
||||
upload_translations: false
|
||||
download_translations: true
|
||||
|
||||
- name: Build for Linux
|
||||
run: dotnet publish ${{ env.project-path }}.sln -f ${{ env.target-version }} -r linux-x64 ${{ env.compile-flags }}
|
||||
localization_branch_name: l10n_master
|
||||
create_pull_request: false
|
||||
push_translations: false
|
||||
|
||||
- name: Zip Linux Build
|
||||
run: zip -qq -r linux.zip *
|
||||
working-directory: ${{ env.linux-out-path }}
|
||||
|
||||
- name: Build for ARM64 Linux
|
||||
run: dotnet publish ${{ env.project-path }}.sln -f ${{ env.target-version }} -r linux-arm64 ${{ env.compile-flags }}
|
||||
base_path: ${{ github.workspace }}
|
||||
env:
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
CROWDIN_PROJECT_ID: ${{ secrets.CROWDIN_PROJECT_ID }}
|
||||
CROWDIN_PERSONAL_TOKEN: ${{ secrets.CROWDIN_TOKEN }}
|
||||
|
||||
- name: Zip ARM64 Linux Build
|
||||
run: zip -qq -r linux-arm64.zip *
|
||||
working-directory: ${{ env.linux-arm64-out-path }}
|
||||
|
||||
- name: Build for OSX
|
||||
run: dotnet publish ${{ env.project-path }}.sln -f ${{ env.target-version }} -r osx-x64 ${{ env.compile-flags }}
|
||||
|
||||
- name: Zip OSX Build
|
||||
run: zip -qq -r osx.zip *
|
||||
working-directory: ${{ env.osx-out-path }}
|
||||
|
||||
- name: Get Release DateTime
|
||||
id: date-release
|
||||
uses: nanzm/get-time-action@v1.0
|
||||
- name: Save cache
|
||||
uses: actions/cache/save@v3
|
||||
if: steps.cache-check.outputs.cache-hit != 'true'
|
||||
with:
|
||||
timeZone: 0
|
||||
format: 'YYYYMMDD'
|
||||
path: ${{ github.workspace }}/*
|
||||
key: "translation-${{ github.sha }}"
|
||||
|
||||
- name: Windows Release
|
||||
uses: tix-factory/release-manager@v1
|
||||
with:
|
||||
github_token: ${{ secrets.GITHUB_TOKEN }}
|
||||
mode: uploadReleaseAsset
|
||||
filePath: ${{ env.win-out-path }}windows.zip
|
||||
assetName: ${{ env.PROJECT }}-windows.zip
|
||||
tag: ${{ format('{0}-{1}', steps.date-release.outputs.time, github.run_number) }}
|
||||
create-tag:
|
||||
runs-on: ubuntu-slim
|
||||
timeout-minutes: 5 # Wait 5 minutes in case of network issues/etc
|
||||
needs: determine-build
|
||||
if: ${{ needs.determine-build.outputs.skip != 'true' }}
|
||||
steps:
|
||||
- id: make-tag
|
||||
run: |
|
||||
TAG="$(date -u +'%Y%m%d')-${{ github.run_number }}"
|
||||
echo "tag=$TAG" >> $GITHUB_OUTPUT
|
||||
echo "TAG=$TAG" >> $GITHUB_ENV
|
||||
|
||||
- name: Create Release Tag
|
||||
uses: actions/github-script@v7
|
||||
with:
|
||||
script: |
|
||||
const tag = process.env.TAG;
|
||||
try {
|
||||
await github.rest.git.createRef({
|
||||
owner: context.repo.owner,
|
||||
repo: context.repo.repo,
|
||||
ref: `refs/tags/${tag}`,
|
||||
sha: context.sha
|
||||
});
|
||||
} catch(error) {
|
||||
if (error.message.includes('already exists')) {
|
||||
console.log(`Tag ${tag} already exists`);
|
||||
} else {
|
||||
throw error;
|
||||
}
|
||||
}
|
||||
outputs:
|
||||
build-tag: ${{ steps.make-tag.outputs.tag }}
|
||||
|
||||
- name: Linux Release
|
||||
uses: tix-factory/release-manager@v1
|
||||
with:
|
||||
github_token: ${{ secrets.GITHUB_TOKEN }}
|
||||
mode: uploadReleaseAsset
|
||||
filePath: ${{ env.linux-out-path }}linux.zip
|
||||
assetName: ${{ env.PROJECT }}-linux.zip
|
||||
tag: ${{ format('{0}-{1}', steps.date-release.outputs.time, github.run_number) }}
|
||||
build:
|
||||
runs-on: ubuntu-latest
|
||||
# Check if we're not skipping build, tag is created, and translations successfully fetched (or skipped)
|
||||
if: ${{ needs.determine-build.outputs.skip != 'true' &&
|
||||
needs.create-tag.result == 'success' &&
|
||||
(needs.fetch-translations.result == 'success' || needs.fetch-translations.result == 'skipped')
|
||||
}}
|
||||
needs: [determine-build, fetch-translations, create-tag]
|
||||
timeout-minutes: 15
|
||||
strategy:
|
||||
matrix:
|
||||
target: [win-x86, win-x64, win-arm64, linux-x64, linux-arm, linux-arm64, osx-x64, osx-arm64]
|
||||
|
||||
- name: Linux ARM64 Release
|
||||
uses: tix-factory/release-manager@v1
|
||||
steps:
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v3
|
||||
with:
|
||||
github_token: ${{ secrets.GITHUB_TOKEN }}
|
||||
mode: uploadReleaseAsset
|
||||
filePath: ${{ env.linux-arm64-out-path }}linux-arm64.zip
|
||||
assetName: ${{ env.PROJECT }}-linux-arm64.zip
|
||||
tag: ${{ format('{0}-{1}', steps.date-release.outputs.time, github.run_number) }}
|
||||
fetch-depth: 0
|
||||
submodules: 'true'
|
||||
|
||||
- name: Get Current Date
|
||||
run: |
|
||||
echo date_dashed=$(date -u +'%Y-%m-%d') >> $GITHUB_ENV
|
||||
|
||||
- name: Restore Translations (if available)
|
||||
uses: actions/cache/restore@v3
|
||||
with:
|
||||
path: ${{ github.workspace }}/*
|
||||
key: "translation-${{ github.sha }}"
|
||||
restore-keys: "translation-"
|
||||
|
||||
- name: Setup .NET SDK
|
||||
uses: actions/setup-dotnet@v4
|
||||
with:
|
||||
dotnet-version: ${{ env.dotnet-version }}
|
||||
|
||||
- name: Setup Environment Variables
|
||||
run: |
|
||||
PROJECT_PATH=${{ github.workspace }}/${{ env.PROJECT }}
|
||||
FILE_EXT=${{ (startsWith(matrix.target, 'win') && '.exe') || '' }}
|
||||
TARGET_OUT_PATH=$PROJECT_PATH/bin/Release/${{ env.target-version }}/${{ matrix.target }}/publish/
|
||||
|
||||
echo "project-path=$PROJECT_PATH" >> $GITHUB_ENV
|
||||
echo "file-ext=$FILE_EXT" >> $GITHUB_ENV
|
||||
echo "target-out-path=$TARGET_OUT_PATH" >> $GITHUB_ENV
|
||||
echo "assembly-info=$PROJECT_PATH/Properties/AssemblyInfo.cs" >> $GITHUB_ENV
|
||||
echo "build-version-info=${{ needs.create-tag.outputs.build-tag }}" >> $GITHUB_ENV
|
||||
echo "commit=$(echo ${{ github.sha }} | cut -c 1-7)" >> $GITHUB_ENV
|
||||
|
||||
- name: Setup Binaries Path
|
||||
run: |
|
||||
echo built-executable-path=${{ env.target-out-path }}${{ env.PROJECT }}${{ env.file-ext }} >> $GITHUB_ENV
|
||||
|
||||
- name: OSX Release
|
||||
uses: tix-factory/release-manager@v1
|
||||
- name: Set Version Info
|
||||
run: |
|
||||
echo '' >> ${{ env.assembly-info }}
|
||||
echo "[assembly: AssemblyConfiguration(\"GitHub build ${{ github.run_number }}, built on ${{ env.date_dashed }} from commit ${{ env.commit }}\")]" >> ${{ env.assembly-info }}
|
||||
|
||||
- name: Inject Sentry DSN (if applicable)
|
||||
if: ${{ github.repository == 'MCCTeam/Minecraft-Console-Client' }}
|
||||
run: |
|
||||
grep -q 'SentryDSN = "";' ${{ env.project-path }}/Program.cs || { echo "SentryDSN pattern not found in Program.cs"; exit 1; }
|
||||
sed -i -e 's|SentryDSN = "";|SentryDSN = "${{ secrets.SENTRY_DSN }}";|g' ${{ env.project-path }}/Program.cs
|
||||
|
||||
- name: Build Target
|
||||
run: dotnet publish ${{ env.project-path }}/${{ env.PROJECT }}.csproj -f ${{ env.target-version }} -r ${{ matrix.target }} ${{ env.compile-flags }}
|
||||
env:
|
||||
DOTNET_NOLOGO: true
|
||||
|
||||
- name: Rename Binary
|
||||
run: |
|
||||
mv ${{ env.built-executable-path }} ${{ env.PROJECT }}-${{ env.build-version-info }}-${{ matrix.target }}${{ env.file-ext }}
|
||||
|
||||
- name: Upload Artifact
|
||||
uses: actions/upload-artifact@v4
|
||||
with:
|
||||
github_token: ${{ secrets.GITHUB_TOKEN }}
|
||||
mode: uploadReleaseAsset
|
||||
filePath: ${{ env.osx-out-path }}osx.zip
|
||||
assetName: ${{ env.PROJECT }}-osx.zip
|
||||
tag: ${{ format('{0}-{1}', steps.date-release.outputs.time, github.run_number) }}
|
||||
name: ${{ env.PROJECT }}-${{ env.build-version-info }}-${{ matrix.target }}
|
||||
path: ${{ env.PROJECT }}-${{ env.build-version-info }}-${{ matrix.target }}${{ env.file-ext }}
|
||||
if-no-files-found: error
|
||||
|
||||
create-release:
|
||||
runs-on: ubuntu-slim
|
||||
needs: [create-tag, build]
|
||||
if: ${{ needs.build.result == 'success' && needs.create-tag.result == 'success' }}
|
||||
steps:
|
||||
- name: Download All Artifacts
|
||||
uses: actions/download-artifact@v4
|
||||
with:
|
||||
path: artifacts/
|
||||
merge-multiple: true
|
||||
|
||||
- name: Truncate commit message for release name
|
||||
id: release-name
|
||||
run: |
|
||||
SUBJECT=$(echo "$COMMIT_MSG" | head -n 1)
|
||||
MAX=220
|
||||
TRUNCATED="${SUBJECT:0:$MAX}"
|
||||
echo "name=${BUILD_TAG}: $TRUNCATED" >> $GITHUB_OUTPUT
|
||||
env:
|
||||
COMMIT_MSG: ${{ github.event.head_commit.message }}
|
||||
BUILD_TAG: ${{ needs.create-tag.outputs.build-tag }}
|
||||
|
||||
- name: Create Release
|
||||
uses: ncipollo/release-action@v1.14.0
|
||||
with:
|
||||
token: ${{ secrets.GITHUB_TOKEN }}
|
||||
artifacts: "artifacts/**/*"
|
||||
tag: ${{ needs.create-tag.outputs.build-tag }}
|
||||
name: ${{ steps.release-name.outputs.name }}
|
||||
generateReleaseNotes: true
|
||||
artifactErrorsFailBuild: true
|
||||
allowUpdates: true
|
||||
makeLatest: true
|
||||
omitBodyDuringUpdate: true
|
||||
omitNameDuringUpdate: true
|
||||
replacesArtifacts: true
|
||||
|
|
|
|||
79
.github/workflows/deploy-doc-only.yml
vendored
Normal file
79
.github/workflows/deploy-doc-only.yml
vendored
Normal file
|
|
@ -0,0 +1,79 @@
|
|||
name: Build Documents
|
||||
|
||||
on:
|
||||
schedule:
|
||||
- cron: '45 0 * * *'
|
||||
workflow_dispatch:
|
||||
|
||||
jobs:
|
||||
check-secrets:
|
||||
runs-on: ubuntu-latest
|
||||
outputs:
|
||||
has-deploy-token: ${{ steps.check.outputs.has-deploy-token }}
|
||||
steps:
|
||||
- name: Check required secrets
|
||||
id: check
|
||||
run: |
|
||||
if [ -z "$GH_PAGES_TOKEN" ]; then
|
||||
echo "has-deploy-token=false" >> $GITHUB_OUTPUT
|
||||
echo "::warning::GH_PAGES_TOKEN is not set, skipping documentation deployment."
|
||||
else
|
||||
echo "has-deploy-token=true" >> $GITHUB_OUTPUT
|
||||
fi
|
||||
env:
|
||||
GH_PAGES_TOKEN: ${{ secrets.GH_PAGES_TOKEN }}
|
||||
|
||||
Build:
|
||||
runs-on: ubuntu-latest
|
||||
needs: check-secrets
|
||||
if: ${{ needs.check-secrets.outputs.has-deploy-token == 'true' }}
|
||||
|
||||
steps:
|
||||
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v3
|
||||
with:
|
||||
fetch-depth: 0
|
||||
submodules: 'true'
|
||||
|
||||
- name: Check Crowdin secrets
|
||||
id: crowdin-check
|
||||
run: |
|
||||
if [ -z "$CROWDIN_PROJECT_ID" ] || [ -z "$CROWDIN_PERSONAL_TOKEN" ]; then
|
||||
echo "available=false" >> $GITHUB_OUTPUT
|
||||
echo "::warning::Crowdin secrets not set, skipping translation download."
|
||||
else
|
||||
echo "available=true" >> $GITHUB_OUTPUT
|
||||
fi
|
||||
env:
|
||||
CROWDIN_PROJECT_ID: ${{ secrets.CROWDIN_PROJECT_ID }}
|
||||
CROWDIN_PERSONAL_TOKEN: ${{ secrets.CROWDIN_TOKEN }}
|
||||
|
||||
- name: Download translations from crowdin
|
||||
if: ${{ steps.crowdin-check.outputs.available == 'true' }}
|
||||
uses: crowdin/github-action@v2.4.0
|
||||
with:
|
||||
upload_sources: false
|
||||
upload_translations: false
|
||||
download_translations: true
|
||||
|
||||
localization_branch_name: l10n_master
|
||||
create_pull_request: false
|
||||
push_translations: false
|
||||
|
||||
base_path: ${{ github.workspace }}
|
||||
env:
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
CROWDIN_PROJECT_ID: ${{ secrets.CROWDIN_PROJECT_ID }}
|
||||
CROWDIN_PERSONAL_TOKEN: ${{ secrets.CROWDIN_TOKEN }}
|
||||
|
||||
- name: Deploy Documentation Site
|
||||
uses: jenkey2011/vuepress-deploy@master
|
||||
env:
|
||||
ACCESS_TOKEN: ${{ secrets.GH_PAGES_TOKEN }}
|
||||
TARGET_REPO: MCCTeam/MCCTeam.github.io
|
||||
TARGET_BRANCH: master
|
||||
BUILD_SCRIPT: export NODE_OPTIONS=--max-old-space-size=8192 && yarn --cwd ./docs/ && yarn --cwd ./docs/ docs:build
|
||||
BUILD_DIR: docs/.vuepress/dist
|
||||
COMMIT_MESSAGE: Build from ${{ github.sha }}
|
||||
CNAME: https://mccteam.github.io
|
||||
33
.github/workflows/upload-translation-source-only.yml
vendored
Normal file
33
.github/workflows/upload-translation-source-only.yml
vendored
Normal file
|
|
@ -0,0 +1,33 @@
|
|||
name: Upload translation sources
|
||||
|
||||
on:
|
||||
workflow_dispatch:
|
||||
|
||||
jobs:
|
||||
Sync:
|
||||
runs-on: ubuntu-latest
|
||||
|
||||
steps:
|
||||
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v3
|
||||
with:
|
||||
fetch-depth: 0
|
||||
submodules: 'true'
|
||||
|
||||
- name: Upload translation sources to crowdin
|
||||
uses: crowdin/github-action@v1.6.0
|
||||
with:
|
||||
upload_sources: true
|
||||
upload_translations: false
|
||||
download_translations: false
|
||||
|
||||
localization_branch_name: l10n_master
|
||||
create_pull_request: false
|
||||
push_translations: false
|
||||
|
||||
base_path: ${{ github.workspace }}
|
||||
env:
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
CROWDIN_PROJECT_ID: ${{ secrets.CROWDIN_PROJECT_ID }}
|
||||
CROWDIN_PERSONAL_TOKEN: ${{ secrets.CROWDIN_TOKEN }}
|
||||
51
.gitignore
vendored
51
.gitignore
vendored
|
|
@ -8,11 +8,15 @@
|
|||
/Other/
|
||||
/.vs/
|
||||
SessionCache.ini
|
||||
.*
|
||||
!/.github
|
||||
/packages
|
||||
/packages
|
||||
|
||||
# OS-generated files
|
||||
.DS_Store
|
||||
Thumbs.db
|
||||
desktop.ini
|
||||
ehthumbs.db
|
||||
._*
|
||||
|
||||
## Ignore Visual Studio temporary files, build results, and
|
||||
## files generated by popular Visual Studio add-ons.
|
||||
##
|
||||
|
|
@ -384,7 +388,7 @@ FodyWeavers.xsd
|
|||
.vscode/*
|
||||
!.vscode/settings.json
|
||||
!.vscode/tasks.json
|
||||
!.vscode/launch.json
|
||||
!.vscode/launch.json``
|
||||
!.vscode/extensions.json
|
||||
*.code-workspace
|
||||
|
||||
|
|
@ -402,3 +406,42 @@ FodyWeavers.xsd
|
|||
.idea/
|
||||
*.sln.iml
|
||||
*.sln.iml
|
||||
|
||||
# docs
|
||||
!/docs/.vuepress
|
||||
/docs/.vuepress/.cache
|
||||
/docs/.vuepress/.temp
|
||||
/docs/.vuepress/dist
|
||||
|
||||
# translations
|
||||
/MinecraftClient/Resources/Translations/Translations.*.resx
|
||||
/MinecraftClient/Resources/AsciiArt/AsciiArt.*.resx
|
||||
/MinecraftClient/Resources/ConfigComments/ConfigComments.*.resx
|
||||
|
||||
/docs/.vuepress/translations/*.json
|
||||
!/docs/.vuepress/translations/en.json
|
||||
|
||||
/docs/l10n/
|
||||
/docs/.vuepress/public/MCC-README/
|
||||
/docs/superpowers/
|
||||
|
||||
# Floder to store the decompiled Minecraft official source code
|
||||
/MinecraftOfficial/
|
||||
|
||||
# Possible debug files
|
||||
/lang/*
|
||||
/mcc_input.txt
|
||||
/MinecraftClient.ini
|
||||
/MinecraftClient.backup.ini
|
||||
|
||||
# SpecStory files
|
||||
/.specstory/
|
||||
/.vscode/settings.json
|
||||
|
||||
# Other
|
||||
/Sentry/
|
||||
/downloads/
|
||||
server.pid
|
||||
|
||||
# Crowdin translation automation working directory
|
||||
/.crowdin-translate/
|
||||
|
|
|
|||
1165
.skills/csharp-best-practices/SKILL.md
Normal file
1165
.skills/csharp-best-practices/SKILL.md
Normal file
File diff suppressed because it is too large
Load diff
|
|
@ -0,0 +1,82 @@
|
|||
---
|
||||
description: >-
|
||||
Context-specific async guidance for library code, ui apps, asp.net core,
|
||||
background work, task.run, configureawait, and performance-sensitive design.
|
||||
metadata:
|
||||
tags: [configureawait, task.run, asp.net core, ui, library, performance]
|
||||
source: mixed
|
||||
---
|
||||
|
||||
# Context and Tradeoffs
|
||||
|
||||
## Library code versus app code
|
||||
|
||||
### General-purpose library code
|
||||
- Prefer APIs that expose true async for I/O-bound work.
|
||||
- Do not add async wrappers around purely compute-bound methods just to look modern. Expose sync compute APIs and let callers decide whether to offload.
|
||||
- `ConfigureAwait(false)` is a strong default when the library does not need the caller’s context.
|
||||
- Avoid ambient assumptions about a UI thread, request context, or test framework behavior.
|
||||
|
||||
### App code
|
||||
- Prefer the style that fits the app model.
|
||||
- UI code often needs the original context after `await`.
|
||||
- ASP.NET Core request code normally does not need `Task.Run` just to stay responsive, because it already runs on thread pool threads.
|
||||
- Do not present “ASP.NET Core has no synchronization context” as proof that every `ConfigureAwait(false)` discussion is obsolete.
|
||||
|
||||
## `Task.Run` boundaries
|
||||
|
||||
### Good uses
|
||||
- Offload CPU-bound work so a UI thread can stay responsive.
|
||||
- Offload CPU work from a caller when that scheduling boundary is deliberate.
|
||||
|
||||
### Weak uses
|
||||
- Wrapping synchronous I/O to pretend it is true async I/O.
|
||||
- Calling `Task.Run` and immediately awaiting it in ASP.NET Core request handling when no CPU offload goal exists.
|
||||
- Using `Task.Run` to hide blocking APIs instead of fixing the underlying API choice.
|
||||
|
||||
## Fire-and-forget
|
||||
|
||||
### Assume unsafe until proven otherwise
|
||||
A background task needs answers for all of these:
|
||||
- Who owns its lifetime?
|
||||
- How are exceptions observed?
|
||||
- How does shutdown cancel it?
|
||||
- Does it touch scoped services or request-bound objects?
|
||||
- Does work need retries, backpressure, or queueing?
|
||||
|
||||
### Safer alternatives
|
||||
- Await the task normally.
|
||||
- Queue work to an owned background component.
|
||||
- In ASP.NET Core, prefer hosted services or a dedicated background queue pattern for long-lived work.
|
||||
- If scoped services are required in background processing, create an explicit scope instead of capturing request scope objects.
|
||||
|
||||
## `ConfigureAwait`
|
||||
|
||||
### Strong recommendation
|
||||
- In general-purpose libraries, use `ConfigureAwait(false)` unless the continuation must run in the captured context.
|
||||
|
||||
### Weak recommendation
|
||||
- “Always use it in app code.”
|
||||
- “Never use it on .NET Core.”
|
||||
- “Use it once at the first await and you are done.”
|
||||
|
||||
### Review note
|
||||
If code after the `await` needs a specific context, say so explicitly. If it does not, the recommendation depends on whether the code is app-level or general-purpose library code.
|
||||
|
||||
## Performance guidance
|
||||
|
||||
### Correctness first
|
||||
Do not trade API clarity for speculative micro-optimizations.
|
||||
|
||||
### `ValueTask` is performance-specialized
|
||||
Recommend it only when most of these are true:
|
||||
1. the method is called very frequently
|
||||
2. it often completes synchronously or from a reusable source
|
||||
3. allocation reduction matters on measurements
|
||||
4. consumers can respect single-consumer semantics
|
||||
5. task combinator ergonomics are not central to the API
|
||||
|
||||
### Throttling and concurrency control
|
||||
- `Task.WhenAll` expresses concurrency; it does not limit it.
|
||||
- For bounded concurrency, use an async gate such as `SemaphoreSlim.WaitAsync`, or platform helpers such as `Parallel.ForEachAsync` when the workload fits.
|
||||
- Always define what happens to remaining work after the first completion or first failure.
|
||||
105
.skills/csharp-best-practices/references/async-core-guidance.md
Normal file
105
.skills/csharp-best-practices/references/async-core-guidance.md
Normal file
|
|
@ -0,0 +1,105 @@
|
|||
---
|
||||
description: >-
|
||||
Source-backed core guidance for task, valuetask, cancellation, exception flow,
|
||||
blocking, and concurrency in c# async code reviews and implementations.
|
||||
metadata:
|
||||
tags: [csharp, async, task, valuetask, cancellation, exceptions, concurrency]
|
||||
source: mixed
|
||||
---
|
||||
|
||||
# Core Guidance
|
||||
|
||||
## Facts from official .NET documentation
|
||||
|
||||
### 1. Return types and `async void`
|
||||
- Async methods should normally return `Task` or `Task<T>`.
|
||||
- `async void` is intended for event handlers; callers cannot await it and exception handling differs.
|
||||
- TAP methods that return awaitable types conventionally use the `Async` suffix.
|
||||
|
||||
### 2. Blocking on async
|
||||
- `Task<T>.Result` is blocking. Prefer `await` in most cases.
|
||||
- Blocking can deadlock in context-bound environments and reduces scalability even when it does not deadlock.
|
||||
- `await` on a faulted task rethrows one exception directly; `.Wait()` and `.Result` wrap failures in `AggregateException`.
|
||||
|
||||
### 3. `Task` versus `ValueTask`
|
||||
- Default to `Task` or `Task<T>` unless there is a demonstrated reason not to.
|
||||
- `ValueTask` has stricter usage rules. A given instance should generally be awaited only once.
|
||||
- Do not await the same `ValueTask` multiple times, call `AsTask()` multiple times, or mix consumption techniques on the same instance.
|
||||
- For synchronously successful `Task`-returning methods, `Task.CompletedTask` is the normal zero-result completion value.
|
||||
|
||||
### 4. Cancellation
|
||||
- If a TAP method supports cancellation, expose a `CancellationToken`.
|
||||
- Pass the token to nested operations that should participate in cancellation.
|
||||
- If an async method throws `OperationCanceledException` associated with the method’s token, the returned task transitions to `Canceled`.
|
||||
- After a method has completed its work successfully, do not report cancellation instead of success.
|
||||
|
||||
### 5. Exception flow and task combinators
|
||||
- `Task.WhenAll` does not block the calling thread.
|
||||
- If any supplied task faults, the `WhenAll` task faults and aggregates the unwrapped exceptions from the component tasks.
|
||||
- If none fault and at least one is canceled, the `WhenAll` task is canceled.
|
||||
- `Task.WhenAny` returns a task that completes successfully with the first completed task as its result, even when that winning task itself is faulted or canceled.
|
||||
- After `WhenAny`, await the returned winner task to propagate its outcome.
|
||||
- The remaining tasks continue unless you cancel or otherwise handle them.
|
||||
|
||||
## Expert guidance that is strong and technically grounded
|
||||
|
||||
### Stephen Toub
|
||||
- Use `ConfigureAwait(false)` as the general default for general-purpose library code, because library code should not depend on an app model’s context.
|
||||
- App-level code is different. UI code often needs the captured context. ASP.NET Core also changes the deadlock discussion because it does not install the classic ASP.NET style synchronization context, but that does not make blanket `ConfigureAwait` advice strong.
|
||||
- `ValueTask<T>` exists mainly to avoid allocations on frequently synchronous success paths. It is not a general replacement for `Task<T>` because `Task` is more flexible for multiple awaits, caching, and combinators.
|
||||
|
||||
### Andrew Arnott
|
||||
- Propagate the token until the point of no cancellation.
|
||||
- Validate arguments before cancellation checks when argument validation should always run.
|
||||
- Prefer catching `OperationCanceledException` rather than `TaskCanceledException` in general-purpose logic.
|
||||
- Keep `CancellationToken` last in the parameter list; make it optional mainly on public APIs, not necessarily on internal methods.
|
||||
|
||||
### Stephen Cleary
|
||||
- “Async all the way” is a strong design guideline, not an absolute law of physics. Sync bridges exist, but they are specialized boundary decisions, not a normal code review recommendation.
|
||||
- `async void` and sync-over-async both create real observability and composition problems even when a sample appears to work.
|
||||
|
||||
## Naming and testability
|
||||
|
||||
### Naming
|
||||
- TAP methods that return awaitable types conventionally use the `Async` suffix. Do not force renames when an interface, base class, or event pattern already dictates the name.
|
||||
|
||||
### Testability
|
||||
- Favor awaitable APIs over hidden work so tests can await completion, assert faults, and drive cancellation deterministically.
|
||||
- Prefer explicit background components, injected clocks, and owned queues over ad hoc fire-and-forget logic that tests cannot observe.
|
||||
|
||||
## Synthesis for agents
|
||||
|
||||
### Code review defaults
|
||||
- Treat `.Result`, `.Wait()`, and `GetAwaiter().GetResult()` as likely defects unless the code is a deliberate sync boundary and the caller explicitly cannot be async.
|
||||
- Prefer `Task`/`Task<T>` for API design. Require an explicit reason before recommending `ValueTask`.
|
||||
- Require cancellation behavior to be coherent: accepted, propagated, and not silently dropped.
|
||||
- Prefer `await Task.WhenAll(...)` for independent operations started before awaiting.
|
||||
- Treat `Task.WhenAny(...)` as incomplete until the winner is awaited and losers are canceled, observed, or intentionally left running.
|
||||
|
||||
### Minimal examples
|
||||
|
||||
#### Avoid sync-over-async
|
||||
```csharp
|
||||
// bad
|
||||
var user = client.GetUserAsync(id).Result;
|
||||
|
||||
// better
|
||||
var user = await client.GetUserAsync(id);
|
||||
```
|
||||
|
||||
#### Use `Task.WhenAll` for parallel I/O
|
||||
```csharp
|
||||
var userTask = repo.GetUserAsync(id, ct);
|
||||
var ordersTask = repo.GetOrdersAsync(id, ct);
|
||||
await Task.WhenAll(userTask, ordersTask);
|
||||
return new Dashboard(await userTask, await ordersTask);
|
||||
```
|
||||
|
||||
#### Be conservative with `ValueTask`
|
||||
```csharp
|
||||
// default
|
||||
Task<Item?> GetAsync(string key, CancellationToken ct);
|
||||
|
||||
// specialized hot path only when justified
|
||||
ValueTask<Item?> TryGetCachedAsync(string key);
|
||||
```
|
||||
|
|
@ -0,0 +1,59 @@
|
|||
---
|
||||
description: >-
|
||||
Authority notes and citations for the c# async best practices skill, separating
|
||||
official documentation, expert interpretation, and synthesized guidance.
|
||||
metadata:
|
||||
tags: [sources, citations, authority, notes]
|
||||
source: external
|
||||
---
|
||||
|
||||
# Source Notes
|
||||
|
||||
## Official facts
|
||||
|
||||
- Microsoft Learn, "Implementing the Task-based Asynchronous Pattern"
|
||||
- https://learn.microsoft.com/en-us/dotnet/standard/asynchronous-programming-patterns/implementing-the-task-based-asynchronous-pattern
|
||||
- Return types, cancellation behavior, `Task.Run` boundaries, and TAP implementation guidance.
|
||||
- Microsoft Learn, "Consuming the Task-based Asynchronous Pattern"
|
||||
- https://learn.microsoft.com/en-us/dotnet/standard/asynchronous-programming-patterns/consuming-the-task-based-asynchronous-pattern
|
||||
- `await`, `WhenAll`, `WhenAny`, cancellation propagation, and exception behavior.
|
||||
- Microsoft Learn, "Async return types"
|
||||
- https://learn.microsoft.com/en-us/dotnet/csharp/asynchronous-programming/async-return-types
|
||||
- `Task`, `Task<T>`, `async void`, generalized async return types.
|
||||
- Microsoft Learn, `ValueTask` API reference
|
||||
- https://learn.microsoft.com/en-us/dotnet/api/system.threading.tasks.valuetask
|
||||
- single-consumer warnings and default-to-`Task` guidance.
|
||||
- Microsoft Learn, ASP.NET Core best practices
|
||||
- https://learn.microsoft.com/en-us/aspnet/core/fundamentals/best-practices
|
||||
- avoid blocking calls, avoid unnecessary `Task.Run`, background-work cautions.
|
||||
- Microsoft Learn, hosted services in ASP.NET Core
|
||||
- https://learn.microsoft.com/en-us/aspnet/core/fundamentals/host/hosted-services
|
||||
- safe long-lived background work and cancellation during shutdown.
|
||||
|
||||
## Expert guidance used only when technically grounded
|
||||
|
||||
- Stephen Toub, ".NET Blog: ConfigureAwait FAQ"
|
||||
- https://devblogs.microsoft.com/dotnet/configureawait-faq/
|
||||
- best source for context capture semantics and library-vs-app guidance.
|
||||
- Stephen Toub, ".NET Blog: Understanding the Whys, Whats, and Whens of ValueTask"
|
||||
- https://devblogs.microsoft.com/dotnet/understanding-the-whys-whats-and-whens-of-valuetask/
|
||||
- performance rationale and tradeoffs behind `ValueTask<T>`.
|
||||
- Stephen Toub, ".NET Blog: Await, and UI, and deadlocks! Oh my!"
|
||||
- https://devblogs.microsoft.com/dotnet/await-and-ui-and-deadlocks-oh-my/
|
||||
- canonical deadlock explanation for context-bound code.
|
||||
- Stephen Toub, ".NET Blog: Task Exception Handling in .NET 4.5"
|
||||
- https://devblogs.microsoft.com/dotnet/task-exception-handling-in-net-4-5/
|
||||
- explains `await` versus blocking exception shape and why `WhenAll` matters.
|
||||
- Andrew Arnott, "Recommended patterns for CancellationToken"
|
||||
- https://devblogs.microsoft.com/premier-developer/recommended-patterns-for-cancellationtoken/
|
||||
- practical cancellation design heuristics; useful, but not treated as a language/runtime spec.
|
||||
- Stephen Cleary, "Async/Await - Best Practices in Asynchronous Programming"
|
||||
- https://learn.microsoft.com/en-us/archive/msdn-magazine/2013/march/async-await-best-practices-in-asynchronous-programming
|
||||
- useful design interpretation, but older and treated as contextual guidance rather than current official policy.
|
||||
|
||||
## Where the skill is intentionally cautious
|
||||
|
||||
- `ConfigureAwait`: strong guidance exists for libraries, weaker guidance for app code. Blanket rules are rejected.
|
||||
- `Task.Run`: valid for deliberate CPU offload, weak as a server-side patch for blocking I/O.
|
||||
- `ValueTask`: supported and useful, but easy to misuse. The skill defaults to `Task` unless evidence is present.
|
||||
- Fire-and-forget: acceptable only with explicit ownership and lifecycle design, especially in server code.
|
||||
342
.skills/dotnet-performance-profiling-and-optimization/SKILL.md
Normal file
342
.skills/dotnet-performance-profiling-and-optimization/SKILL.md
Normal file
|
|
@ -0,0 +1,342 @@
|
|||
---
|
||||
name: dotnet-performance-profiling-and-optimization
|
||||
description: >-
|
||||
Use when a .NET process is slow, hung, memory-heavy, or deadlocked, or when
|
||||
analyzing C#/ASP.NET Core code for performance anti-patterns across memory,
|
||||
async, LINQ, database, JSON, caching, DI, concurrency, HttpClient, exceptions,
|
||||
response, strings, startup, and metrics.
|
||||
metadata:
|
||||
category: technique
|
||||
triggers:
|
||||
- dotnet-counters
|
||||
- dotnet-trace
|
||||
- dotnet-dump
|
||||
- dotnet-gcdump
|
||||
- dotnet-stack
|
||||
- benchmarkdotnet
|
||||
- allocations
|
||||
- gc pressure
|
||||
- memory leak
|
||||
- hot path
|
||||
- linq
|
||||
- stackalloc
|
||||
- span
|
||||
- boxing
|
||||
- heap
|
||||
- latency
|
||||
- throughput
|
||||
- slow
|
||||
- hang
|
||||
- deadlock
|
||||
- optimize
|
||||
- performance
|
||||
- async
|
||||
- caching
|
||||
- di lifetime
|
||||
- ef core
|
||||
- cosmosdb
|
||||
- json serialization
|
||||
- httpclient
|
||||
- middleware
|
||||
version: 1.1.0
|
||||
platform: ".NET 8 and .NET 10 (no .NET 9 projects in scope)"
|
||||
---
|
||||
|
||||
# .NET Performance: Diagnostic & Code Review
|
||||
|
||||
Unified C#/.NET performance skill targeting **.NET 8 and .NET 10**. Two modes: live process diagnostics (Mode A) and static code optimization review with fixes (Mode B).
|
||||
|
||||
## Step 0 — Detect the target framework
|
||||
|
||||
Before recommending APIs, follow `../../references/detect-target-framework.md`. Many .NET 9+ APIs (`HybridCache`, `MemoryExtensions.Split` for spans, `Dictionary.GetAlternateLookup`, `params ReadOnlySpan<T>`) do **not** exist on .NET 8 — the references below mark the floor for each pattern, and you must downgrade to the .NET 8 fallback when the target is `net8.0`. Stephen Toub's posts are the primary benchmark source:
|
||||
[Performance Improvements in .NET 8](https://devblogs.microsoft.com/dotnet/performance-improvements-in-net-8/) ·
|
||||
[Performance Improvements in .NET 10](https://devblogs.microsoft.com/dotnet/performance-improvements-in-net-10/).
|
||||
|
||||
## References
|
||||
|
||||
Load on demand:
|
||||
- [references/memory-model-gc.md](references/memory-model-gc.md) — stack vs heap, generations, LOH, boxing, GC tuning
|
||||
- [references/code-patterns.md](references/code-patterns.md) — LINQ, Span/Memory, stackalloc, structs, boxing, pooling, strings (conceptual framing)
|
||||
- [references/categories.md](references/categories.md) — 14 optimization category definitions with checks and grep patterns (Mode B spine)
|
||||
- [references/grep-patterns.md](references/grep-patterns.md) — consolidated anti-pattern grep library for Phase 1 scanning
|
||||
- [references/measurement-guide.md](references/measurement-guide.md) — BenchmarkDotNet, k6, dotnet-counters, KPI targets, CI/CD
|
||||
|
||||
Pattern catalogs (with measured impact numbers, ❌/✅ pairs, and per-topic Detection recipes):
|
||||
- [references/critical-patterns.md](references/critical-patterns.md) — 17 🔴 patterns: deadlocks, order-of-magnitude regressions, excessive allocations
|
||||
- [references/async-patterns.md](references/async-patterns.md) — sync-over-async, ValueTask hot paths, Channels, false sharing
|
||||
- [references/memory-and-strings.md](references/memory-and-strings.md) — `u8` literals, `Span.Split`/`TryWrite`, compound `+=`, chained `.Replace()`
|
||||
- [references/collections-and-linq.md](references/collections-and-linq.md) — `FrozenDictionary`, `GetAlternateLookup`, `CollectionsMarshal.GetValueRefOrAddDefault`, hoisting static data
|
||||
- [references/regex-patterns.md](references/regex-patterns.md) — `[GeneratedRegex]`, `IsMatch`, `EnumerateMatches`, `NonBacktracking`
|
||||
- [references/io-and-serialization.md](references/io-and-serialization.md) — `HttpCompletionOption.ResponseHeadersRead`, `useAsync` `FileStream`, `Memory<byte>` overloads
|
||||
- [references/structural-patterns.md](references/structural-patterns.md) — sealed-class devirtualization (absence pattern, scale-based severity)
|
||||
|
||||
Reference loading guide for Mode B by signal:
|
||||
|
||||
| Signal in Code | Load |
|
||||
|---|---|
|
||||
| `async`, `await`, `Task`, `ValueTask` | `async-patterns.md` |
|
||||
| `Span<`, `Memory<`, `stackalloc`, `string.Substring`, `+=` in loops, `params` | `memory-and-strings.md` |
|
||||
| `Regex`, `[GeneratedRegex]`, `Regex.Match`, `RegexOptions.Compiled` | `regex-patterns.md` |
|
||||
| `Dictionary<`, `List<`, `.ToList()`, LINQ chains, `static readonly Dictionary<` | `collections-and-linq.md` |
|
||||
| `JsonSerializer`, `HttpClient`, `Stream`, `FileStream` | `io-and-serialization.md` |
|
||||
| Any code review on a hot path | always check `critical-patterns.md` first |
|
||||
| Codebase-wide scans (sealed classes, static `Dictionary` → `FrozenDictionary`) | `structural-patterns.md` |
|
||||
|
||||
## Iron Rule
|
||||
|
||||
**Always measure first, change second, and re-measure third.**
|
||||
|
||||
Never claim an optimization without before/after evidence from the same scenario.
|
||||
|
||||
| Rationalization | Reality |
|
||||
|---|---|
|
||||
| "This is obviously slow" | The runtime, JIT, and libraries often invalidate intuition. |
|
||||
| "struct means stack" | Value types are stored inline — not always on the stack. |
|
||||
| "All LINQ is slow" | .NET 9+ improved many LINQ paths. Measure before rewriting. |
|
||||
| "GC.Collect will fix it" | Forced collection treats symptoms, not cause. |
|
||||
| "Too small to matter" | MEDIUM+ impact is cumulative across the request pipeline. |
|
||||
| "I'll change the DI lifetime while I'm here" | DI lifetime changes require explicit user approval. |
|
||||
| "Need to refactor to optimize" | Optimization fixes must be surgical. Refactoring is a separate task. |
|
||||
| "Tests pass so fix is correct" | Tests passing = behavior preserved. Still verify the metric improved. |
|
||||
|
||||
## Mode Selection
|
||||
|
||||
| Situation | Mode |
|
||||
|---|---|
|
||||
| Live process: slow, high CPU/memory, hung, deadlocked, GC pauses | **A – Diagnostic** |
|
||||
| Asking how GC, heap, boxing, or LINQ overhead works in .NET | **A – Diagnostic** (conceptual) |
|
||||
| Code to analyze for anti-patterns, then fix | **B – Code Review** |
|
||||
| Both a running process AND code to fix | Start with **A**, then **B** on hot paths identified |
|
||||
|
||||
**Not for:** Visual Studio, Rider, PerfView, speedscope, or GUI-first workflows.
|
||||
|
||||
---
|
||||
|
||||
## Mode A: Diagnostic (Live Process)
|
||||
|
||||
### Investigation Order
|
||||
|
||||
1. **`dotnet-counters`** — always start here for live triage.
|
||||
2. **`dotnet-stack`** — immediately if process is stuck, hung, or deadlocked.
|
||||
3. **`dotnet-trace`** — if CPU or allocation hot paths matter.
|
||||
4. **`dotnet-gcdump`** — if heap growth matters more than call paths.
|
||||
5. **`dotnet-dump`** — if SOS heap inspection or postmortem analysis is needed.
|
||||
6. After live evidence identifies a candidate routine, apply patterns from [references/code-patterns.md](references/code-patterns.md) and [references/memory-model-gc.md](references/memory-model-gc.md).
|
||||
7. Use BenchmarkDotNet if the change is isolated and needs microbenchmark comparison.
|
||||
8. Re-run the original live capture to prove the real workload improved.
|
||||
|
||||
### CLI Tool Selection
|
||||
|
||||
| Question | Tool | What it answers |
|
||||
|---|---|---|
|
||||
| Is the process allocating, GCing, or saturating CPU? | `dotnet-counters` | Live counters and trend direction |
|
||||
| Is the process hung or deadlocked right now? | `dotnet-stack` | Current managed stack snapshot |
|
||||
| Which call paths consume CPU or allocate heavily? | `dotnet-trace` | Sampled execution and runtime events |
|
||||
| Which object types dominate managed heap? | `dotnet-gcdump` | Heap composition and type totals |
|
||||
| Need SOS heap inspection or thread state? | `dotnet-dump` | Full dump plus CLI analysis |
|
||||
| Did a code change improve an isolated routine? | BenchmarkDotNet | Reproducible microbenchmark comparison |
|
||||
|
||||
### Minimal CLI Commands
|
||||
|
||||
```bash
|
||||
dotnet-counters monitor -p <PID> --counters System.Runtime
|
||||
dotnet-counters monitor -n <ProcessName> --counters System.Runtime,Microsoft.AspNetCore.Hosting
|
||||
dotnet-stack report -p <PID>
|
||||
dotnet-trace collect -p <PID> --duration 00:00:30
|
||||
dotnet-trace report <trace.nettrace> topN
|
||||
dotnet-gcdump collect -p <PID>
|
||||
dotnet-gcdump report <file.gcdump>
|
||||
dotnet-dump collect -p <PID> --type Heap
|
||||
dotnet-dump analyze <dump> -c "dumpheap -stat" -c "exit"
|
||||
```
|
||||
|
||||
Minimal BenchmarkDotNet pattern:
|
||||
```csharp
|
||||
[MemoryDiagnoser]
|
||||
[SimpleJob(RuntimeMoniker.Net90)]
|
||||
public class CandidateBench
|
||||
{
|
||||
[Benchmark(Baseline = true)]
|
||||
public int Original() => OriginalImpl();
|
||||
|
||||
[Benchmark]
|
||||
public int Candidate() => CandidateImpl();
|
||||
}
|
||||
```
|
||||
```bash
|
||||
dotnet run -c Release
|
||||
```
|
||||
|
||||
### Reference Loading Guide
|
||||
|
||||
| User question | Load first |
|
||||
|---|---|
|
||||
| "How do stack and heap really work in .NET?" | `references/memory-model-gc.md` |
|
||||
| "Why is GC pausing or why is LOH churn hurting us?" | `references/memory-model-gc.md` |
|
||||
| "How should I optimize this LINQ?" | `references/code-patterns.md` |
|
||||
| "Can I move this to the stack with stackalloc or Span?" | `references/code-patterns.md` |
|
||||
| "Should this be a struct, ref struct, readonly struct, or class?" | Both |
|
||||
| "Why is this boxing?" | `references/code-patterns.md` |
|
||||
|
||||
### Diagnostic Output
|
||||
|
||||
Report: measured symptom + evidence (counter values, trace hotspots, heap stats) · chosen tool and why · relevant tradeoff (allocation vs copy, deferred vs eager, stack vs pool) · before/after result, or explicitly state if still unverified.
|
||||
|
||||
---
|
||||
|
||||
## Mode B: Code Review (Static Analysis)
|
||||
|
||||
### Target
|
||||
|
||||
`$ARGUMENTS` is the optimization target:
|
||||
- **File path**: Analyze that file and its close dependencies.
|
||||
- **Directory**: Analyze all C# files in that directory.
|
||||
- **"all"**: Scan the solution with Grep, deep-dive the worst offenders.
|
||||
- **`--fix`** anywhere: Skip confirmation and apply fixes after analysis.
|
||||
- **Empty**: Check `git diff --name-only HEAD~5 -- '*.cs'` for recently changed files. If none, ask the user.
|
||||
|
||||
### Phase 1: Discovery
|
||||
|
||||
1. **Glob** to find `.cs` files matching the target.
|
||||
2. **Read** file contents. For files under 500 lines, read the whole file first — visual inspection catches patterns faster than grep, then grep confirms counts.
|
||||
3. **Detect signals** in the code (async, Span, Regex, Dictionary, JsonSerializer, etc.) and **load matching pattern catalogs** from the per-topic references listed at the top of this file.
|
||||
4. **Grep** for anti-patterns. Run the recipes in [references/grep-patterns.md](references/grep-patterns.md) plus the per-topic Detection sections in the catalogs you loaded.
|
||||
5. **Emit a scan execution checklist** before classifying — list each recipe and the hit count. **0 hits is valid and valuable** (confirms good practice).
|
||||
|
||||
### Phase 2: Analysis (Read-Only)
|
||||
|
||||
Check each file against all 14 categories. Record per finding: **file path, line number, current pattern, recommended pattern, impact level, category**.
|
||||
|
||||
Read [references/categories.md](references/categories.md) for detailed check definitions.
|
||||
|
||||
#### Compound Allocation Check
|
||||
|
||||
Single-line grep recipes miss multi-allocation patterns. After running scan recipes, look for:
|
||||
|
||||
1. **Branched `.Replace()` chains** — methods that call `.Replace()` across multiple `if/else` branches. Report total allocation count across all branches, not just per-line.
|
||||
2. **Cross-method chaining** — public method A calls B (which does 3 regex replaces) then calls C (which allocates). Report the total chain cost as one finding, not per-method.
|
||||
3. **Compound `+=` with embedded allocating calls** — `result += $"...{Foo().ToLower()}"` is 2+ allocations (interpolation + `ToLower` + concatenation). Flag the compound cost, not just `.ToLower()`.
|
||||
4. **`string.Format` specificity** — distinguish resource-loaded format strings (not fixable) from compile-time literal format strings (fixable with interpolation). Enumerate only the actionable sites.
|
||||
|
||||
#### Cross-File Consistency Check
|
||||
|
||||
If an optimized pattern is found in one file, check whether sibling files (same directory, same interface, same base class) use the un-optimized equivalent. Flag as MEDIUM with the optimized file as evidence.
|
||||
|
||||
#### Verify-the-Inverse Rule
|
||||
|
||||
For absence patterns (e.g., unsealed classes, static `Dictionary` not converted to `FrozenDictionary`, `RegexOptions.Compiled` not migrated to `[GeneratedRegex]`), always count both sides and report the **N-of-M ratio**, not just the count of bad cases. The ratio determines severity:
|
||||
|
||||
- 0/185 sealed → systematic codebase-wide issue
|
||||
- 12/15 sealed → consistency fix on the remaining 3
|
||||
- 50/100 sealed → mid-migration; flag the laggards
|
||||
|
||||
| # | Category | Code | Focus |
|
||||
|---|---|---|---|
|
||||
| 1 | Memory Allocation | MEM | Span, ArrayPool, pooling, stackalloc, string optimization, collections |
|
||||
| 2 | Async Anti-Patterns | ASYNC | Blocking, ValueTask, CancellationToken, IAsyncEnumerable, Channel |
|
||||
| 3 | LINQ Inefficiencies | LINQ | Count vs Any, multiple enumeration, filter/project order |
|
||||
| 4 | Database | DB | EF Core, CosmosDB patterns, N+1, partition keys, RU cost |
|
||||
| 5 | JSON Serialization | JSON | Options reuse, source generators, serializer boundaries |
|
||||
| 6 | Caching | CACHE | HybridCache, stampede protection, output cache, size limits |
|
||||
| 7 | DI Lifetimes | DI | Captive dependencies, lifetime mismatches, IOptions patterns |
|
||||
| 8 | Concurrency | CONC | Lock contention, throttling, thread safety, Channel patterns |
|
||||
| 9 | HttpClient | HTTP | IHttpClientFactory, resilience, response disposal |
|
||||
| 10 | Exception Control Flow | EXC | Try/catch for expected paths, broad catches |
|
||||
| 11 | Response Optimization | RESP | Compression, pagination, ETags |
|
||||
| 12 | String Optimization | STR | Concatenation loops, ToLower/ToUpper, String.Format |
|
||||
| 13 | Startup & Pipeline | STARTUP | Middleware ordering, compression, health checks, PGO |
|
||||
| 14 | Metrics & Observability | METRICS | IMeterFactory, histograms, tag cardinality, OpenTelemetry |
|
||||
|
||||
### Phase 3: Report
|
||||
|
||||
```
|
||||
## Performance Analysis Report
|
||||
|
||||
### Summary
|
||||
- Files analyzed: N
|
||||
- Total findings: N
|
||||
- Critical (HIGH): N | Moderate (MEDIUM): N | Minor (LOW): N
|
||||
|
||||
### Findings by Category
|
||||
|
||||
#### [CATEGORY_NAME] (N findings)
|
||||
|
||||
| # | Impact | File:Line | Issue | Recommendation |
|
||||
|---|--------|-----------|-------|----------------|
|
||||
| 1 | HIGH | `path/File.cs:42` | Current anti-pattern | Recommended fix |
|
||||
|
||||
### Prioritized Action List
|
||||
1. [HIGH] Fix blocking async calls in X — thread pool starvation risk
|
||||
2. [MEDIUM] Switch to ArrayPool in Z — reduces GC pressure on upload path
|
||||
```
|
||||
|
||||
**Impact levels:**
|
||||
- **HIGH**: Measurable gain, prevents starvation, fixes correctness, reduces P95 latency. Examples: blocking async, missing CancellationToken, N+1 queries, captive dependencies.
|
||||
- **MEDIUM**: Reduces allocations, GC pressure, or unnecessary work. Examples: ArrayPool, StringBuilder, FrozenDictionary.
|
||||
- **LOW**: Minor improvements, cold-path optimizations. Examples: initial collection capacity, Count() vs Any().
|
||||
|
||||
**Scale-based severity escalation.** When the same anti-pattern appears across many instances, escalate:
|
||||
|
||||
- 1–10 instances → report at the pattern's base severity
|
||||
- 11–50 instances → escalate LOW patterns to MEDIUM
|
||||
- 50+ instances → MEDIUM with elevated priority; flag as a codebase-wide systematic issue
|
||||
|
||||
Always report **exact counts from scan recipes**, not estimates. Group findings by severity (HIGH → MEDIUM → LOW), not by file. Merge related findings that share the same fix (e.g., all `.ToLower()` calls in one finding, not split per file).
|
||||
|
||||
### Phase 4: Optimization (Apply Fixes)
|
||||
|
||||
After presenting the report:
|
||||
- If `--fix` in `$ARGUMENTS`, proceed directly.
|
||||
- Otherwise ask: "Would you like me to apply these optimizations? I'll work one category at a time, starting with HIGH impact. You can specify categories or findings (e.g., 'fix ASYNC and MEM' or 'fix #1, #3')."
|
||||
|
||||
**Before any fix:**
|
||||
1. **Read actual code context** around the grep match — false positives exist (`.Result` in `Task.FromResult` is NOT blocking).
|
||||
2. Confirm the finding is real. If uncertain, flag as "needs manual review."
|
||||
|
||||
**Applying fixes:**
|
||||
1. One category at a time, highest impact first. Use Edit tool with brief before/after summary.
|
||||
2. After each category: `dotnet build --no-restore`
|
||||
3. After all changes: `dotnet test`
|
||||
4. If build or tests fail, diagnose before continuing.
|
||||
|
||||
---
|
||||
|
||||
## Analyzer Radar
|
||||
|
||||
- `CA1826`, `CA1827`, `CA1829`, `CA1836`, `CA1851`, `CA1860` — LINQ and enumeration
|
||||
- `CA1845`, `CA1846`, `CA1858` — string and span-friendly APIs
|
||||
- `CA1834`, `CA1865`–`CA1867` — StringBuilder char overloads
|
||||
- `CA1870` — cached `SearchValues<T>`
|
||||
|
||||
These are clues, not goals. Apply where measured hot paths justify it.
|
||||
|
||||
## Pattern Guardrails
|
||||
|
||||
- Do not say "put it on the stack" as a blanket goal. Explain lifetime, copies, boxing, and escape rules.
|
||||
- Do not suggest `stackalloc` for unbounded sizes, large buffers, or loop-carried allocations.
|
||||
- Do not recommend `Span<T>` for data that crosses `await`, escapes to the heap, or lives in object fields — use `Memory<T>`.
|
||||
- Do not recommend converting every `class` to a `struct` — large, mutable, or frequently boxed types often get worse.
|
||||
- Do not blanket-rewrite LINQ to loops — use analyzer-backed fixes first.
|
||||
- Do not recommend pooling without ownership rules — returned pooled arrays must not be reused by the caller.
|
||||
- Do not recommend `GC.Collect()` except for rare justified lifecycle boundaries with measurement.
|
||||
|
||||
## Red Flags — STOP and Confirm
|
||||
|
||||
Stop and ask before:
|
||||
- Changing `Program.cs` or the middleware pipeline
|
||||
- Adding a new NuGet package
|
||||
- Changing any DI service lifetime registration
|
||||
- Replacing the serializer in the HTTP pipeline
|
||||
- Modifying API response shapes or route patterns
|
||||
- Changing error handling patterns
|
||||
|
||||
## Constraints
|
||||
|
||||
- NEVER add NuGet packages without user approval
|
||||
- NEVER change DI lifetimes without explaining implications and getting confirmation
|
||||
- NEVER modify `Program.cs` or middleware pipeline without explicit approval
|
||||
- NEVER change API contracts, route patterns, or response shapes
|
||||
- ALWAYS preserve existing tests; update only if behavior intentionally changes
|
||||
- ALWAYS use the Grep tool for searches, never bash `grep` or `find`
|
||||
|
||||
Consult the project's CLAUDE.md or AGENTS.md for project-specific rules and constraints.
|
||||
|
|
@ -0,0 +1,112 @@
|
|||
# Async & Concurrency Patterns
|
||||
|
||||
### Don't Expose Async Wrappers for Sync Methods
|
||||
🟡 **AVOID** wrapping sync methods with `Task.Run` in libraries | .NET Core+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
public Task<int> ComputeHashAsync(byte[] data) =>
|
||||
Task.Run(() => ComputeHash(data));
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
public int ComputeHash(byte[] data) { /* CPU-bound work */ }
|
||||
// Consumer decides: var hash = await Task.Run(() => lib.ComputeHash(data));
|
||||
```
|
||||
|
||||
**Impact: Eliminates unnecessary thread pool queue/dequeue overhead per call.**
|
||||
|
||||
### Don't Expose Sync Wrappers for Async Methods
|
||||
🟡 **AVOID** creating sync wrappers that block on async implementations | .NET Core+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
public string GetData() => GetDataAsync().Result;
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
public async Task<string> GetDataAsync() { /* ... */ }
|
||||
```
|
||||
|
||||
**Impact: Prevents deadlocks and thread pool starvation from hidden sync-over-async blocking.**
|
||||
|
||||
### Use ValueTask for Hot Paths with Frequent Sync Completion
|
||||
🟡 **DO** use `ValueTask<T>` on hot paths where sync completion is common | .NET Core 2.1+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
public async Task<int> ReadAsync(Memory<byte> buffer)
|
||||
{
|
||||
if (_bufferedCount > 0)
|
||||
return ReadFromBuffer(buffer.Span);
|
||||
return await ReadAsyncCore(buffer);
|
||||
}
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
public ValueTask<int> ReadAsync(Memory<byte> buffer)
|
||||
{
|
||||
if (_bufferedCount > 0)
|
||||
return new ValueTask<int>(ReadFromBuffer(buffer.Span));
|
||||
return new ValueTask<int>(ReadAsyncCore(buffer));
|
||||
}
|
||||
```
|
||||
|
||||
**Impact: Eliminates Task\<T\> allocation on synchronous completion — the struct stores results inline.**
|
||||
|
||||
### Use Channels for Producer/Consumer
|
||||
🟡 **DO** use `System.Threading.Channels` for producer-consumer patterns | .NET Core 3.0+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
var queue = new BlockingCollection<WorkItem>();
|
||||
var item = queue.Take();
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
var channel = Channel.CreateUnbounded<WorkItem>();
|
||||
|
||||
// Producer
|
||||
await channel.Writer.WriteAsync(item);
|
||||
|
||||
// Consumer
|
||||
await foreach (var item in channel.Reader.ReadAllAsync())
|
||||
Process(item);
|
||||
```
|
||||
|
||||
**Impact: ~25% faster, ~95% fewer GC collections vs manual approaches.**
|
||||
|
||||
### Avoid False Sharing with Thread-Local State
|
||||
🟡 **AVOID** adjacent mutable fields written by different threads | .NET 7+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
class SharedCounters
|
||||
{
|
||||
public long Counter1;
|
||||
public long Counter2;
|
||||
}
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
[StructLayout(LayoutKind.Explicit, Size = 128)]
|
||||
struct PaddedCounter
|
||||
{
|
||||
[FieldOffset(0)] public long Value;
|
||||
}
|
||||
```
|
||||
|
||||
**Impact: Eliminates cross-core cache invalidation — can improve multi-threaded throughput by 10x+.**
|
||||
|
||||
## Detection
|
||||
|
||||
Scan recipes for async anti-patterns. Run these and report exact counts.
|
||||
|
||||
```bash
|
||||
# async void methods (correctness issue — crashes on exception)
|
||||
grep -rn --include='*.cs' 'async void' --exclude-dir=bin --exclude-dir=obj . | grep -v 'event' | wc -l
|
||||
```
|
||||
|
||||
### Patterns Requiring Manual Review
|
||||
|
||||
- **Sync-over-async** (`.Result`, `.Wait()`): `.Result` matches any property named Result — needs type context to confirm it's `Task.Result`
|
||||
|
|
@ -0,0 +1,297 @@
|
|||
# Optimization Category Definitions
|
||||
|
||||
Complete check definitions for all 14 optimization categories. Each category lists specific checks to perform, with grep patterns at the end for automated scanning.
|
||||
|
||||
---
|
||||
|
||||
## Category 1: Memory Allocation (MEM)
|
||||
|
||||
Checks:
|
||||
|
||||
- **Unnecessary string allocations**: `Substring()` calls that could use `Span<char>` or `AsSpan()`. String concatenation with `+` inside loops (should use `StringBuilder` or `string.Create`).
|
||||
- **Missing ArrayPool/MemoryPool usage**: `new byte[...]` for temporary buffers, especially in I/O paths. Should use `ArrayPool<byte>.Shared.Rent()` with try/finally Return.
|
||||
- **Large Object Heap triggers**: Allocations of objects >= 85,000 bytes (arrays, large strings, `MemoryStream` without `RecyclableMemoryStream`).
|
||||
- **Missing object pooling**: Frequently created/disposed objects (like `StringBuilder`) that could use `ObjectPool<T>`.
|
||||
- **Record class vs record struct**: Small, immutable DTOs that are `record class` but could be `readonly record struct` to avoid heap allocation.
|
||||
- **Boxing**: Value types cast to `object` or non-generic interfaces. Structs without `IEquatable<T>`.
|
||||
- **Collection inefficiencies**: `new List<T>()` or `new Dictionary<K,V>()` without initial capacity when size is known or estimable. Double-lookup patterns (`TryGetValue` + indexer set) that could use `CollectionsMarshal.GetValueRefOrAddDefault`. Read-only dictionaries populated once that could be `FrozenDictionary<K,V>` (.NET 8+).
|
||||
- **stackalloc for small buffers**: Flag `new byte[N]` where N <= 256 in synchronous methods. Recommend `Span<byte> buffer = stackalloc byte[N]` for short-lived stack allocation with zero GC pressure.
|
||||
- **string.Create for pre-sized construction**: When output string length is known at call time, `string.Create(length, state, action)` avoids intermediate allocations by writing directly into the final buffer.
|
||||
- **Interpolated strings in logging** *(Impact: LOW — only matters when the log level is inactive at runtime)*: `_logger.LogXxx($"...")` allocates the interpolated string even when the log level is disabled. Use structured logging parameters `_logger.LogXxx("Message {Param}", value)` or the `[LoggerMessage]` source generator for high-frequency hot paths.
|
||||
- **ReadOnlySpan for string parsing**: Flag `.Split()` and `.Substring()` in hot paths where `AsSpan()` slicing avoids allocation. Common in string parsing, normalization, and URL handling.
|
||||
- **CollectionsMarshal.GetValueRefOrAddDefault**: Flag the TryGetValue + indexer set double-lookup pattern. Single-lookup alternative reduces dictionary operations by 50%.
|
||||
- **RecyclableMemoryStream**: Flag `new MemoryStream()` in I/O-heavy paths (blob upload/download, log writes). `Microsoft.IO.RecyclableMemoryStream` pools internal buffers and avoids LOH fragmentation.
|
||||
|
||||
Grep patterns:
|
||||
```
|
||||
\.Substring\(
|
||||
new byte\[
|
||||
new MemoryStream\(\)
|
||||
new StringBuilder\(\)
|
||||
new List<.*>\(\)
|
||||
new Dictionary<.*>\(\)
|
||||
_logger\.Log(Debug|Trace|Information|Warning|Error|Critical)\(\$"
|
||||
\.Split\(
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Category 2: Async Anti-Patterns (ASYNC)
|
||||
|
||||
Checks:
|
||||
|
||||
- **Blocking async calls**: `.Result`, `.Wait()`, `.GetAwaiter().GetResult()` -- causes thread pool starvation. CRITICAL finding.
|
||||
- **async void**: Methods declared `async void` (except event handlers) -- swallows exceptions and cannot be awaited.
|
||||
- **Missing CancellationToken**: Async methods that do not accept or propagate `CancellationToken`. Every async method should have `CancellationToken cancellationToken = default`.
|
||||
- **Missing ValueTask**: Methods that frequently return cached/synchronous results but use `Task<T>` instead of `ValueTask<T>`. Look for `if (cache.TryGetValue(...)) return Task.FromResult(...)`. Benchmark data: ValueTask 7.41ns/0B vs Task 15.23ns/72B.
|
||||
- **Sequential awaits that could parallelize**: Multiple independent `await` calls in sequence that could use `Task.WhenAll`.
|
||||
- **Async over sync**: Methods that use `Task.Run` to wrap synchronous code in an ASP.NET Core context (unnecessary and wastes a thread).
|
||||
- **ConfigureAwait(false) in library projects**: LOW/INFO level. Not required in ASP.NET Core (no sync context), but recommended if library assemblies may be reused outside ASP.NET Core.
|
||||
- **IAsyncEnumerable opportunities**: Methods returning `Task<List<T>>` where the caller iterates sequentially. If the caller processes items one-by-one, `IAsyncEnumerable<T>` reduces memory and improves time-to-first-byte.
|
||||
- **Channel verification**: Verify `BoundedChannelOptions` has `SingleReader`/`SingleWriter` hints set for performance. Verify `CancellationToken` is propagated on `WriteAsync` and `ReadAllAsync`.
|
||||
|
||||
Grep patterns:
|
||||
```
|
||||
\.Result[^s]
|
||||
\.Wait\(\)
|
||||
\.GetAwaiter\(\)\.GetResult\(\)
|
||||
async void
|
||||
Task\.Run\(
|
||||
\.WriteAsync\([^,]*\)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Category 3: LINQ Inefficiencies (LINQ)
|
||||
|
||||
Checks:
|
||||
|
||||
- **Count() > 0 or Count() == 0**: Should use `Any()` or `!Any()`. `Count()` may enumerate the entire collection.
|
||||
- **Multiple enumeration**: An `IEnumerable<T>` variable used more than once without materializing.
|
||||
- **Filter after projection**: `.Select(...).Where(...)` -- should filter first, then project.
|
||||
- **ToList() too early**: `.ToList().Where(...)` or `.ToList().Select(...)` -- materializes before filtering.
|
||||
- **OrderBy before Where**: Sorting the full collection before filtering it down.
|
||||
|
||||
Grep patterns:
|
||||
```
|
||||
\.Count\(\) [><=!]
|
||||
\.ToList\(\)\.Where\(
|
||||
\.ToList\(\)\.Select\(
|
||||
\.Select\(.*\)\.Where\(
|
||||
\.OrderBy.*\.Where\(
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Category 4: Database (DB)
|
||||
|
||||
Check for both EF Core and CosmosDB patterns depending on what the project uses. Consult the project's CLAUDE.md or AGENTS.md for the data access strategy.
|
||||
|
||||
Checks:
|
||||
|
||||
- **N+1 queries**: Loops that call the database inside each iteration.
|
||||
- **Missing AsNoTracking**: EF Core read-only queries without `.AsNoTracking()`.
|
||||
- **Missing compiled queries**: Frequently executed EF Core queries on hot paths without `EF.CompileAsyncQuery`.
|
||||
- **Full entity loading**: Fetching entire entities when only a few fields are needed (should project to DTOs).
|
||||
- **Missing AsSplitQuery**: EF Core queries with multiple `.Include()` calls without `.AsSplitQuery()`.
|
||||
- **CosmosDB partition key misuse**: Operations not specifying the partition key, or using cross-partition queries unnecessarily.
|
||||
- **CosmosDB point reads**: Using queries instead of `ReadItemAsync` when both `id` and partition key are known.
|
||||
- **RU cost awareness**: Flag discarded `ItemResponse<T>` without logging `RequestCharge`. Recommend tracking RU cost via metrics for cost visibility.
|
||||
- **Cross-partition query detection**: `GetItemLinqQueryable()` without partition key option leads to fan-out queries. Verify all LINQ queryables specify the partition key.
|
||||
- **Indexing policy review**: Flag if queries filter on fields that likely lack composite indexes.
|
||||
- **Redundant round-trips**: Flag patterns where a query fetches an ID, then a separate point read fetches the full document. Recommend a single query.
|
||||
- **EnableContentResponseOnWrite = false**: On write operations where the response body is not needed, setting this option reduces RU cost.
|
||||
|
||||
Grep patterns:
|
||||
```
|
||||
\.Include\(.*\.Include\(
|
||||
await.*foreach.*await.*Async
|
||||
ReadItemAsync
|
||||
GetItemQueryIterator
|
||||
GetItemLinqQueryable
|
||||
\.RequestCharge
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Category 5: JSON Serialization (JSON)
|
||||
|
||||
Checks:
|
||||
|
||||
- **New JsonSerializerOptions per call**: `new JsonSerializerOptions { ... }` inside method bodies -- rebuilds the metadata cache every time. Should use a static readonly instance.
|
||||
- **Missing source generators**: High-throughput serialization paths without `[JsonSerializable]` source generation context.
|
||||
- **Newtonsoft.Json in hot paths**: If the project uses Newtonsoft.Json for MVC, flag any hot-path internal serialization that could benefit from `System.Text.Json` with source generators. NEVER suggest replacing the controller/DTO serializer without checking the project's documented constraints.
|
||||
- **System.Text.Json source generators for internal serialization**: Internal paths (database serialization, audit logs, blob metadata) that don't affect API contracts are candidates for `System.Text.Json` with source generators.
|
||||
- **JsonSerializerSettings singleton**: Flag `new JsonSerializerSettings()` in method bodies. The contract resolver cache is rebuilt each time. Use a static readonly instance or inject via DI.
|
||||
- **CosmosDB SDK serializer**: The Cosmos SDK supports custom serializers. `CosmosSystemTextJsonSerializer` with source generators reduces allocation on read/write operations.
|
||||
|
||||
Grep patterns:
|
||||
```
|
||||
new JsonSerializerOptions
|
||||
JsonConvert\.Serialize
|
||||
JsonConvert\.Deserialize
|
||||
new JsonSerializer
|
||||
new JsonSerializerSettings
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Category 6: Caching (CACHE)
|
||||
|
||||
Checks:
|
||||
|
||||
- **Repeated expensive calls without caching**: Service methods that call external APIs on every request without caching the result.
|
||||
- **Missing HybridCache pattern**: Look for manual cache-aside patterns that could use `HybridCache` for built-in stampede protection, two-level caching (L1 memory + L2 distributed), and tag invalidation. **HybridCache requires .NET 10 (or .NET 9) — the `Microsoft.Extensions.Caching.Hybrid` package does not target .NET 8.** On .NET 8 implement `IMemoryCache` (L1) + `IDistributedCache` (L2) manually with a `SemaphoreSlim` keyed by cache key for stampede protection. See [HybridCache GA announcement](https://devblogs.microsoft.com/dotnet/hybrid-cache-is-now-ga/).
|
||||
- **Static data fetched repeatedly**: Configuration, taxonomies, or lookup data fetched from external APIs that rarely changes.
|
||||
- **Missing output caching**: Read-only GET endpoints that return the same data for all callers -- candidates for `[OutputCache]`.
|
||||
- **HybridCache upgrade path (.NET 10 only)**: Manual L1+L2 caching with `SemaphoreSlim` stampede protection can migrate to `HybridCache` `GetOrCreateAsync` with built-in stampede protection and tag invalidation **when the project target is `net10.0`**. On `net8.0` the manual pattern is the correct end state, not a stepping stone.
|
||||
- **Cache stampede detection**: Cache-aside without locking -- `TryGetValue` followed by expensive call followed by `Set` without `SemaphoreSlim` or equivalent stampede guard.
|
||||
- **Tag-based invalidation with RemoveByTagAsync**: When using HybridCache, group related entries by tag for efficient bulk invalidation instead of tracking individual keys.
|
||||
- **IMemoryCache size limits**: Flag `AddMemoryCache()` without `SizeLimit` in `MemoryCacheOptions`. Unbounded in-memory cache can grow until the process runs out of memory.
|
||||
|
||||
Grep patterns:
|
||||
```
|
||||
GetAsync\(
|
||||
SendAsync\(
|
||||
_cache\.TryGetValue
|
||||
DistributedCache
|
||||
AddMemoryCache\(\)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Category 7: DI Lifetime Issues (DI)
|
||||
|
||||
Consult the project's CLAUDE.md or AGENTS.md for the expected DI lifetime registrations.
|
||||
|
||||
Checks:
|
||||
|
||||
- **Transient services that should be Singleton**: Stateless, thread-safe services registered as Transient that have no per-request state (could be Singleton for zero allocation).
|
||||
- **Scoped injected into Singleton**: A Scoped service captured in a Singleton constructor -- captive dependency bug.
|
||||
- **IOptions vs IOptionsMonitor vs IOptionsSnapshot**: `IOptions<T>` in Singleton services that need to react to config changes should use `IOptionsMonitor<T>`. `IOptionsSnapshot<T>` in Singleton is a captive dependency.
|
||||
|
||||
---
|
||||
|
||||
## Category 8: Concurrency Issues (CONC)
|
||||
|
||||
Checks:
|
||||
|
||||
- **Lock contention**: `lock` statements that guard async operations (should use `SemaphoreSlim`).
|
||||
- **Missing throttling**: Unbounded parallel calls to external APIs without `SemaphoreSlim` or concurrency limits.
|
||||
- **Thread-unsafe patterns**: Shared mutable state without synchronization. `HttpContext` accessed from background threads.
|
||||
- **Channel usage patterns**: If the project uses `Channel<T>` for background tasks, verify `SingleReader`/`SingleWriter` hints are set correctly for performance.
|
||||
|
||||
Grep patterns:
|
||||
```
|
||||
lock\s*\(
|
||||
new SemaphoreSlim
|
||||
HttpContext.*Task\.Run
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Category 9: HttpClient Misuse (HTTP)
|
||||
|
||||
Checks:
|
||||
|
||||
- **new HttpClient()**: Direct instantiation instead of `IHttpClientFactory`. Causes socket exhaustion and DNS caching issues.
|
||||
- **Missing resilience**: HTTP calls without retry/circuit-breaker policies. Verify `Microsoft.Extensions.Http.Resilience` or Polly is applied to external API clients.
|
||||
- **Missing response disposal**: `HttpResponseMessage` not disposed after reading.
|
||||
|
||||
Grep patterns:
|
||||
```
|
||||
new HttpClient\(
|
||||
new HttpClient\b
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Category 10: Exception-Driven Control Flow (EXC)
|
||||
|
||||
Checks:
|
||||
|
||||
- **Try/catch for expected paths**: Using exceptions for normal control flow (e.g., catching `KeyNotFoundException` instead of `TryGetValue`, catching `FormatException` instead of `TryParse`).
|
||||
- **Broad catch blocks**: `catch (Exception)` that swallow errors or use exceptions as branching logic.
|
||||
- **Exception allocation in hot paths**: Throwing exceptions on paths that execute frequently.
|
||||
|
||||
Grep patterns:
|
||||
```
|
||||
catch\s*\(Exception\b
|
||||
catch\s*\(KeyNotFoundException
|
||||
catch\s*\(FormatException
|
||||
catch\s*\(InvalidOperationException
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Category 11: Response Optimization (RESP)
|
||||
|
||||
Checks:
|
||||
|
||||
- **Missing compression**: No response compression middleware, or JSON responses served uncompressed.
|
||||
- **Missing pagination**: Endpoints returning unbounded collections.
|
||||
- **Missing ETags for conditional requests**: GET endpoints without ETag support where the data has a natural version (e.g., database ETags or row versions).
|
||||
|
||||
---
|
||||
|
||||
## Category 12: String Optimization (STR)
|
||||
|
||||
Checks:
|
||||
|
||||
- **String concatenation in loops**: `+=` on strings inside `for`/`foreach`/`while` loops.
|
||||
- **String.Format in hot paths**: Could use interpolated string handlers or `StringBuilder`.
|
||||
- **Repeated string operations**: Multiple `ToLower()`/`ToUpper()` calls on the same value. Should use `StringComparison.OrdinalIgnoreCase` instead.
|
||||
|
||||
Grep patterns:
|
||||
```
|
||||
\+= "
|
||||
\+= \$"
|
||||
\.ToLower\(\)
|
||||
\.ToUpper\(\)
|
||||
String\.Format\(
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Category 13: Startup & Pipeline Optimization (STARTUP)
|
||||
|
||||
Checks:
|
||||
|
||||
- **Middleware ordering**: Verify Program.cs follows the recommended sequence: ExceptionHandler, ResponseCompression, OutputCache, Routing, RateLimiter, CORS, Authentication, Authorization, MapControllers. Incorrect ordering degrades performance (e.g., compression after routing skips static responses).
|
||||
- **Response compression**: Flag missing `AddResponseCompression`/`UseResponseCompression`. Without it, all JSON responses are uncompressed. Recommend Brotli (optimal ratio) + GZip (compatibility) providers with `EnableForHttps = true`.
|
||||
- **Health check optimization**: `.ShortCircuit()` (.NET 8+) bypasses the entire middleware pipeline for health endpoints. `.DisableHttpMetrics()` prevents health check traffic from skewing request duration metrics.
|
||||
- **PGO/ReadyToRun**: Check .csproj for `<TieredPGO>true</TieredPGO>` (dynamic PGO for runtime hot-path optimization) and `<PublishReadyToRun>true</PublishReadyToRun>` (pre-compiled code for faster startup). Both should be present for production builds.
|
||||
- **Warm-up pattern**: `ApplicationStarted` callback to warm expensive singletons (database connections, cache, external API health). Cold-start latency without warm-up can spike P99 for the first requests after deployment.
|
||||
|
||||
Grep patterns:
|
||||
```
|
||||
UseResponseCompression
|
||||
AddResponseCompression
|
||||
ShortCircuit
|
||||
AddOutputCache
|
||||
UseOutputCache
|
||||
TieredPGO
|
||||
PublishReadyToRun
|
||||
ApplicationStarted
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Category 14: Metrics & Observability (METRICS)
|
||||
|
||||
Checks:
|
||||
|
||||
- **IMeterFactory vs static new Meter()**: Services should inject `IMeterFactory` from DI rather than using `new Meter(...)`. `IMeterFactory` enables testability with `MetricCollector<T>` and proper meter lifecycle management.
|
||||
- **Missing histograms**: If the project only has `Counter<long>` instruments, recommend `Histogram<double>` for request processing duration, external API latency, and database query duration. Histograms enable percentile analysis (P50/P95/P99).
|
||||
- **Tag cardinality**: Verify metric tags have bounded cardinality. NEVER use request IDs, user IDs, or unbounded strings as metric tags. Tags like `operation_name`, `status_code`, `endpoint` are acceptable (bounded). Unbounded tags cause metric explosion and memory issues.
|
||||
- **OpenTelemetry AddMeter() registration**: Verify custom meter names are registered with `.AddMeter("YourMeterName")` in the OpenTelemetry metrics configuration. Without this, custom counters are silently dropped.
|
||||
|
||||
Grep patterns:
|
||||
```
|
||||
new Meter\(
|
||||
CreateCounter
|
||||
CreateHistogram
|
||||
AddMeter
|
||||
\.Record\(
|
||||
\.Add\(
|
||||
```
|
||||
|
|
@ -0,0 +1,384 @@
|
|||
---
|
||||
description: Code-level performance patterns for csharp-dotnet-cli-optimization.
|
||||
metadata:
|
||||
tags: [linq, span, stackalloc, boxing, pooling, strings, analyzers]
|
||||
---
|
||||
|
||||
# Code Patterns
|
||||
|
||||
> **See also:** for pattern-by-pattern detection recipes with measured impact numbers and ❌/✅ pairs, see the topic catalogs: [critical-patterns.md](critical-patterns.md), [async-patterns.md](async-patterns.md), [memory-and-strings.md](memory-and-strings.md), [collections-and-linq.md](collections-and-linq.md), [regex-patterns.md](regex-patterns.md), [io-and-serialization.md](io-and-serialization.md), [structural-patterns.md](structural-patterns.md). This file covers the conceptual framing (when/why), those files cover the catalog (what/how-much).
|
||||
|
||||
Use this reference after a profile or benchmark identifies a hot path. Do not apply these patterns speculatively.
|
||||
|
||||
## Table Of Contents
|
||||
|
||||
- [LINQ And Enumeration](#linq-and-enumeration)
|
||||
- [Stack Allocation, Span, And Memory](#stack-allocation-span-and-memory)
|
||||
- [Structs, Boxing, And Copies](#structs-boxing-and-copies)
|
||||
- [Buffer Reuse And Advanced Helpers](#buffer-reuse-and-advanced-helpers)
|
||||
- [Strings](#strings)
|
||||
- [What Not To Suggest](#what-not-to-suggest)
|
||||
|
||||
## LINQ And Enumeration
|
||||
|
||||
### Property or indexer over LINQ when the concrete collection is known
|
||||
|
||||
Wrong:
|
||||
|
||||
```csharp
|
||||
if (items.Count() > 0)
|
||||
{
|
||||
return items.First();
|
||||
}
|
||||
```
|
||||
|
||||
Better:
|
||||
|
||||
```csharp
|
||||
if (items.Count > 0)
|
||||
{
|
||||
return items[0];
|
||||
}
|
||||
```
|
||||
|
||||
Use `Count`, `Length`, `IsEmpty`, or an indexer when you already have a concrete collection with that API. Relevant analyzers: `CA1826`, `CA1829`, `CA1836`, `CA1860`.
|
||||
|
||||
### `Any()` over `Count() > 0` when all you know is `IEnumerable<T>`
|
||||
|
||||
Wrong:
|
||||
|
||||
```csharp
|
||||
if (source.Count() != 0)
|
||||
{
|
||||
Process(source);
|
||||
}
|
||||
```
|
||||
|
||||
Better:
|
||||
|
||||
```csharp
|
||||
if (source.Any())
|
||||
{
|
||||
Process(source);
|
||||
}
|
||||
```
|
||||
|
||||
Relevant analyzer: `CA1827`.
|
||||
|
||||
### Avoid multiple enumeration of deferred queries
|
||||
|
||||
Wrong:
|
||||
|
||||
```csharp
|
||||
var query = source.Where(Filter);
|
||||
return query.Count() + query.Last().Id;
|
||||
```
|
||||
|
||||
Better:
|
||||
|
||||
```csharp
|
||||
var materialized = source.Where(Filter).ToArray();
|
||||
return materialized.Length + materialized[^1].Id;
|
||||
```
|
||||
|
||||
Materialize once only when you truly need multiple passes or random access and can afford the extra memory. Relevant analyzer: `CA1851`.
|
||||
|
||||
### Avoid premature materialization
|
||||
|
||||
Wrong:
|
||||
|
||||
```csharp
|
||||
var projected = source.ToList().Select(Map);
|
||||
```
|
||||
|
||||
Better:
|
||||
|
||||
```csharp
|
||||
var projected = source.Select(Map);
|
||||
```
|
||||
|
||||
Keep deferred execution unless you need a snapshot, repeated traversal, indexing, or a boundary between expensive stages.
|
||||
|
||||
### Do not blanket-rewrite LINQ to loops
|
||||
|
||||
- .NET 10 improved many LINQ operations substantially through JIT array-interface devirtualisation — operations like `Skip`/`Take`/`Sum` on arrays and `ReadOnlyCollection<T>` got roughly 50% faster *for free*. See Stephen Toub, [Performance Improvements in .NET 10](https://devblogs.microsoft.com/dotnet/performance-improvements-in-net-10/).
|
||||
- On `net8.0` LINQ has the historical performance characteristics described in [Performance Improvements in .NET 8](https://devblogs.microsoft.com/dotnet/performance-improvements-in-net-8/) — fast paths exist for `Count`, `ToList`, `ToArray` on `ICollection<T>`, but indexed access via `ElementAt`/`Skip`/`Take` is not as cheap as on .NET 10.
|
||||
- Start with analyzer-backed fixes and measurement.
|
||||
- Replace LINQ with hand-written loops only when a benchmark or trace shows that the remaining cost matters — on either target.
|
||||
|
||||
### Use `TryGetNonEnumeratedCount` when count is optional
|
||||
|
||||
```csharp
|
||||
if (source.TryGetNonEnumeratedCount(out int count))
|
||||
{
|
||||
LogCount(count);
|
||||
}
|
||||
```
|
||||
|
||||
This avoids forcing enumeration when the underlying type already knows its size.
|
||||
|
||||
## Stack Allocation, Span, And Memory
|
||||
|
||||
### `stackalloc` only for small, bounded, temporary buffers
|
||||
|
||||
Wrong:
|
||||
|
||||
```csharp
|
||||
for (int i = 0; i < items.Length; i++)
|
||||
{
|
||||
Span<byte> buffer = stackalloc byte[4096];
|
||||
Use(buffer);
|
||||
}
|
||||
```
|
||||
|
||||
Better:
|
||||
|
||||
```csharp
|
||||
Span<byte> buffer = stackalloc byte[256];
|
||||
for (int i = 0; i < items.Length; i++)
|
||||
{
|
||||
buffer.Clear();
|
||||
Use(buffer);
|
||||
}
|
||||
```
|
||||
|
||||
Guidance:
|
||||
|
||||
- keep sizes conservative
|
||||
- avoid `stackalloc` inside loops
|
||||
- initialize the memory before use
|
||||
- fall back to heap or pooling for larger or variable-sized buffers
|
||||
|
||||
### Prefer span-based APIs over substring copies
|
||||
|
||||
Wrong:
|
||||
|
||||
```csharp
|
||||
int.TryParse(line.Substring(7), out int value);
|
||||
```
|
||||
|
||||
Better:
|
||||
|
||||
```csharp
|
||||
int.TryParse(line.AsSpan(7), out int value);
|
||||
```
|
||||
|
||||
Relevant analyzers: `CA1845`, `CA1846`.
|
||||
|
||||
### Use `Span<T>` for sync work and `Memory<T>` for async or heap-stored state
|
||||
|
||||
Wrong:
|
||||
|
||||
```csharp
|
||||
// Wrong: Span<T> cannot cross await safely.
|
||||
public async Task<int> ReadAsync(Span<byte> buffer)
|
||||
{
|
||||
await socket.ReceiveAsync(buffer);
|
||||
return buffer[0];
|
||||
}
|
||||
```
|
||||
|
||||
Better:
|
||||
|
||||
```csharp
|
||||
public async Task<int> ReadAsync(Memory<byte> buffer)
|
||||
{
|
||||
await socket.ReceiveAsync(buffer);
|
||||
return buffer.Span[0];
|
||||
}
|
||||
```
|
||||
|
||||
`Span<T>` is stack-only. If the lifetime crosses `await`, callbacks, or object storage, move to `Memory<T>`.
|
||||
|
||||
### `ref struct` is for stack-bound wrappers, not a general optimization badge
|
||||
|
||||
- Use `ref struct` when the type itself contains spans or must never escape to the heap.
|
||||
- Do not use it if you need arrays of that type, boxing, interface conversions, or heap fields.
|
||||
|
||||
## Structs, Boxing, And Copies
|
||||
|
||||
### Use `readonly struct` or `readonly record struct` for small immutable values
|
||||
|
||||
Wrong:
|
||||
|
||||
```csharp
|
||||
public struct Measurement
|
||||
{
|
||||
public double A;
|
||||
public double B;
|
||||
public void Normalize() => A /= B;
|
||||
}
|
||||
```
|
||||
|
||||
Better:
|
||||
|
||||
```csharp
|
||||
public readonly record struct Measurement(double A, double B);
|
||||
```
|
||||
|
||||
Prefer value types for small, copyable, data-only values. Avoid large, mutable structs.
|
||||
|
||||
### Pass large structs by `in`
|
||||
|
||||
Wrong:
|
||||
|
||||
```csharp
|
||||
double Distance(Vector4 value) => value.X + value.Y + value.Z + value.W;
|
||||
```
|
||||
|
||||
Better:
|
||||
|
||||
```csharp
|
||||
double Distance(in Vector4 value) => value.X + value.Y + value.Z + value.W;
|
||||
```
|
||||
|
||||
This avoids copying large struct values on each call.
|
||||
|
||||
### Avoid boxing in hot paths
|
||||
|
||||
Wrong:
|
||||
|
||||
```csharp
|
||||
object boxed = valueStruct;
|
||||
```
|
||||
|
||||
Wrong:
|
||||
|
||||
```csharp
|
||||
IFormattable f = valueStruct;
|
||||
```
|
||||
|
||||
Better:
|
||||
|
||||
```csharp
|
||||
Use(in valueStruct);
|
||||
```
|
||||
|
||||
Boxing allocates a heap object and copies the value. Interface conversions can box too.
|
||||
|
||||
### Mark readonly members on structs
|
||||
|
||||
- Non-readonly instance members on a readonly receiver can trigger defensive copies.
|
||||
- Mark the whole struct `readonly` when possible, or mark readonly members explicitly.
|
||||
|
||||
## Buffer Reuse And Advanced Helpers
|
||||
|
||||
### Use `ArrayPool<T>` when the buffer is too large or variable for `stackalloc`
|
||||
|
||||
Wrong:
|
||||
|
||||
```csharp
|
||||
byte[] temp = new byte[inputLength];
|
||||
```
|
||||
|
||||
Better:
|
||||
|
||||
```csharp
|
||||
byte[] temp = ArrayPool<byte>.Shared.Rent(inputLength);
|
||||
try
|
||||
{
|
||||
Use(temp);
|
||||
}
|
||||
finally
|
||||
{
|
||||
ArrayPool<byte>.Shared.Return(temp);
|
||||
}
|
||||
```
|
||||
|
||||
Rules:
|
||||
|
||||
- return to the same pool once
|
||||
- never use the buffer after return
|
||||
- rented arrays may be larger than requested
|
||||
- rented arrays are not guaranteed to be zeroed
|
||||
|
||||
### Prevent accidental closure capture
|
||||
|
||||
Wrong:
|
||||
|
||||
```csharp
|
||||
return values.Select(v => v * 2).ToArray();
|
||||
```
|
||||
|
||||
Better when no capture is needed:
|
||||
|
||||
```csharp
|
||||
return values.Select(static v => v * 2).ToArray();
|
||||
```
|
||||
|
||||
Use `static` lambdas or static local functions to prevent capture when the delegate does not need outer state.
|
||||
|
||||
### Cache `SearchValues<T>` for repeated searches
|
||||
|
||||
Wrong:
|
||||
|
||||
```csharp
|
||||
int index = text.IndexOfAny(":/?&=".AsSpan());
|
||||
```
|
||||
|
||||
Better:
|
||||
|
||||
```csharp
|
||||
private static readonly SearchValues<char> s_delims =
|
||||
SearchValues.Create(":/?&=".AsSpan());
|
||||
```
|
||||
|
||||
```csharp
|
||||
int index = text.IndexOfAny(s_delims);
|
||||
```
|
||||
|
||||
Relevant analyzer: `CA1870`.
|
||||
|
||||
### `CollectionsMarshal.AsSpan` is advanced and ownership-sensitive
|
||||
|
||||
```csharp
|
||||
Span<int> span = CollectionsMarshal.AsSpan(list);
|
||||
```
|
||||
|
||||
Use this only when:
|
||||
|
||||
- you own the `List<T>`
|
||||
- you will not add or remove items while the span is in use
|
||||
- a measured hot path justifies bypassing normal list APIs
|
||||
|
||||
## Strings
|
||||
|
||||
### `StartsWith` over `IndexOf(...) == 0`
|
||||
|
||||
Wrong:
|
||||
|
||||
```csharp
|
||||
return text.IndexOf("abc", StringComparison.Ordinal) == 0;
|
||||
```
|
||||
|
||||
Better:
|
||||
|
||||
```csharp
|
||||
return text.StartsWith("abc", StringComparison.Ordinal);
|
||||
```
|
||||
|
||||
Relevant analyzer: `CA1858`.
|
||||
|
||||
### `Append(char)` over `Append("x")`
|
||||
|
||||
Wrong:
|
||||
|
||||
```csharp
|
||||
builder.Append("]");
|
||||
```
|
||||
|
||||
Better:
|
||||
|
||||
```csharp
|
||||
builder.Append(']');
|
||||
```
|
||||
|
||||
Relevant analyzers: `CA1834`, `CA1865-CA1867`.
|
||||
|
||||
## What Not To Suggest
|
||||
|
||||
- Do not suggest unsafe code first.
|
||||
- Do not suggest pooling tiny objects by default.
|
||||
- Do not suggest `stackalloc` because "heap bad, stack good".
|
||||
- Do not suggest converting APIs to `Span<T>` if the lifetime model does not fit.
|
||||
- Do not suggest loop rewrites without a profile or benchmark showing LINQ still matters after simpler fixes.
|
||||
|
|
@ -0,0 +1,217 @@
|
|||
# Collections & LINQ Patterns
|
||||
|
||||
### Use FrozenDictionary/FrozenSet for Read-Heavy Lookup Tables
|
||||
🟡 **DO** use `FrozenDictionary`/`FrozenSet` for collections created once and read many times | .NET 8+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
private static readonly Dictionary<string, int> s_statusCodes = new()
|
||||
{
|
||||
["OK"] = 200, ["NotFound"] = 404, ["InternalServerError"] = 500
|
||||
};
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
private static readonly FrozenDictionary<string, int> s_statusCodes =
|
||||
new Dictionary<string, int>
|
||||
{
|
||||
["OK"] = 200, ["NotFound"] = 404, ["InternalServerError"] = 500
|
||||
}.ToFrozenDictionary();
|
||||
```
|
||||
|
||||
**Impact: ~50% faster lookups than Dictionary, ~14x faster than ImmutableDictionary.**
|
||||
|
||||
### Use Dictionary Alternate Lookup for Span-Based Keys
|
||||
🟡 **DO** use `GetAlternateLookup<ReadOnlySpan<char>>()` to avoid string allocation on lookups | **.NET 10 (or .NET 9) only — NOT available on .NET 8**
|
||||
|
||||
❌ (allocates on every lookup; the only option on .NET 8)
|
||||
```csharp
|
||||
string key = headerLine.Substring(0, colonIndex);
|
||||
if (s_dict.TryGetValue(key, out int value)) { /* ... */ }
|
||||
```
|
||||
✅ .NET 10 / C# 14
|
||||
```csharp
|
||||
var lookup = s_dict.GetAlternateLookup<ReadOnlySpan<char>>();
|
||||
ReadOnlySpan<char> key = headerLine.AsSpan(0, colonIndex);
|
||||
if (lookup.TryGetValue(key, out int value)) { }
|
||||
```
|
||||
✅ .NET 8 fallback — keep the allocation but minimise it
|
||||
```csharp
|
||||
// On net8.0 GetAlternateLookup does not exist (added in .NET 9 BCL).
|
||||
// Pre-intern frequent keys, or accept the allocation. If the hot path is
|
||||
// truly critical, store keys as ReadOnlyMemory<char> and write a custom
|
||||
// IEqualityComparer<string> that compares against a span via string.Compare.
|
||||
string key = headerLine.Substring(0, colonIndex);
|
||||
if (s_dict.TryGetValue(key, out int value)) { /* ... */ }
|
||||
```
|
||||
|
||||
**Impact: Avoids string allocation per lookup on .NET 10 — especially valuable in parser/protocol hot paths.**
|
||||
|
||||
### Use CollectionsMarshal.GetValueRefOrNullRef for Lookup-and-Update
|
||||
🟡 **DO** use `CollectionsMarshal.GetValueRefOrAddDefault` for dictionary update patterns | .NET 6+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
_counts.TryGetValue(key, out int count);
|
||||
_counts[key] = count + 1;
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
ref int count = ref CollectionsMarshal.GetValueRefOrAddDefault(_counts, key, out _);
|
||||
count++;
|
||||
```
|
||||
|
||||
**Impact: ~48% faster for lookup-and-update patterns (95µs → 49µs).**
|
||||
|
||||
### Use Collection Expressions [] for Zero-Allocation Span Creation
|
||||
🟡 **DO** use collection expressions for `Span<T>` targets | C# 12 / .NET 8+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
int[] values = new int[] { a, b, c, d };
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
Span<int> values = [a, b, c, d];
|
||||
ReadOnlySpan<int> daysInMonth = [31, 28, 31, 30, 31, 30, 31, 31, 30, 31, 30, 31];
|
||||
```
|
||||
|
||||
**Impact: Zero heap allocation for span-targeted collection expressions.**
|
||||
|
||||
### Use EnsureCapacity on List/Stack/Queue Before Bulk Adds
|
||||
🟡 **DO** call `EnsureCapacity` before bulk insertions | .NET 6+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
var list = new List<int>();
|
||||
for (int i = 0; i < 10000; i++)
|
||||
list.Add(i);
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
var list = new List<int>();
|
||||
list.EnsureCapacity(10000);
|
||||
for (int i = 0; i < 10000; i++)
|
||||
list.Add(i);
|
||||
```
|
||||
|
||||
**Impact: Reduces reallocations and array copies during bulk operations.**
|
||||
|
||||
### Use TryGetNonEnumeratedCount for Pre-Sizing
|
||||
🟡 **DO** use `TryGetNonEnumeratedCount` to pre-size destination collections | .NET 6+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
var results = new List<int>();
|
||||
foreach (var item in source)
|
||||
results.Add(Transform(item));
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
var results = source.TryGetNonEnumeratedCount(out int count)
|
||||
? new List<int>(count)
|
||||
: new List<int>();
|
||||
foreach (var item in source)
|
||||
results.Add(Transform(item));
|
||||
```
|
||||
|
||||
**Impact: Avoids O(n) enumeration for counting; eliminates resizing allocations.**
|
||||
|
||||
### Hoist Static Data Out of Method Bodies
|
||||
🟡 **AVOID** creating collections with static/deterministic data inside method bodies | .NET Core+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
public string Convert(long number)
|
||||
{
|
||||
var groupsMap = new Dictionary<long, Func<long, string>>
|
||||
{
|
||||
{ 1_000_000_000, n => $"{Convert(n)} billion" },
|
||||
{ 1_000_000, n => $"{Convert(n)} million" },
|
||||
{ 1_000, n => $"{Convert(n)} thousand" },
|
||||
};
|
||||
}
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
private static readonly FrozenDictionary<long, Func<long, string>> s_groupsMap =
|
||||
new Dictionary<long, Func<long, string>>
|
||||
{
|
||||
{ 1_000_000_000, n => $"{Convert(n)} billion" },
|
||||
{ 1_000_000, n => $"{Convert(n)} million" },
|
||||
{ 1_000, n => $"{Convert(n)} thousand" },
|
||||
}.ToFrozenDictionary();
|
||||
|
||||
public string Convert(long number)
|
||||
{
|
||||
// ... use s_groupsMap
|
||||
}
|
||||
```
|
||||
|
||||
**Impact: Eliminates collection + internal storage + closure allocations per call. For a Dictionary with N entries, saves ~N+3 allocations per invocation.**
|
||||
|
||||
### Add Overloads to Avoid params Array Allocation
|
||||
🟡 **DO** add 1- and 2-argument overloads for methods that accept `params T[]`. On .NET 10 also expose a `params ReadOnlySpan<T>` overload | works on .NET 8 and .NET 10
|
||||
|
||||
❌ (single `params T[]` overload allocates a new array on every call, including the common 1-argument case)
|
||||
```csharp
|
||||
public static string Transform(this string input, params IStringTransformer[] transformers) =>
|
||||
transformers.Aggregate(input, (current, t) => t.Transform(current));
|
||||
|
||||
"hello".Transform(To.TitleCase);
|
||||
```
|
||||
✅ Option A — explicit overloads for common arities (works on .NET 8 and .NET 10)
|
||||
```csharp
|
||||
public static string Transform(this string input, IStringTransformer transformer) =>
|
||||
transformer.Transform(input);
|
||||
|
||||
public static string Transform(this string input, IStringTransformer t1, IStringTransformer t2) =>
|
||||
t2.Transform(t1.Transform(input));
|
||||
|
||||
public static string Transform(this string input, params IStringTransformer[] transformers) =>
|
||||
transformers.Aggregate(input, (current, t) => t.Transform(current));
|
||||
```
|
||||
✅ Option B — `.NET 10 / C# 14` adds a span overload (eliminates the allocation for all arities)
|
||||
```csharp
|
||||
public static string Transform(this string input, params ReadOnlySpan<IStringTransformer> transformers)
|
||||
{
|
||||
foreach (var t in transformers)
|
||||
input = t.Transform(input);
|
||||
return input;
|
||||
}
|
||||
```
|
||||
⚠️ Option B does **not** compile on `net8.0`: `params ReadOnlySpan<T>` requires C# 13 (default on .NET 9+). On .NET 8 ship only Option A.
|
||||
|
||||
**Impact: Option A eliminates the array allocation for 1- and 2-argument calls on every target. Option B eliminates it for all arities on .NET 10.**
|
||||
|
||||
## Detection
|
||||
|
||||
Scan recipes for collection and LINQ anti-patterns. Run these and report exact counts.
|
||||
|
||||
```bash
|
||||
# Static Dictionary not using FrozenDictionary (read-only after init)
|
||||
grep -rn --include='*.cs' 'static readonly Dictionary<' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
|
||||
# Static FrozenDictionary (already optimized — verify the inverse)
|
||||
grep -rn --include='*.cs' 'static readonly FrozenDictionary<' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
|
||||
# Per-call List allocation (inside method bodies, not static/readonly fields)
|
||||
grep -rn --include='*.cs' 'new List<' --exclude-dir=bin --exclude-dir=obj . | grep -v 'static\|readonly' | wc -l
|
||||
|
||||
# Per-call Dictionary allocation (inside method bodies, not static/readonly fields)
|
||||
grep -rn --include='*.cs' 'new Dictionary<' --exclude-dir=bin --exclude-dir=obj . | grep -v 'static\|readonly' | wc -l
|
||||
|
||||
# StringComparer.CurrentCulture usage (almost always wrong in library code — use Ordinal)
|
||||
grep -rn --include='*.cs' 'StringComparer.CurrentCulture' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
|
||||
# LINQ chains in extension/hot-path files (.Select, .Where, .Cast, .Take, .Aggregate)
|
||||
grep -rn --include='*.cs' -E '\.(Select|Where|Cast|Take|Aggregate)\(' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
```
|
||||
|
||||
For the LINQ chain recipe: any hit in a file whose name ends in `Extensions.cs`, `Formatter.cs`, or implements a method called from a public extension method is a hot-path candidate. Inspect each hit in these files and flag LINQ chains that allocate delegates, enumerators, or intermediate collections on every call. Hits in localization converters or one-time initialization are lower priority.
|
||||
|
||||
### Patterns Requiring Manual Review
|
||||
|
||||
- **ContainsKey + indexer double-lookup**: Requires verifying the same key is used in a subsequent indexer access — multi-line/multi-statement context
|
||||
- **LINQ on hot paths**: The LINQ chain recipe above catches call sites, but distinguishing hot-path from cold-path requires context. Prioritize hits in `*Extensions.cs` and `*Formatter.cs` files, which are typically called on every user invocation
|
||||
- **`new Dictionary/List<` in method bodies vs fields**: The grep heuristic (`grep -v 'static\|readonly'`) catches most cases but may include false positives from field initializers without `static`/`readonly` — spot-check flagged lines
|
||||
|
|
@ -0,0 +1,288 @@
|
|||
# Critical .NET Performance Anti-Patterns
|
||||
|
||||
17 patterns that cause deadlocks, order-of-magnitude regressions, or excessive allocations.
|
||||
|
||||
## Async / Tasks
|
||||
|
||||
### Never Block on Async (Sync-over-Async)
|
||||
🔴 **AVOID** | .NET Core+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
public string GetData()
|
||||
=> GetDataAsync().Result;
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
public async Task<string> GetDataAsync()
|
||||
=> await GetDataInternalAsync();
|
||||
```
|
||||
**Impact: Deadlocks or thread pool starvation; wastes threads, destroys scalability.**
|
||||
|
||||
### Never Await a ValueTask Multiple Times
|
||||
🔴 **AVOID** | .NET Core 2.1+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
ValueTask<int> vt = SomeMethodAsync();
|
||||
int a = await vt;
|
||||
int b = await vt;
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
int result = await SomeMethodAsync();
|
||||
```
|
||||
**Impact: Undefined behavior — silent data corruption or exceptions.**
|
||||
|
||||
## Memory / Allocation
|
||||
|
||||
### Use Span\<T\> / AsSpan Instead of Substring for Slicing
|
||||
🔴 **DO** | .NET Core 2.1+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
string sub = input.Substring(5, 10);
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
ReadOnlySpan<char> sub = input.AsSpan(5, 10);
|
||||
```
|
||||
**Impact: Eliminates per-slice allocations; 2-4x faster via vectorization.**
|
||||
|
||||
### Use ArrayPool\<T\> for Temporary Buffers
|
||||
🔴 **DO** | .NET Core+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
byte[] buf = new byte[4096];
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
byte[] buf = ArrayPool<byte>.Shared.Rent(4096);
|
||||
Process(buf);
|
||||
ArrayPool<byte>.Shared.Return(buf);
|
||||
```
|
||||
**Impact: Dramatically reduces GC pressure for buffer-heavy workloads.**
|
||||
|
||||
### Avoid stackalloc in Loops
|
||||
🔴 **AVOID** | .NET 5+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
for (int i = 0; i < 10_000; i++)
|
||||
Span<byte> buf = stackalloc byte[1024];
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
Span<byte> buf = stackalloc byte[1024];
|
||||
for (int i = 0; i < 10_000; i++) { Process(buf); }
|
||||
```
|
||||
**Impact: StackOverflowException — unrecoverable, no catch possible.**
|
||||
|
||||
### Avoid Boxing Value Types
|
||||
🔴 **AVOID** | .NET 6+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
string s = string.Format("{0}.{1}", major, minor);
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
string s = $"{major}.{minor}";
|
||||
```
|
||||
**Impact: When replacing `string.Format` with C# 10+ interpolation, typical improvements are ~40% faster with significantly less allocation. Actual gains vary by call site.**
|
||||
|
||||
## Strings
|
||||
|
||||
### Use StringComparison.Ordinal for Non-Linguistic Comparisons
|
||||
🔴 **DO** | .NET Core+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
bool found = text.IndexOf("Content-Type") >= 0;
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
bool found = text.Contains("Content-Type", StringComparison.Ordinal);
|
||||
```
|
||||
**Impact: 2-3x faster; OrdinalIgnoreCase hash codes ~3.3x faster.**
|
||||
|
||||
### Use AsSpan Instead of Substring
|
||||
🔴 **DO** | .NET Core 2.1+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
int val = int.Parse(str.Substring(5, 3));
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
int val = int.Parse(str.AsSpan(5, 3));
|
||||
```
|
||||
**Impact: Eliminates one string allocation per parse operation.**
|
||||
|
||||
## Regular Expressions
|
||||
|
||||
### Use Source-Generated Regex [GeneratedRegex]
|
||||
🔴 **ALWAYS** use `[GeneratedRegex]` for all static regex patterns | .NET 7+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
private static readonly Regex s_re =
|
||||
new(@"\w+@\w+\.\w+", RegexOptions.Compiled);
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
[GeneratedRegex(@"\w+@\w+\.\w+")]
|
||||
private static partial Regex EmailRegex();
|
||||
```
|
||||
**Impact: Always beneficial or neutral for static patterns — near-zero startup, better throughput, and required for AOT/trimming scenarios.**
|
||||
|
||||
### Avoid Nested Quantifiers (Catastrophic Backtracking)
|
||||
🔴 **AVOID** | .NET Core+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
var r = new Regex(@"^(\w+)+$");
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
var r = new Regex(@"^\w+$", RegexOptions.NonBacktracking);
|
||||
```
|
||||
**Impact: Can hang process indefinitely on crafted input.**
|
||||
|
||||
### Use TryGetValue Instead of ContainsKey + Indexer
|
||||
🔴 **DO** | .NET Core+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
if (dict.ContainsKey(key))
|
||||
Use(dict[key]);
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
if (dict.TryGetValue(key, out var value))
|
||||
Use(value);
|
||||
```
|
||||
**Impact: ~2x faster (50% reduction in lookup time).**
|
||||
|
||||
### Avoid LINQ in Hot Paths
|
||||
🔴 **AVOID** | .NET Core+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
bool found = items.Any(x => x.Name == target);
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
bool found = false;
|
||||
foreach (var item in items)
|
||||
if (item.Name == target) { found = true; break; }
|
||||
```
|
||||
**Impact: Eliminates 1-3 allocations per call; measurable in tight loops.**
|
||||
|
||||
### Don't Iterate IEnumerable Multiple Times
|
||||
🔴 **AVOID** | .NET Core+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
foreach (Type t in types) { Validate(t); }
|
||||
_types = types.ToArray();
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
Type[] arr = types.ToArray();
|
||||
foreach (Type t in arr) { Validate(t); }
|
||||
_types = arr;
|
||||
```
|
||||
**Impact: Halves enumeration cost; prevents bugs from re-executing deferred queries.**
|
||||
|
||||
## JSON Serialization
|
||||
|
||||
### Use System.Text.Json Source Generator
|
||||
🔴 **DO** | .NET 6+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
string json = JsonSerializer.Serialize(post);
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
[JsonSerializable(typeof(BlogPost))]
|
||||
internal partial class AppJsonCtx : JsonSerializerContext { }
|
||||
string json = JsonSerializer.Serialize(post, AppJsonCtx.Default.BlogPost);
|
||||
```
|
||||
**Impact: 37-44% faster; enables trimming and Native AOT.**
|
||||
|
||||
### Cache JsonSerializerOptions
|
||||
🔴 **DO** | .NET 5+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
JsonSerializer.Serialize(obj, new JsonSerializerOptions());
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
private static readonly JsonSerializerOptions s_opts = new();
|
||||
JsonSerializer.Serialize(obj, s_opts);
|
||||
```
|
||||
**Impact: Up to 592x slower without caching (.NET 6); always cache or use defaults.**
|
||||
|
||||
## Networking
|
||||
|
||||
### Reuse HttpClient Instances
|
||||
🔴 **DO** | .NET Core 2.1+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
using var client = new HttpClient();
|
||||
await client.GetStringAsync(url);
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
private static readonly HttpClient s_http = new(new SocketsHttpHandler
|
||||
{ PooledConnectionLifetime = TimeSpan.FromMinutes(5) });
|
||||
await s_http.GetStringAsync(url);
|
||||
```
|
||||
**Impact: Prevents socket exhaustion; 6-12x faster concurrent HTTPS.**
|
||||
|
||||
## General
|
||||
|
||||
### Use SearchValues\<T\> for Repeated Set Searches
|
||||
🔴 **DO** | .NET 8+ (works on both targets; .NET 10 adds multi-string overloads)
|
||||
|
||||
❌
|
||||
```csharp
|
||||
int pos = text.IndexOfAny("ABCDEF".ToCharArray());
|
||||
```
|
||||
✅ (.NET 8 and .NET 10 — `SearchValues<char>` is the same API on both)
|
||||
```csharp
|
||||
private static readonly SearchValues<char> s_hex = SearchValues.Create("ABCDEF");
|
||||
int pos = text.AsSpan().IndexOfAny(s_hex);
|
||||
```
|
||||
✅ (.NET 10 only — multi-string `SearchValues<string>`)
|
||||
```csharp
|
||||
private static readonly SearchValues<string> s_keywords =
|
||||
SearchValues.Create(["error", "warning", "fatal"], StringComparison.OrdinalIgnoreCase);
|
||||
int pos = log.AsSpan().IndexOfAny(s_keywords); // SearchValues<string> overload is .NET 9+ BCL
|
||||
```
|
||||
On `net8.0` use a `SearchValues<char>` with the first letter of each keyword and then fall back to `string.IndexOf(StringComparison.Ordinal)`.
|
||||
|
||||
**Impact: 2-10x faster for chars (both targets); 10-30x faster for multi-string on .NET 10.**
|
||||
|
||||
## Detection
|
||||
|
||||
Scan recipes for critical anti-patterns. Run these and report exact counts of issues found in each case.
|
||||
|
||||
```bash
|
||||
# .IndexOf(string) without StringComparison (culture-aware, 2-3x slower)
|
||||
grep -rn --include='*.cs' -E '\.IndexOf\("[^"]+"\)' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
|
||||
# .Substring( calls (allocates new string — consider AsSpan)
|
||||
grep -rn --include='*.cs' '\.Substring(' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
|
||||
# .StartsWith/.EndsWith without StringComparison (culture-aware, 2-3x slower)
|
||||
grep -rn --include='*.cs' -E '\.(StartsWith|EndsWith)\("[^"]+"\)' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
|
||||
# .Contains(string) without StringComparison — NOTE: will also match collection .Contains() calls; filter to string receivers
|
||||
grep -rn --include='*.cs' -E '\.Contains\("[^"]+"\)' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
```
|
||||
|
|
@ -0,0 +1,165 @@
|
|||
# Anti-Pattern Grep Library
|
||||
|
||||
Consolidated grep patterns for automated scanning. Run these with the Grep tool using `type: "cs"` filter.
|
||||
|
||||
All patterns use ripgrep regex syntax. Run during Phase 1 (Discovery) broad scan.
|
||||
|
||||
> **See also:** each topic catalog ([critical-patterns.md](critical-patterns.md), [async-patterns.md](async-patterns.md), [memory-and-strings.md](memory-and-strings.md), [collections-and-linq.md](collections-and-linq.md), [regex-patterns.md](regex-patterns.md), [io-and-serialization.md](io-and-serialization.md), [structural-patterns.md](structural-patterns.md)) has its own Detection section with topic-specific recipes and ratio-counting guidance. This file is the consolidated cross-cutting library; load topic files when their signals are present.
|
||||
|
||||
---
|
||||
|
||||
## ASYNC anti-patterns (CRITICAL)
|
||||
|
||||
```
|
||||
\.Result\b
|
||||
\.Wait\(\)
|
||||
\.GetAwaiter\(\)\.GetResult\(\)
|
||||
async void\b
|
||||
Task\.Run\(
|
||||
\.WriteAsync\([^,]*\)
|
||||
```
|
||||
|
||||
**False positive note**: `.Result` matches `Task.FromResult` -- verify actual blocking before flagging.
|
||||
|
||||
---
|
||||
|
||||
## Memory anti-patterns (MEM)
|
||||
|
||||
```
|
||||
new byte\[\d{4,}\]
|
||||
new byte\[
|
||||
new MemoryStream\(\)
|
||||
\.Substring\(
|
||||
new StringBuilder\(\)
|
||||
new List<.*>\(\)
|
||||
new Dictionary<.*>\(\)
|
||||
_logger\.Log(Debug|Trace|Information|Warning|Error|Critical)\(\$"
|
||||
\.Split\(
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## LINQ anti-patterns
|
||||
|
||||
```
|
||||
\.Count\(\)\s*[><=!]
|
||||
\.ToList\(\)\.Where\(
|
||||
\.ToList\(\)\.Select\(
|
||||
\.Select\(.*\)\.Where\(
|
||||
\.OrderBy.*\.Where\(
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Database / CosmosDB anti-patterns (DB)
|
||||
|
||||
```
|
||||
\.Include\(.*\.Include\(
|
||||
await.*foreach.*await.*Async
|
||||
ReadItemAsync
|
||||
GetItemQueryIterator
|
||||
GetItemLinqQueryable
|
||||
\.RequestCharge
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## JSON anti-patterns
|
||||
|
||||
```
|
||||
new JsonSerializerOptions
|
||||
new JsonSerializerSettings
|
||||
JsonConvert\.Serialize
|
||||
JsonConvert\.Deserialize
|
||||
new JsonSerializer
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Caching anti-patterns (CACHE)
|
||||
|
||||
```
|
||||
GetAsync\(
|
||||
SendAsync\(
|
||||
_cache\.TryGetValue
|
||||
DistributedCache
|
||||
AddMemoryCache\(\)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## HttpClient misuse (HTTP)
|
||||
|
||||
```
|
||||
new HttpClient\(
|
||||
new HttpClient\b
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Exception control flow (EXC)
|
||||
|
||||
```
|
||||
catch\s*\(Exception\b
|
||||
catch\s*\(KeyNotFoundException
|
||||
catch\s*\(FormatException
|
||||
catch\s*\(InvalidOperationException
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## String anti-patterns (STR)
|
||||
|
||||
```
|
||||
\+= "
|
||||
\+= \$"
|
||||
\.ToLower\(\)
|
||||
\.ToUpper\(\)
|
||||
String\.Format\(
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Concurrency anti-patterns (CONC)
|
||||
|
||||
```
|
||||
lock\s*\(
|
||||
new SemaphoreSlim
|
||||
HttpContext.*Task\.Run
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Startup & Pipeline (STARTUP)
|
||||
|
||||
```
|
||||
UseResponseCompression
|
||||
AddResponseCompression
|
||||
ShortCircuit
|
||||
AddOutputCache
|
||||
UseOutputCache
|
||||
TieredPGO
|
||||
PublishReadyToRun
|
||||
ApplicationStarted
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Metrics & Observability (METRICS)
|
||||
|
||||
```
|
||||
new Meter\(
|
||||
CreateCounter
|
||||
CreateHistogram
|
||||
AddMeter
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## CancellationToken coverage
|
||||
|
||||
```
|
||||
async Task[<\s].*\)\s*$
|
||||
```
|
||||
|
||||
This pattern finds async methods whose signature ends without a CancellationToken parameter. Verify each match -- some may be interface implementations where the token is propagated differently.
|
||||
|
|
@ -0,0 +1,124 @@
|
|||
# I/O, Serialization & General Patterns
|
||||
|
||||
### Use HttpCompletionOption.ResponseHeadersRead for Streaming
|
||||
🟡 **DO** use `ResponseHeadersRead` when downloading large responses | .NET Core 3.0+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
var response = await client.GetAsync(uri);
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
using var response = await client.GetAsync(uri, HttpCompletionOption.ResponseHeadersRead);
|
||||
using var stream = await response.Content.ReadAsStreamAsync();
|
||||
await stream.CopyToAsync(destinationStream);
|
||||
```
|
||||
|
||||
**Impact: ~2x faster for large downloads (10MB+), dramatically reduced memory usage.**
|
||||
|
||||
### Use Async FileStream Operations
|
||||
🟡 **DO** use `FileStream` with `useAsync: true` for scalable file I/O | .NET 6+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
using var fs = new FileStream(path, FileMode.Open);
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
await using var fs = new FileStream(path, FileMode.Open, FileAccess.Read,
|
||||
FileShare.Read, bufferSize: 4096, useAsync: true);
|
||||
|
||||
byte[] buffer = new byte[1024];
|
||||
while (await fs.ReadAsync(buffer) != 0) { /* process */ }
|
||||
```
|
||||
|
||||
**Impact: Up to 3x faster async reads; allocation reduced from megabytes to hundreds of bytes.**
|
||||
|
||||
### Use Memory\<byte\> Overloads for Stream.ReadAsync/WriteAsync
|
||||
🟡 **DO** use `Memory<byte>`-based stream overloads instead of `byte[]` overloads | .NET 5+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
await stream.ReadAsync(buffer, 0, buffer.Length);
|
||||
await stream.WriteAsync(buffer, 0, buffer.Length);
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
await stream.ReadAsync(buffer.AsMemory());
|
||||
await stream.WriteAsync(buffer.AsMemory());
|
||||
```
|
||||
|
||||
**Impact: Eliminates ~72 KB allocation per 1,000 read/write pairs on NetworkStream.**
|
||||
|
||||
### Use Span-Based TryFormat for Number Formatting
|
||||
🟡 **DO** use `TryFormat` to format numbers into `Span<char>` buffers | .NET Core 2.1+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
string formatted = value.ToString();
|
||||
destination.Write(formatted);
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
Span<char> buffer = stackalloc char[20];
|
||||
if (value.TryFormat(buffer, out int charsWritten))
|
||||
destination.Write(buffer[..charsWritten]);
|
||||
```
|
||||
|
||||
**Impact: Int32.ToString() ~2x faster in .NET Core 2.1, Int32 parsing ~5x faster in .NET Core 3.0.**
|
||||
|
||||
### Use static readonly for Runtime Devirtualization
|
||||
🟡 **DO** store implementations in `static readonly` fields for JIT devirtualization | .NET Core 3.0+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
private static Base s_impl = new DerivedImpl();
|
||||
s_impl.Process();
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
private static readonly Base s_impl = new DerivedImpl();
|
||||
s_impl.Process();
|
||||
|
||||
private static readonly bool s_feature =
|
||||
Environment.GetEnvironmentVariable("Feature") == "1";
|
||||
```
|
||||
|
||||
**Impact: Virtual call eliminated entirely — can be inlined to zero overhead. Dead code elimination in tier 1.**
|
||||
|
||||
### Avoid Explicit Static Constructors — Use Field Initializers
|
||||
🟡 **AVOID** explicit `static` constructors when field initializers suffice | .NET Core 3.0+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
class Foo
|
||||
{
|
||||
static readonly int s_value;
|
||||
static Foo() { s_value = ComputeValue(); }
|
||||
}
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
class Foo
|
||||
{
|
||||
static readonly int s_value = ComputeValue();
|
||||
}
|
||||
```
|
||||
|
||||
**Impact: Enables better JIT optimization and reduces potential lock overhead on static method access.**
|
||||
|
||||
## Detection
|
||||
|
||||
Scan recipes for I/O and serialization anti-patterns. Run these and report exact counts.
|
||||
|
||||
```bash
|
||||
# new HttpClient() (socket exhaustion risk)
|
||||
grep -rn --include='*.cs' 'new HttpClient(' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
|
||||
# new JsonSerializerOptions() not cached (592x slower in .NET 6)
|
||||
grep -rn --include='*.cs' 'new JsonSerializerOptions' --exclude-dir=bin --exclude-dir=obj . | grep -v 'static\|readonly' | wc -l
|
||||
```
|
||||
|
||||
### Patterns Requiring Manual Review
|
||||
|
||||
- **`JsonSerializer.Serialize/Deserialize` without source-gen context**: Can't determine from grep if a context parameter is passed
|
||||
|
|
@ -0,0 +1,193 @@
|
|||
---
|
||||
description: >-
|
||||
Performance measurement guide for dotnet-performance skill. Covers tool
|
||||
selection per category, KPI targets, BenchmarkDotNet, k6 load testing,
|
||||
CI/CD integration, and live process CLI commands.
|
||||
metadata:
|
||||
tags: [measurement, benchmarkdotnet, k6, dotnet-counters, kpi]
|
||||
---
|
||||
|
||||
# Measurement Guide
|
||||
|
||||
How to measure performance before and after applying optimizations. Never optimize without baseline data.
|
||||
|
||||
---
|
||||
|
||||
## Tool Selection Decision Table
|
||||
|
||||
Map each optimization category to the appropriate measurement tools:
|
||||
|
||||
| Category | Primary Tool | Secondary Tool | What to Measure |
|
||||
|---|---|---|---|
|
||||
| MEM | `dotnet-counters` (gc-heap-size, alloc-rate) | `dotnet-gcdump` comparison | Allocation rate reduction, GC collection frequency |
|
||||
| ASYNC | `dotnet-counters` (threadpool-queue-length, thread-count) | App Insights dependency tracking | Thread pool starvation, blocked threads |
|
||||
| LINQ | BenchmarkDotNet `[MemoryDiagnoser]` | `dotnet-trace` hot path | Allocation per operation, throughput |
|
||||
| DB | `response.RequestCharge` logging | App Insights DB dependency | RU cost per operation, query latency |
|
||||
| JSON | BenchmarkDotNet serialization benchmark | `dotnet-counters` alloc-rate | Throughput (ops/sec), bytes allocated |
|
||||
| CACHE | App Insights dependency duration | Custom hit ratio counter | Cache hit rate, dependency call reduction |
|
||||
| DI | `dotnet-counters` alloc-rate | Load test comparison | Object creation overhead |
|
||||
| CONC | `dotnet-counters` (monitor-lock-contention-count) | `dotnet-trace` contention events | Lock wait time, throughput under load |
|
||||
| HTTP | `dotnet-counters` Microsoft.AspNetCore.Hosting | k6/NBomber load test | Request duration, throughput |
|
||||
| EXC | `dotnet-counters` exception-count | App Insights exceptions | Exception rate per interval |
|
||||
| RESP | Network tab / curl with timing | k6 response size check | Response size (bytes), transfer time |
|
||||
| STR | BenchmarkDotNet `[MemoryDiagnoser]` | `dotnet-counters` alloc-rate | String allocations per operation |
|
||||
| STARTUP | Startup time measurement | `dotnet-trace` startup events | Time to first request, cold start latency |
|
||||
| METRICS | `MetricCollector<T>` in tests | Prometheus/Grafana dashboard | Metric emission, cardinality |
|
||||
|
||||
---
|
||||
|
||||
## KPI Targets
|
||||
|
||||
Standard targets for ASP.NET Core APIs. Use as thresholds when evaluating optimization impact:
|
||||
|
||||
| Metric | Target | Red Flag |
|
||||
|---|---|---|
|
||||
| P50 response time | < 100ms | > 200ms |
|
||||
| P95 response time | < 500ms | > 1000ms |
|
||||
| P99 response time | < 1000ms | > 2000ms |
|
||||
| Error rate (5xx) | < 0.1% | > 1% |
|
||||
| CPU utilization | < 70% sustained | > 85% |
|
||||
| Memory working set | < 80% | > 90% |
|
||||
| Thread pool queue length | < 10 sustained | > 50 |
|
||||
| GC time percentage | < 10% | > 20% |
|
||||
| Allocation rate | Trend down after optimization | Sustained increase |
|
||||
|
||||
---
|
||||
|
||||
## BenchmarkDotNet Guidance
|
||||
|
||||
Use for micro-optimizations on hot paths (MEM, LINQ, JSON, STR categories).
|
||||
|
||||
**When to benchmark**: Hot-path changes where the difference is in nanoseconds or bytes allocated. Not needed for architectural changes (caching, DI lifetime) — use load testing instead.
|
||||
|
||||
**Minimum setup**:
|
||||
```csharp
|
||||
[MemoryDiagnoser]
|
||||
[SimpleJob(RuntimeMoniker.Net90)]
|
||||
public class MyBenchmark
|
||||
{
|
||||
[Benchmark(Baseline = true)]
|
||||
public void Original() { /* original code */ }
|
||||
|
||||
[Benchmark]
|
||||
public void Optimized() { /* optimized code */ }
|
||||
}
|
||||
```
|
||||
|
||||
**Run command**: `dotnet run -c Release --project path/to/benchmark`
|
||||
|
||||
**Common pitfalls**:
|
||||
- Running in Debug mode (JIT optimizations disabled, results meaningless)
|
||||
- Not returning computed values (JIT eliminates dead code)
|
||||
- Ignoring allocation metrics (throughput may improve but allocations increase)
|
||||
- Benchmarking with a debugger attached
|
||||
- Including setup costs in the measured method
|
||||
|
||||
---
|
||||
|
||||
## Load Testing
|
||||
|
||||
For HIGH-impact optimizations, perform before/after load testing to validate real-world improvement.
|
||||
|
||||
**k6 template**:
|
||||
```javascript
|
||||
import http from 'k6/http';
|
||||
import { check, sleep } from 'k6';
|
||||
|
||||
export const options = {
|
||||
stages: [
|
||||
{ duration: '30s', target: 20 },
|
||||
{ duration: '1m', target: 20 },
|
||||
{ duration: '10s', target: 0 },
|
||||
],
|
||||
thresholds: {
|
||||
http_req_duration: ['p(50)<100', 'p(95)<500', 'p(99)<1000'],
|
||||
http_req_failed: ['rate<0.01'],
|
||||
},
|
||||
};
|
||||
|
||||
export default function () {
|
||||
const res = http.get('http://localhost:5000/your-endpoint');
|
||||
check(res, {
|
||||
'status is 200': (r) => r.status === 200,
|
||||
'p95 under 500ms': (r) => r.timings.duration < 500,
|
||||
});
|
||||
sleep(1);
|
||||
}
|
||||
```
|
||||
|
||||
**While load testing, monitor simultaneously**:
|
||||
```bash
|
||||
dotnet-counters monitor -n <ProcessName> --counters System.Runtime,Microsoft.AspNetCore.Hosting
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## CI/CD Integration
|
||||
|
||||
For PR regression detection, use `benchmark-action/github-action-benchmark`:
|
||||
|
||||
```yaml
|
||||
- uses: benchmark-action/github-action-benchmark@v1
|
||||
with:
|
||||
tool: 'benchmarkdotnet'
|
||||
output-file-path: BenchmarkDotNet.Artifacts/results/*.json
|
||||
alert-threshold: '150%'
|
||||
comment-on-alert: true
|
||||
fail-on-alert: true
|
||||
```
|
||||
|
||||
This fails the PR if any benchmark regresses by more than 50% compared to the baseline.
|
||||
|
||||
---
|
||||
|
||||
## Code Review Mode: Quick Reference Commands
|
||||
|
||||
```bash
|
||||
# Baseline runtime health
|
||||
dotnet-counters monitor -n <ProcessName> --counters System.Runtime
|
||||
|
||||
# ASP.NET Core request metrics
|
||||
dotnet-counters monitor -n <ProcessName> --counters Microsoft.AspNetCore.Hosting
|
||||
|
||||
# Full monitoring (runtime + HTTP + custom meters)
|
||||
dotnet-counters monitor -n <ProcessName> --counters System.Runtime,Microsoft.AspNetCore.Hosting,Microsoft.AspNetCore.Server.Kestrel
|
||||
|
||||
# GC heap snapshot for before/after comparison
|
||||
dotnet-gcdump collect -n <ProcessName> -o before.gcdump
|
||||
# ... apply optimization ...
|
||||
dotnet-gcdump collect -n <ProcessName> -o after.gcdump
|
||||
|
||||
# 30-second CPU trace
|
||||
dotnet-trace collect -n <ProcessName> --duration 00:00:30
|
||||
dotnet-trace convert trace.nettrace --format speedscope
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Diagnostic Mode: Full CLI Commands
|
||||
|
||||
When profiling a live process (Mode A), use these commands by investigation stage:
|
||||
|
||||
```bash
|
||||
# Stage 1: Live triage
|
||||
dotnet-counters monitor -p <PID> --counters System.Runtime
|
||||
dotnet-counters monitor -n <ProcessName> --counters System.Runtime,Microsoft.AspNetCore.Hosting,Microsoft.AspNetCore.Server.Kestrel
|
||||
|
||||
# Stage 2: Stuck/hung process — get stacks immediately
|
||||
dotnet-stack report -p <PID>
|
||||
|
||||
# Stage 3: CPU + allocation hot paths
|
||||
dotnet-trace collect -p <PID> --duration 00:00:30
|
||||
dotnet-trace report <trace.nettrace> topN
|
||||
|
||||
# Stage 4: Heap composition
|
||||
dotnet-gcdump collect -p <PID> -o before.gcdump
|
||||
# ... apply optimization ...
|
||||
dotnet-gcdump collect -p <PID> -o after.gcdump
|
||||
dotnet-gcdump report <file.gcdump>
|
||||
|
||||
# Stage 5: Full dump for SOS analysis
|
||||
dotnet-dump collect -p <PID> --type Heap
|
||||
dotnet-dump analyze <dump> -c "dumpheap -stat" -c "exit"
|
||||
```
|
||||
|
|
@ -0,0 +1,223 @@
|
|||
# Memory & String Patterns
|
||||
|
||||
### Use ReadOnlySpan\<byte\> for Constant Byte Data
|
||||
🟡 **DO** assign constant byte arrays to `ReadOnlySpan<byte>` | .NET 5+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
byte[] data = new byte[] { 0x48, 0x65, 0x6C, 0x6C, 0x6F };
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
ReadOnlySpan<byte> data = [0x48, 0x65, 0x6C, 0x6C, 0x6F];
|
||||
ReadOnlySpan<int> primes = [2, 3, 5, 7, 11, 13];
|
||||
```
|
||||
|
||||
**Impact: ~100x faster access than static byte[] field, zero allocation.**
|
||||
|
||||
### Use stackalloc for Small Temporary Buffers
|
||||
🟡 **DO** use `stackalloc` for small, fixed-size temporary buffers | .NET Core+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
char[] buffer = new char[64];
|
||||
guid.TryFormat(buffer, out int written);
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
Span<char> buffer = stackalloc char[64];
|
||||
guid.TryFormat(buffer, out int written);
|
||||
```
|
||||
|
||||
**Impact: Zero heap allocation, no GC pressure, instant alloc/dealloc.**
|
||||
|
||||
### Use Span.TryWrite for Allocation-Free Interpolation
|
||||
🟡 **DO** use `MemoryExtensions.TryWrite` to format into `Span<char>` buffers | .NET 6+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
string formatted = $"Date: {dt:R}";
|
||||
destination.Write(formatted);
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
Span<char> buffer = stackalloc char[64];
|
||||
buffer.TryWrite($"Date: {dt:R}", out int charsWritten);
|
||||
```
|
||||
|
||||
**Impact: Zero heap allocation for formatting operations.**
|
||||
|
||||
### Use Span.Split() for Zero-Allocation Splitting
|
||||
🟡 **DO** use `MemoryExtensions.Split` for allocation-free string splitting | **.NET 10 (or .NET 9) only — NOT available on .NET 8**
|
||||
|
||||
❌ (allocates `string[]` — the only built-in option on .NET 8)
|
||||
```csharp
|
||||
string[] parts = input.Split(',');
|
||||
```
|
||||
✅ .NET 10
|
||||
```csharp
|
||||
foreach (Range range in input.AsSpan().Split(','))
|
||||
{
|
||||
ReadOnlySpan<char> segment = input.AsSpan(range);
|
||||
}
|
||||
```
|
||||
✅ .NET 8 fallback — manual `IndexOf` loop on the span (no allocation)
|
||||
```csharp
|
||||
ReadOnlySpan<char> remaining = input.AsSpan();
|
||||
while (!remaining.IsEmpty)
|
||||
{
|
||||
int idx = remaining.IndexOf(',');
|
||||
ReadOnlySpan<char> segment = idx < 0 ? remaining : remaining[..idx];
|
||||
// ... use segment ...
|
||||
remaining = idx < 0 ? default : remaining[(idx + 1)..];
|
||||
}
|
||||
```
|
||||
|
||||
**Impact: 208 bytes → 0 bytes per split, 2x faster on .NET 10. The manual .NET 8 loop is also zero-allocation but more verbose.**
|
||||
|
||||
### Use UTF8 String Literals (u8 suffix)
|
||||
🟡 **DO** use the `u8` suffix for compile-time UTF8 `ReadOnlySpan<byte>` | .NET 7+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
byte[] header = Encoding.UTF8.GetBytes("Content-Type");
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
ReadOnlySpan<byte> header = "Content-Type"u8;
|
||||
```
|
||||
|
||||
**Impact: 17ns → 0.006ns — eliminates runtime transcoding entirely.**
|
||||
|
||||
### Use ReadOnlySpan\<char\> Pattern Matching with switch
|
||||
🟡 **DO** use `switch` on `ReadOnlySpan<char>` for allocation-free string matching | C# 11+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
switch (attr.Value.Trim()) { case "preserve": /* ... */ break; }
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
switch (attr.Value.AsSpan().Trim())
|
||||
{
|
||||
case "preserve": return Preserve;
|
||||
case "default": return Default;
|
||||
}
|
||||
```
|
||||
|
||||
**Impact: Eliminates string allocation from Trim() in switch-based dispatch.**
|
||||
|
||||
### Use params ReadOnlySpan\<T\> to Eliminate Array Allocations
|
||||
🟡 **DO** add `params ReadOnlySpan<T>` overloads to library methods | **.NET 10 (or .NET 9) only — requires C# 13**
|
||||
|
||||
❌ (the only option on .NET 8 — accept the array allocation, or add explicit 1/2/3-argument overloads)
|
||||
```csharp
|
||||
public static void Log(params string[] messages) { /* ... */ }
|
||||
Log("Starting", "Processing", "Done");
|
||||
```
|
||||
✅ .NET 10
|
||||
```csharp
|
||||
public static void Log(params ReadOnlySpan<string> messages) { /* ... */ }
|
||||
Log("Starting", "Processing", "Done");
|
||||
```
|
||||
✅ .NET 8 fallback — keep `params string[]` and add fixed-arity overloads for the hot common cases
|
||||
```csharp
|
||||
public static void Log(string m) { /* ... */ }
|
||||
public static void Log(string m1, string m2) { /* ... */ }
|
||||
public static void Log(string m1, string m2, string m3) { /* ... */ }
|
||||
public static void Log(params string[] messages) { /* fallback for 4+ args */ }
|
||||
```
|
||||
|
||||
**Impact: Eliminates params array allocation on .NET 10. On .NET 8 fixed-arity overloads cover the hot 1–3 argument cases.**
|
||||
|
||||
### Avoid Chained String-Returning Operations
|
||||
🟡 **AVOID** chains of 3+ string-returning method calls that each allocate intermediates | .NET Core+
|
||||
|
||||
**Pattern 1: Chained .Replace() calls**
|
||||
|
||||
❌
|
||||
```csharp
|
||||
string result = input.Replace("a", "b").Replace("c", "d").Replace("e", "f");
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
var sb = new StringBuilder(input.Length);
|
||||
// single pass replacing all patterns
|
||||
```
|
||||
|
||||
**Pattern 2: Chained Regex.Replace() calls**
|
||||
|
||||
❌
|
||||
```csharp
|
||||
public static string Underscore(this string input) =>
|
||||
Regex3.Replace(Regex2.Replace(Regex1.Replace(input, "$1_$2"), "$1_$2"), "_").ToLower();
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
return string.Create(totalLength, state, (span, s) => { /* write directly */ });
|
||||
```
|
||||
|
||||
**Pattern 3: += string concatenation in loops**
|
||||
|
||||
❌
|
||||
```csharp
|
||||
string result = "";
|
||||
foreach (var part in parts)
|
||||
result += separator + part;
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
var sb = new StringBuilder();
|
||||
foreach (var part in parts)
|
||||
sb.Append(separator).Append(part);
|
||||
return sb.ToString();
|
||||
```
|
||||
|
||||
**Impact: Eliminates N-1 intermediate string allocations per chain. For `+=` in loops, eliminates O(n²) total allocation.**
|
||||
|
||||
### Cache char.ToString() for Known Character Sets
|
||||
🟡 **DO** cache `char.ToString()` results when the set of characters is small and known | .NET Core+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
return symbol.ToString();
|
||||
|
||||
foreach (var prefix in UnitPrefixes)
|
||||
input = input.Replace(prefix.Value.Name, prefix.Key.ToString());
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
private static readonly FrozenDictionary<char, string> s_charStrings =
|
||||
new Dictionary<char, string>
|
||||
{
|
||||
['k'] = "k", ['M'] = "M", ['G'] = "G",
|
||||
}.ToFrozenDictionary();
|
||||
|
||||
return s_charStrings[symbol];
|
||||
```
|
||||
|
||||
**Impact: Eliminates one string allocation per char.ToString() call. Significant when called in loops or on hot paths.**
|
||||
|
||||
## Detection
|
||||
|
||||
Scan recipes for memory and string anti-patterns. Run these and report exact counts.
|
||||
|
||||
```bash
|
||||
# .ToLower()/.ToUpper() without culture parameter (allocates + culture-sensitive)
|
||||
grep -rn --include='*.cs' -E '\.(ToLower|ToUpper)\(\)' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
|
||||
# Chained .Replace( calls (3+ on one line — intermediate string allocations)
|
||||
grep -rn --include='*.cs' '\.Replace(.*\.Replace(.*\.Replace(' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
|
||||
# params in method signatures (array allocation per call)
|
||||
grep -rn --include='*.cs' 'params ' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
|
||||
# LINQ on strings — .All/.Any on IEnumerable<char> (replace with foreach loop)
|
||||
grep -rn --include='*.cs' -E '\.(All|Any)\(char\.' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
```
|
||||
|
||||
### Patterns Requiring Manual Review
|
||||
|
||||
- **Boxing via string.Format**: Can't determine argument types from grep — needs type analysis
|
||||
- **`+=` string concatenation in loops**: `+=` matches all types (int, list, event, string) — needs type context to confirm string
|
||||
- **`char.ToString()`**: Requires knowing the variable type is `char` — not reliably greppable
|
||||
|
|
@ -0,0 +1,117 @@
|
|||
---
|
||||
description: CLR memory model and GC guidance for csharp-dotnet-cli-optimization.
|
||||
metadata:
|
||||
tags: [clr, gc, heap, stack, loh, memory-model]
|
||||
---
|
||||
|
||||
# CLR Memory Model And GC
|
||||
|
||||
Use this reference when the user asks why allocations, boxing, heap growth, GC pauses, or stack-based techniques behave the way they do.
|
||||
|
||||
## Core Model
|
||||
|
||||
- Reference types allocate objects on the managed heap. Local variables and fields hold references to those objects.
|
||||
- Value types store their data directly. A value type local is often stored in stack storage, but value types also live inline inside object fields and array elements.
|
||||
- "`struct` means stack" is false. The useful distinction is inline storage plus copy semantics, not "stack forever".
|
||||
- Boxing converts a value type to `object` or an interface by allocating a new heap object and copying the value into it.
|
||||
- `ref struct` types, including `Span<T>` and `ReadOnlySpan<T>`, are stack-constrained wrappers that can't escape to the managed heap.
|
||||
- `Memory<T>` and `ReadOnlyMemory<T>` are the heap-storable counterparts when data must live across `await`, callbacks, or object fields.
|
||||
|
||||
## How The GC Works
|
||||
|
||||
- The GC is generational: gen0 for young objects, gen1 as a buffer, gen2 for long-lived survivors.
|
||||
- The large object heap (LOH) is used for allocations at or above 85,000 bytes by default.
|
||||
- Background GC is enabled by default. It reduces pause impact for full collections but does not make them free.
|
||||
- Server GC and workstation GC are process-level choices. The defaults are usually right unless measurement says otherwise.
|
||||
- On modern 64-bit Windows and Linux, the GC internally uses regions, but the optimization model for application code is still about generations, allocation rate, survivor rate, LOH churn, and pinning.
|
||||
|
||||
## What Usually Makes GC Expensive
|
||||
|
||||
- High allocation rate on hot paths
|
||||
- Objects surviving long enough to promote into older generations
|
||||
- Large transient allocations that churn the LOH
|
||||
- Excessive pinning that increases fragmentation
|
||||
- Finalizers on objects that should have been deterministic `Dispose` calls instead
|
||||
|
||||
## Wrong vs Better
|
||||
|
||||
| Wrong | Better | Why |
|
||||
|---|---|---|
|
||||
| Assume a `struct` is always stack allocated | Explain whether it will be copied, boxed, stored inline, or escape | That is what actually drives cost |
|
||||
| Allocate large temporary arrays repeatedly | Reuse, pool, or redesign the algorithm if measurement shows LOH churn | LOH allocations are cleared and collected with gen2 work |
|
||||
| Call `GC.Collect()` to "fix" memory pressure | Lower allocation rate and object lifetime first | Forced GC usually adds pause time and hides the real problem |
|
||||
| Pin many buffers for long periods | Minimize pin count and pin duration | Pinning can fragment the heap |
|
||||
| Use finalizers for routine cleanup | Use `IDisposable`, `using`, and `SafeHandle` for unmanaged resources | Finalization is slower and delays reclamation |
|
||||
|
||||
## GC Configuration Rules
|
||||
|
||||
- Treat GC configuration changes as process-wide tuning, not local fixes.
|
||||
- Prefer runtime defaults unless counters and traces show a clear reason to change them.
|
||||
- Choose server GC for throughput-oriented workloads only after measurement.
|
||||
- Use low-latency modes sparingly and for bounded windows. They reduce GC intrusiveness by letting memory grow and can increase fragmentation.
|
||||
- If you are tuning in containers or hard memory limits, treat heap hard-limit settings as operational controls, not code-level optimizations.
|
||||
|
||||
## Bad vs Good Examples
|
||||
|
||||
Bad:
|
||||
|
||||
```csharp
|
||||
for (int i = 0; i < 10_000; i++)
|
||||
{
|
||||
DoWork(new byte[200_000]);
|
||||
}
|
||||
GC.Collect();
|
||||
```
|
||||
|
||||
Better:
|
||||
|
||||
```csharp
|
||||
byte[] buffer = ArrayPool<byte>.Shared.Rent(200_000);
|
||||
try
|
||||
{
|
||||
for (int i = 0; i < 10_000; i++)
|
||||
{
|
||||
DoWork(buffer);
|
||||
}
|
||||
}
|
||||
finally
|
||||
{
|
||||
ArrayPool<byte>.Shared.Return(buffer);
|
||||
}
|
||||
```
|
||||
|
||||
Bad:
|
||||
|
||||
```csharp
|
||||
public sealed class NativeThing
|
||||
{
|
||||
~NativeThing() => ReleaseHandle();
|
||||
}
|
||||
```
|
||||
|
||||
Better:
|
||||
|
||||
```csharp
|
||||
public sealed class NativeThing : IDisposable
|
||||
{
|
||||
public void Dispose()
|
||||
{
|
||||
ReleaseHandle();
|
||||
GC.SuppressFinalize(this);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Practical Heuristics
|
||||
|
||||
- If counters show rising allocation rate and frequent gen0 collections, start by eliminating short-lived allocations.
|
||||
- If gen2 collections or LOH size are the problem, look for survivor growth, pinned buffers, and large transient objects.
|
||||
- If a change turns classes into structs, verify both allocation wins and copy costs.
|
||||
- If the process is memory-constrained, inspect runtime GC settings before changing code blindly.
|
||||
|
||||
## What Not To Claim
|
||||
|
||||
- Do not claim that moving code to `struct` always reduces memory.
|
||||
- Do not claim that stack allocation is always faster than pooling.
|
||||
- Do not claim that background GC removes pause concerns.
|
||||
- Do not claim that the GC is the problem unless counters, traces, or dumps support that diagnosis.
|
||||
|
|
@ -0,0 +1,95 @@
|
|||
# Regex Patterns
|
||||
|
||||
### Choose the Right Regex Engine Mode
|
||||
🟡 **DO** use `[GeneratedRegex]` for all static regex patterns, but never remove `NonBacktracking` if present | .NET 7+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
var r = new Regex(dynamicPattern, RegexOptions.Compiled);
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
[GeneratedRegex("pattern")]
|
||||
private static partial Regex MyRegex();
|
||||
|
||||
var safe = new Regex(untrustedPattern, RegexOptions.NonBacktracking);
|
||||
|
||||
var oneOff = new Regex("pattern");
|
||||
```
|
||||
|
||||
**Impact: Source generator is always beneficial for static patterns. NonBacktracking prevents O(2^N) worst case — never remove it if present.**
|
||||
|
||||
### Use IsMatch When You Only Need a Boolean Result
|
||||
🟡 **DO** use `IsMatch` instead of `Match(...).Success` | .NET 7+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
bool found = Regex.Match(input, pattern).Success;
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
bool found = Regex.IsMatch(input, pattern);
|
||||
```
|
||||
|
||||
**Impact: Avoids Match object allocation; with NonBacktracking, ~3x faster by skipping capture computation.**
|
||||
|
||||
### Use Regex.Count/EnumerateMatches Instead of Matches
|
||||
🟡 **DO** use `Count()` and `EnumerateMatches()` for allocation-free match processing | .NET 7+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
int count = 0;
|
||||
Match m = regex.Match(text);
|
||||
while (m.Success) { count++; m = m.NextMatch(); }
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
int count = regex.Count(text);
|
||||
|
||||
foreach (ValueMatch m in Regex.EnumerateMatches(text, @"\b\w+\b"))
|
||||
{
|
||||
ReadOnlySpan<char> word = text.AsSpan(m.Index, m.Length);
|
||||
}
|
||||
```
|
||||
|
||||
**Impact: ~3x faster than Match/NextMatch with NonBacktracking. Zero allocations for both Count and EnumerateMatches.**
|
||||
|
||||
### Use Span-Based Regex APIs for Allocation-Free Matching
|
||||
🟡 **DO** use `ReadOnlySpan<char>` overloads for regex matching on spans | .NET 7+
|
||||
|
||||
❌
|
||||
```csharp
|
||||
string sub = largeBuffer.Substring(start, length);
|
||||
bool found = Regex.IsMatch(sub, pattern);
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
ReadOnlySpan<char> text = largeBuffer.AsSpan(start, length);
|
||||
foreach (ValueMatch m in Regex.EnumerateMatches(text, @"\b\w+\b"))
|
||||
{
|
||||
ReadOnlySpan<char> word = text.Slice(m.Index, m.Length);
|
||||
}
|
||||
```
|
||||
|
||||
**Impact: Eliminates string allocations when working with spans — particularly valuable in high-throughput parsing pipelines.**
|
||||
|
||||
## Detection
|
||||
|
||||
Scan recipes for regex anti-patterns. Run these and report exact counts.
|
||||
|
||||
```bash
|
||||
# Compiled regex count (startup cost budget — compare ratio to GeneratedRegex)
|
||||
grep -rn --include='*.cs' 'RegexOptions.Compiled' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
|
||||
# GeneratedRegex count (already optimized — verify the inverse)
|
||||
grep -rn --include='*.cs' 'GeneratedRegex' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
|
||||
# Uncached new Regex() calls (construction cost per call)
|
||||
grep -rn --include='*.cs' 'new Regex(' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
```
|
||||
|
||||
When `RegexOptions.Compiled` appears inside a class constructor or field initializer of an instantiated class (not a static singleton), count how many instances of that class are created at startup to determine total compiled regex budget. For example, if a `Rule` class compiles a regex in its constructor and 122 rules are registered, that is 122 compiled regexes at startup.
|
||||
|
||||
### Patterns Requiring Manual Review
|
||||
|
||||
- **`new Regex(` uncached**: Field assignment may span multiple lines — grep on one line is unreliable. Verify that matched instances are stored in `static readonly` fields or `[GeneratedRegex]`.
|
||||
|
|
@ -0,0 +1,38 @@
|
|||
# Structural Patterns
|
||||
|
||||
Patterns detected by the **absence** of a keyword or interface. These require codebase-wide counting scans, not single-file matching.
|
||||
|
||||
### Seal Classes for Devirtualization
|
||||
🟡 **DO** seal all leaf classes (those not subclassed) | .NET Core 3.0+
|
||||
|
||||
Sealing lets the JIT devirtualize/inline virtual calls and use pointer comparison for type checks. Every non-abstract, non-static class that is not subclassed should be sealed.
|
||||
|
||||
**Detection:** This is an absence pattern — scan for classes that are NOT sealed.
|
||||
|
||||
```bash
|
||||
# Count unsealed (non-abstract, non-static) classes
|
||||
grep -rn --include='*.cs' -E '^\s*((public|internal|private|protected|file)\s+)?(partial\s+)?class ' --exclude-dir=bin --exclude-dir=obj . | grep -v 'sealed' | grep -v 'abstract' | grep -v 'static' | wc -l
|
||||
|
||||
# Count already-sealed classes (verify the inverse)
|
||||
grep -rn --include='*.cs' 'sealed class' --exclude-dir=bin --exclude-dir=obj . | wc -l
|
||||
```
|
||||
|
||||
**Exclusions:** Do not seal classes that are subclassed elsewhere in the codebase. Identifying base classes requires manual review — grep for `: ClassName` patterns and cross-reference, but expect false positives from interface implementations and generic constraints.
|
||||
|
||||
❌
|
||||
```csharp
|
||||
internal class MyHandler : Base
|
||||
{ public override int Run() => 42; }
|
||||
```
|
||||
✅
|
||||
```csharp
|
||||
internal sealed class MyHandler : Base
|
||||
{ public override int Run() => 42; }
|
||||
```
|
||||
|
||||
**Impact: Virtual calls up to 500x faster; type checks ~25x faster. Severity scales with count.**
|
||||
|
||||
**Scale-based severity:**
|
||||
- 1-10 unsealed leaf classes → ℹ️ Info
|
||||
- 11-50 unsealed leaf classes → 🟡 Moderate
|
||||
- 50+ unsealed leaf classes → 🟡 Moderate (elevated priority)
|
||||
145
.skills/dotnet-security-review/SKILL.md
Normal file
145
.skills/dotnet-security-review/SKILL.md
Normal file
|
|
@ -0,0 +1,145 @@
|
|||
---
|
||||
name: dotnet-security-review
|
||||
description: >-
|
||||
Performs a systematic C#/ASP.NET Core security code review on .NET 8 (C# 12)
|
||||
and .NET 10 (C# 14) codebases. Covers OWASP Top 10, authentication/authorization
|
||||
audit, input validation, cryptography, dependency vulnerabilities, security
|
||||
headers, middleware pipeline, and CI/CD security posture.
|
||||
metadata:
|
||||
platform: ".NET 8 and .NET 10 (no .NET 9 projects in scope)"
|
||||
---
|
||||
|
||||
# Security Code Review for C# / ASP.NET Core (.NET 8 + .NET 10)
|
||||
|
||||
You are a security auditor performing a thorough, evidence-based code review. Every finding MUST include file path, line number, severity, impact, and a concrete fix.
|
||||
|
||||
## Step 0 — Detect the target framework
|
||||
|
||||
Before scoring findings, follow `../../references/detect-target-framework.md`. The security guidance below applies to **both** .NET 8 and .NET 10 unless explicitly marked. A few items are .NET 10-only — when reviewing a .NET 8 project, don't recommend them as "fixes":
|
||||
|
||||
- **ASP.NET Core Identity passkeys** (`AddPasskeys()`) — .NET 10 only. On .NET 8, recommend external IdP / `Fido2NetLib` or password+TOTP.
|
||||
- **Minimal-API built-in validation** (`AddValidation()`) — .NET 10 only. On .NET 8, FluentValidation + `IEndpointFilter` is the safe equivalent.
|
||||
- **First-party `Microsoft.AspNetCore.OpenApi`** — .NET 9+ only. On .NET 8 the project should use `Swashbuckle.AspNetCore`; flag missing OpenAPI security schemes accordingly.
|
||||
- **`HybridCache`** — .NET 9+ only. On .NET 8 verify `IDistributedCache` configurations (encryption-at-rest, key prefixing, TLS to Redis) directly.
|
||||
- The C# 14 `field` keyword, `extension(...)` blocks, null-conditional assignment, and partial constructors **do not compile on net8.0** — never propose security fixes that introduce them on a .NET 8 project.
|
||||
|
||||
All cryptography, JWT, authorization-policy, header, and middleware guidance applies identically on both targets.
|
||||
|
||||
## Target Selection
|
||||
|
||||
The user's arguments are in `$ARGUMENTS`.
|
||||
|
||||
- If `$ARGUMENTS` contains a file path or directory, review that target.
|
||||
- If `$ARGUMENTS` is "all", review the entire codebase starting from the solution root.
|
||||
- If `$ARGUMENTS` is empty, run `git diff --name-only HEAD~5` to find recently changed `.cs` files. If none, ask the user what to review.
|
||||
|
||||
When reviewing a directory or "all", use Glob to find `**/*.cs` files, then prioritize:
|
||||
1. Controllers, filters, middleware (`*Controller.cs`, `*Filter.cs`, `Program.cs`)
|
||||
2. Auth handlers and delegating handlers (`*Handler.cs`, `*DelegatingHandler.cs`)
|
||||
3. Service implementations handling external input or secrets
|
||||
4. Repository and data access code
|
||||
5. Configuration and DI registration (`*Extensions.cs`, `*Options.cs`)
|
||||
6. Validators
|
||||
|
||||
## Review Process
|
||||
|
||||
Execute each phase sequentially. Use the Read tool for files and the Grep tool for pattern searches. NEVER use bash `grep` or `rg` -- always use the Grep tool.
|
||||
|
||||
### Phase 1: Automated Pattern Scanning
|
||||
|
||||
Read `references/scanning-patterns.md` for the full pattern catalog. Run all Grep searches in parallel across `.cs` files in the target scope. Each pattern targets a specific vulnerability class: injection, deserialization, cryptography, async anti-patterns, data exposure, SSRF, missing controls, ReDoS, log injection, open redirect, cookie security, file upload, claims safety, and thread safety.
|
||||
|
||||
### Phase 2: File-by-File Deep Review
|
||||
|
||||
Read `references/deep-review-categories.md` for the complete checklist (Categories A through L). For each file in scope (or top ~20 most security-relevant files when reviewing "all"), check all applicable categories:
|
||||
|
||||
- **A**: Authentication & Authorization (JWT validation, auth schemes, IDOR)
|
||||
- **B**: Input Validation
|
||||
- **C**: Error Handling & Information Leakage
|
||||
- **D**: Cryptography & Secrets
|
||||
- **E**: Data Protection & PII
|
||||
- **F**: Concurrency & State Safety
|
||||
- **G**: CancellationToken Propagation
|
||||
- **H**: HTTP Client Security (resilience handlers, DNS refresh)
|
||||
- **I**: Configuration Security
|
||||
- **J**: Logging & Monitoring Security
|
||||
- **K**: Output Encoding & Response Security
|
||||
- **L**: Supply Chain & Build Security
|
||||
|
||||
### Phase 3: Architecture & Project-Specific Checks
|
||||
|
||||
Read `references/architecture-checks.md` for checks tailored to common ASP.NET Core project patterns. These cover endpoint authorization verification, anonymous endpoint abuse potential, OTP/MFA security, exception handling coverage, optimistic concurrency, state expiry, blob storage SAS security, message queue security, JSON serialization settings, background task queue safety, rate limiting, security headers, middleware ordering, and NuGet audit configuration.
|
||||
|
||||
Read the project's CLAUDE.md or AGENTS.md for project-specific architecture details to inform these checks.
|
||||
|
||||
### Phase 4: Dependency Vulnerability Check
|
||||
|
||||
Read `references/dependencies-and-headers.md` (Phase 4 section) for dependency scanning patterns. Check `.csproj` files for known-vulnerable versions and NuGet audit configuration.
|
||||
|
||||
### Phase 5: Security Headers & Middleware Pipeline
|
||||
|
||||
Read `references/dependencies-and-headers.md` (Phase 5 section) for the 14-item headers checklist and middleware ordering verification.
|
||||
|
||||
## Output Format
|
||||
|
||||
### Security Review Report
|
||||
|
||||
**Scope:** [files/directories reviewed]
|
||||
**Date:** [current date]
|
||||
**Risk Summary:** [X CRITICAL, Y HIGH, Z MEDIUM, W LOW, V INFO]
|
||||
|
||||
#### Findings
|
||||
|
||||
For each finding:
|
||||
|
||||
**[SEVERITY] [SHORT-TITLE]**
|
||||
- **Location:** `file/path.cs:LINE`
|
||||
- **Category:** [OWASP category or security domain]
|
||||
- **Description:** [What the vulnerability is and why it matters]
|
||||
- **Impact:** [What an attacker could achieve]
|
||||
- **Recommendation:** [Specific fix with code example]
|
||||
|
||||
#### Summary Table
|
||||
|
||||
| # | Severity | Category | File | Description |
|
||||
|---|----------|----------|------|-------------|
|
||||
| 1 | CRITICAL | ... | ... | ... |
|
||||
|
||||
#### Recommendations
|
||||
|
||||
1. Immediate fixes (CRITICAL/HIGH)
|
||||
2. Short-term improvements (MEDIUM)
|
||||
3. Long-term hardening (LOW/INFO)
|
||||
4. Tooling recommendations (NuGet audit, SAST integration, etc.)
|
||||
|
||||
## Severity
|
||||
|
||||
Use standard severity: CRITICAL > HIGH > MEDIUM > LOW > INFO. CRITICAL = actively exploitable, HIGH = significant with effort, MEDIUM = increased attack surface, LOW = minor improvement, INFO = hardening suggestion.
|
||||
|
||||
## Anti-Rationalization Table
|
||||
|
||||
| Rationalization | Reality |
|
||||
|---|---|
|
||||
| "This is just a test file" | Test code handling secrets or auth IS production-relevant. Report as INFO. |
|
||||
| "Probably a false positive" | ALWAYS read surrounding code before dismissing. If you cannot prove it safe, report it. |
|
||||
| "The framework handles this" | Verify the protection is actually enabled and configured. Defaults can be overridden. |
|
||||
| "Internal API, not public-facing" | Internal APIs are attacked via SSRF, supply chain, lateral movement. |
|
||||
| "No one would exploit this" | Threat models change. Report it; let the team decide risk acceptance. |
|
||||
|
||||
## Red Flags
|
||||
|
||||
STOP and investigate deeper if you encounter any of these:
|
||||
- Any endpoint without an explicit auth attribute (`[Authorize]` or `[AllowAnonymous]`)
|
||||
- Any `catch` block returning raw exception data to the client
|
||||
- Any hardcoded key, token, password, or connection string literal
|
||||
- Any `new HttpClient()` (should use `IHttpClientFactory`)
|
||||
- Any `TypeNameHandling` value other than `None`
|
||||
|
||||
## Important Guidelines
|
||||
|
||||
1. Only report real findings with evidence (file path and line number). Do not speculate.
|
||||
2. If a pattern search returns no results, note "No issues found" and move on.
|
||||
3. For false positives (e.g., `System.Random` in tests, not production), note as INFO with explanation.
|
||||
4. Prioritize production code over test code.
|
||||
5. When reviewing "all", cap the report at the 30 most significant findings.
|
||||
6. ALWAYS verify context before reporting -- a pattern match alone is not a finding. Read the surrounding code.
|
||||
134
.skills/dotnet-security-review/references/architecture-checks.md
Normal file
134
.skills/dotnet-security-review/references/architecture-checks.md
Normal file
|
|
@ -0,0 +1,134 @@
|
|||
# Architecture & Project-Specific Checks Reference
|
||||
# Phase 3 checks for common ASP.NET Core project patterns. Read the project's CLAUDE.md
|
||||
# or AGENTS.md for project-specific details (endpoint list, service names, DI registrations)
|
||||
# to inform these checks.
|
||||
|
||||
## Check 1: Endpoint Auth Matrix Verification
|
||||
|
||||
Cross-reference the controller's actual `[Authorize]`/`[AllowAnonymous]` attributes against the project's documented auth requirements. Read CLAUDE.md or AGENTS.md for the expected auth matrix. Any mismatch is CRITICAL.
|
||||
|
||||
For projects with multiple auth schemes (e.g., Azure AD + custom JWT), verify each endpoint uses the correct scheme/policy.
|
||||
|
||||
## Check 2: Anonymous Endpoint Abuse Potential
|
||||
|
||||
For each `[AllowAnonymous]` endpoint, verify:
|
||||
- Rate limiting or throttling exists for sensitive operations (e.g., code generation, login attempts)
|
||||
- Enumeration attacks are mitigated (IDs are GUIDs or non-sequential, not auto-increment)
|
||||
- No state modification without prior authentication or verification (e.g., OTP first)
|
||||
|
||||
## Check 3: OTP / MFA Security Review
|
||||
|
||||
If the project implements OTP or MFA, read the service implementation and verify:
|
||||
- Code length is sufficient (6+ characters)
|
||||
- Codes are generated with `RandomNumberGenerator`
|
||||
- Hash is SHA-256 or stronger (not MD5/SHA1)
|
||||
- Expiry is enforced (typically 5-10 minutes)
|
||||
- Wrong attempt counter increments correctly and triggers lockout after a threshold
|
||||
- No timing side-channel in hash comparison
|
||||
|
||||
## Check 4: Exception Handling Coverage
|
||||
|
||||
Grep for `throw new` statements. Verify that all thrown exceptions are either:
|
||||
- The project's structured error type (e.g., `ApiException`, `DomainException`, or the project's custom base exception), OR
|
||||
- Known typed exceptions for external service failures
|
||||
|
||||
Any unstructured exception thrown from handler/service code may bypass error filters and leak internal details.
|
||||
|
||||
## Check 5: Optimistic Concurrency on State Writes
|
||||
|
||||
If using a database with optimistic concurrency (ETags, row versions):
|
||||
- Verify every write/update operation passes the concurrency token
|
||||
- Verify the concurrency token store/tracking mechanism is consulted on every read/write cycle
|
||||
|
||||
## Check 6: Expired State Handling
|
||||
|
||||
If the project uses application-level state expiry (not DB TTL):
|
||||
- Verify expired records are deleted or excluded on read (not returned to callers)
|
||||
- Verify callers cannot act on expired data
|
||||
|
||||
## Check 7: Blob Storage SAS URL Security
|
||||
|
||||
If the project generates SAS URLs for blob storage:
|
||||
- SAS token expiry is short-lived (minutes, not days)
|
||||
- Permission is read-only (not write/delete)
|
||||
- Scoped to the specific blob (not container-level)
|
||||
|
||||
## Check 8: Message Queue Security
|
||||
|
||||
If the project uses message queues (Service Bus, RabbitMQ, etc.):
|
||||
- Messages do not contain secrets or unnecessary PII
|
||||
- Queue connections use managed identity or connection strings from secret stores
|
||||
|
||||
## Check 9: JSON Serialization Settings
|
||||
|
||||
Check that `TypeNameHandling` is set to `None` (default) and not `Auto`/`All` anywhere. This applies to both Newtonsoft.Json and any custom serializer configuration.
|
||||
|
||||
## Check 10: Background Task Queue Safety
|
||||
|
||||
If the project uses a background task queue:
|
||||
- Bounded capacity prevents unbounded memory growth
|
||||
- Backpressure is handled correctly (not silently dropping critical events like audit logs)
|
||||
- Task failures are observed and logged/metered
|
||||
|
||||
## Check 11: Custom Token / JWT Security
|
||||
|
||||
If the project issues its own JWTs (not just validating external tokens):
|
||||
- **Algorithm**: HMAC-SHA256 or stronger (RSA for distributed validation)
|
||||
- **Signing key source**: Key loaded from configuration/secret store, NOT hardcoded
|
||||
- **Signing key length**: Minimum 256 bits (32 bytes) for HMAC-SHA256
|
||||
- **Token expiry**: Appropriately capped (tokens should not outlive the session/resource they protect)
|
||||
- **Claims validation**: Custom claims (e.g., resource IDs) are validated against route parameters by an authorization handler
|
||||
- **TokenValidationParameters**: `ValidateIssuer`, `ValidateAudience`, `ValidateLifetime`, `ValidateIssuerSigningKey` all `true`
|
||||
- **ClockSkew**: Tightened from default 5 minutes to 2 minutes or less
|
||||
|
||||
## Check 12: Response Data Sanitization
|
||||
|
||||
If the project sanitizes response data (e.g., stripping internal paths or fields):
|
||||
- Sanitization handles malformed input gracefully (does not throw/crash)
|
||||
- Only known sensitive fields are stripped (no over-stripping that breaks functionality)
|
||||
- Sanitization is applied on every code path returning the data (not just the happy path)
|
||||
|
||||
## Check 13: Rate Limiting
|
||||
|
||||
Verify rate limiting posture:
|
||||
- Grep for `AddRateLimiter` and `UseRateLimiter` -- if absent, note as finding
|
||||
- Application-level throttling exists for sensitive operations (e.g., SMS/code generation resend limits)
|
||||
- Brute force protection exists for verification endpoints (wrong attempt lockout)
|
||||
- **Recommendation**: Add ASP.NET Core `System.Threading.RateLimiting` middleware for IP-based throttling on public endpoints
|
||||
|
||||
## Check 14: Security Headers Completeness
|
||||
|
||||
Check Program.cs / middleware for these headers (report missing ones):
|
||||
- `X-Content-Type-Options: nosniff`
|
||||
- `X-Frame-Options: DENY`
|
||||
- `Referrer-Policy: strict-origin-when-cross-origin`
|
||||
- `Permissions-Policy: camera=(), microphone=(), geolocation=()`
|
||||
- `X-XSS-Protection: 0` (disable legacy XSS filter; CSP is the modern replacement)
|
||||
- `Content-Security-Policy` (at minimum for APIs: `default-src 'none'`)
|
||||
- `Server` header removed
|
||||
- `X-Powered-By` header removed
|
||||
|
||||
## Check 15: Middleware Pipeline Ordering
|
||||
|
||||
Read Program.cs and verify correct middleware order:
|
||||
1. `UseExceptionHandler` (outermost -- catches everything)
|
||||
2. `UseHsts` (non-development only)
|
||||
3. `UseHttpsRedirection`
|
||||
4. Security headers middleware
|
||||
5. `UseRateLimiter` (if present)
|
||||
6. `UseRouting` (if explicit)
|
||||
7. `UseCors`
|
||||
8. `UseAuthentication`
|
||||
9. `UseAuthorization`
|
||||
10. `MapControllers` / endpoints
|
||||
|
||||
Authentication MUST come before Authorization. CORS MUST come before Authentication. ExceptionHandler MUST be first.
|
||||
|
||||
## Check 16: NuGet Audit & Build Security
|
||||
|
||||
Check for build-level security configuration:
|
||||
- Does `Directory.Build.props` exist? If so, verify `NuGetAudit`, `NuGetAuditMode`, `NuGetAuditLevel` settings.
|
||||
- Are Roslyn security analyzer packages referenced? (`SecurityCodeScan.VS2019`, `SonarAnalyzer.CSharp`, `Meziantou.Analyzer`)
|
||||
- Are `AnalysisLevel` / `AnalysisMode` set in `.csproj` or `Directory.Build.props`?
|
||||
- Run Grep for floating versions: `Version="[^"]*\*"` in `.csproj` files
|
||||
- Recommend `dotnet list package --vulnerable --include-transitive` as a CI step
|
||||
|
|
@ -0,0 +1,107 @@
|
|||
# Deep Review Categories Reference
|
||||
# File-by-file review checklist for Phase 2. For each file in scope (or top ~20 most
|
||||
# security-relevant files when reviewing "all"), read the file and check each applicable category.
|
||||
|
||||
## Category A: Authentication & Authorization
|
||||
|
||||
1. Every controller action has either `[Authorize]` (class or method level) or `[AllowAnonymous]` explicitly.
|
||||
2. No IDOR: when accessing resources by ID, verify the handler checks that the caller owns or is authorized to access that resource.
|
||||
3. JWT validation settings are strict: issuer, audience, lifetime, algorithm all validated.
|
||||
4. Token acquisition uses correct flow: app tokens for backend-to-backend, OBO only where user context is needed.
|
||||
5. No `[AllowAnonymous]` on endpoints that modify sensitive state without alternative authentication (e.g., OTP verification first).
|
||||
6. **Multiple auth scheme verification**: If the project uses multiple auth schemes (e.g., Azure AD + custom JWT), verify correct scheme is applied per endpoint. No scheme confusion between internal and client-facing endpoints.
|
||||
7. **JWT `alg:none` rejection**: Verify `TokenValidationParameters` does NOT allow `alg:none`. All schemes must validate the signing algorithm (`ValidateIssuerSigningKey = true`).
|
||||
8. **HMAC signing key minimum length**: If using HMAC-SHA256 for JWT signing, the key must be at least 256 bits (32 bytes). Check options validation.
|
||||
9. **Structured error responses on auth failure**: `OnChallenge` (401) and `OnForbidden` (403) events should return structured JSON error responses, not default HTML/empty responses.
|
||||
|
||||
## Category B: Input Validation
|
||||
|
||||
1. All DTOs accepted by handlers have corresponding FluentValidation validators registered.
|
||||
2. Route parameters are validated for format before use (e.g., GUID format, positive integers).
|
||||
3. File uploads are validated for content type, size, and extension (not just extension).
|
||||
4. No unvalidated user input flows into file paths, URLs, SQL, commands, or log messages.
|
||||
5. Phone numbers, emails, and other PII are validated and normalized before processing.
|
||||
|
||||
## Category C: Error Handling & Information Leakage
|
||||
|
||||
1. All expected errors use a structured error type -- never return raw exception details to clients.
|
||||
2. Exception filters catch known exception types and return only safe error payloads.
|
||||
3. Unknown exceptions are wrapped as generic 500 errors without stack traces or internal details.
|
||||
4. Error messages returned to clients do not reveal internal architecture, database schema, or file paths.
|
||||
5. Catch blocks never silently swallow exceptions -- they must log or rethrow.
|
||||
|
||||
## Category D: Cryptography & Secrets
|
||||
|
||||
1. OTP/MFA codes use `RandomNumberGenerator` (not `System.Random`).
|
||||
2. Hash comparison uses constant-time comparison to prevent timing attacks.
|
||||
3. Hash storage uses a secure algorithm (SHA-256 minimum; bcrypt/Argon2 for passwords).
|
||||
4. No secrets, connection strings, or API keys appear in source code or `appsettings.json` committed to git.
|
||||
5. Options validation (`ValidateOnStart()`) is configured to reject placeholder secrets in production.
|
||||
|
||||
## Category E: Data Protection & PII
|
||||
|
||||
1. Sensitive fields (phone numbers, etc.) are masked before returning to unauthenticated callers.
|
||||
2. PII (names, addresses, phone numbers, emails) is not logged in full -- use masking.
|
||||
3. Sensitive internal fields (hash values, internal IDs) are excluded from API responses.
|
||||
4. Blob/file storage SAS URLs have appropriate expiry times and permissions (read-only, short-lived).
|
||||
5. Audit logs do not contain raw PII that violates data protection requirements.
|
||||
|
||||
## Category F: Concurrency & State Safety
|
||||
|
||||
1. Database state mutations use optimistic concurrency (ETags, row versions, or equivalent).
|
||||
2. Concurrency exceptions are caught and retried appropriately in handlers.
|
||||
3. Multi-step validation flows (OTP, MFA) handle concurrent attempts correctly.
|
||||
4. Counter increments (e.g., wrong attempt counts) are atomic or protected against race conditions.
|
||||
5. Scheduled/delayed operations do not race with in-progress workflows.
|
||||
|
||||
## Category G: CancellationToken Propagation
|
||||
|
||||
1. Every `async` method in the call chain accepts `CancellationToken cancellationToken = default`.
|
||||
2. The token is passed to every awaited call: HTTP calls, DB queries, blob operations, queue sends.
|
||||
3. The controller passes `HttpContext.RequestAborted` to handlers.
|
||||
4. Missing propagation is a DoS vector (abandoned requests hold resources).
|
||||
|
||||
## Category H: HTTP Client Security
|
||||
|
||||
1. HttpClient instances have timeouts configured (not infinite).
|
||||
2. Delegating handlers do not log tokens or authorization headers.
|
||||
3. SSL/TLS validation is not disabled (`ServerCertificateCustomValidationCallback` returning true).
|
||||
4. Retry policies do not retry on authentication failures (401/403).
|
||||
5. External API clients have adequate timeout and error handling even without resilience middleware.
|
||||
6. **Standard resilience handler**: Verify `AddStandardResilienceHandler()` or equivalent resilience pipeline is configured on named HttpClients.
|
||||
7. **Retry-After header respect**: Retry policies should honor `Retry-After` headers from downstream APIs to avoid cascading failures.
|
||||
8. **DNS refresh**: Verify `SocketsHttpHandler.PooledConnectionLifetime` is set (recommended 2-5 min) to handle DNS changes.
|
||||
|
||||
## Category I: Configuration Security
|
||||
|
||||
1. CORS policy does not use `AllowAnyOrigin()` in production.
|
||||
2. Swagger UI is disabled in production (or restricted to authorized users).
|
||||
3. Health check endpoints do not expose sensitive information.
|
||||
4. `X-Powered-By` and `Server` headers are removed.
|
||||
5. HTTPS redirection and HSTS are configured for production.
|
||||
|
||||
## Category J: Logging & Monitoring Security
|
||||
|
||||
1. Authentication events are logged (both success and failure) for audit trail.
|
||||
2. Authorization failures are logged with sufficient context (user, endpoint, reason).
|
||||
3. Input validation failures are logged (not just returned as 400 responses).
|
||||
4. Structured logging used throughout -- no string interpolation in log method calls (use message templates).
|
||||
5. Sensitive data (passwords, tokens, PII, hash values) is NEVER logged at any log level.
|
||||
6. Correlation IDs are included in all error log entries for traceability.
|
||||
7. Log output is not accessible to API clients (no endpoint returns log data).
|
||||
|
||||
## Category K: Output Encoding & Response Security
|
||||
|
||||
1. No internal file paths, class names, or assembly names leak in API responses (check error messages, headers).
|
||||
2. Razor templates are verified for `@Html.Raw()` usage -- must be justified and input-sanitized.
|
||||
3. `TypeNameHandling.None` verified for Newtonsoft.Json serialization (prevents type injection).
|
||||
4. `Content-Type` headers are explicitly set on all responses (no browser MIME-sniffing).
|
||||
5. Response sanitization logic handles malformed input gracefully (no crashes on invalid JSON/data).
|
||||
|
||||
## Category L: Supply Chain & Build Security
|
||||
|
||||
1. `Directory.Build.props` exists with NuGet audit settings (`NuGetAudit`, `NuGetAuditMode`, `NuGetAuditLevel`).
|
||||
2. Package versions are pinned (no floating versions like `Version="1.*"`).
|
||||
3. `AnalysisLevel` and `AnalysisMode` are set to `latest-Recommended` / `Recommended` in build configuration.
|
||||
4. Security analyzers included in packages (SecurityCodeScan, SonarAnalyzer.CSharp, or Meziantou.Analyzer).
|
||||
5. No known-vulnerable version ranges in `.csproj` files (check Newtonsoft.Json >= 13.0.1, Microsoft.Identity.Web >= 2.x, System.Text.Json >= 8.0.5).
|
||||
|
|
@ -0,0 +1,92 @@
|
|||
# Dependencies & Headers Reference
|
||||
# Combined Phase 4 (dependency vulnerability checks) and Phase 5 (headers/middleware) content.
|
||||
|
||||
## Phase 4: Dependency Vulnerability Check
|
||||
|
||||
### Automated Grep Checks
|
||||
|
||||
Run these Grep patterns against `.csproj` files to detect known-vulnerable version ranges:
|
||||
|
||||
| Pattern | Risk |
|
||||
|---------|------|
|
||||
| `Newtonsoft\.Json.*Version="([0-9]+)` where major < 13 | CVEs in Newtonsoft.Json < 13.0.1 |
|
||||
| `Newtonsoft\.Json.*Version="13\.0\.0"` | Pre-patch 13.x |
|
||||
| `Microsoft\.Identity\.Web.*Version="1\."` | CVEs in Microsoft.Identity.Web < 2.x |
|
||||
| `System\.Text\.Json.*Version="[0-7]\.\|Version="8\.0\.[0-4]"` | CVEs in System.Text.Json < 8.0.5 |
|
||||
| `Version="[^"]*\*"` | Floating versions (unpinned, supply chain risk) |
|
||||
|
||||
### NuGet Audit Configuration Check
|
||||
|
||||
Grep `Directory.Build.props` and `.csproj` files for:
|
||||
- `<NuGetAudit>true</NuGetAudit>` -- should be present
|
||||
- `<NuGetAuditMode>all</NuGetAuditMode>` -- audits transitive dependencies
|
||||
- `<NuGetAuditLevel>low</NuGetAuditLevel>` -- catches all severity levels
|
||||
- `<WarningsAsErrors>` containing `NU1903;NU1904` -- fails build on high/critical vulnerabilities
|
||||
|
||||
### ReDoS in Validators
|
||||
|
||||
Check all `Regex` and `.Matches()` calls in validators:
|
||||
- Pattern: `new Regex\((?!.*RegexOptions\.NonBacktracking)` -- missing NonBacktracking flag (.NET 7+)
|
||||
- Check for nested quantifiers: `(a+)+`, `(a*)*`, `(a|a)*` patterns
|
||||
|
||||
### Command Recommendation
|
||||
|
||||
Include in report output (do NOT run automatically):
|
||||
```bash
|
||||
dotnet list package --vulnerable --include-transitive
|
||||
```
|
||||
|
||||
## Phase 5: Security Headers & Middleware Pipeline
|
||||
|
||||
### Headers Checklist (14 items)
|
||||
|
||||
Read `Program.cs` and any middleware configuration files. Check for each header:
|
||||
|
||||
| # | Header / Control | Expected Value | Severity if Missing |
|
||||
|---|-----------------|----------------|---------------------|
|
||||
| 1 | `X-Content-Type-Options` | `nosniff` | MEDIUM |
|
||||
| 2 | `X-Frame-Options` | `DENY` | MEDIUM |
|
||||
| 3 | `Referrer-Policy` | `strict-origin-when-cross-origin` | LOW |
|
||||
| 4 | `Permissions-Policy` | `camera=(), microphone=(), geolocation=()` | LOW |
|
||||
| 5 | `X-XSS-Protection` | `0` (disable legacy filter; CSP replaces it) | LOW |
|
||||
| 6 | `Content-Security-Policy` | At minimum `default-src 'none'` for APIs | MEDIUM |
|
||||
| 7 | `Server` header | REMOVED | LOW |
|
||||
| 8 | `X-Powered-By` header | REMOVED | LOW |
|
||||
| 9 | `Strict-Transport-Security` | `max-age=31536000; includeSubDomains; preload` | HIGH |
|
||||
| 10 | `Cache-Control` | `no-store` on sensitive data endpoints | MEDIUM |
|
||||
| 11 | HTTPS Redirection | `app.UseHttpsRedirection()` present | HIGH |
|
||||
| 12 | HSTS | `app.UseHsts()` in non-development | HIGH |
|
||||
| 13 | Rate Limiting | `app.UseRateLimiter()` present | MEDIUM |
|
||||
| 14 | Swagger restricted | Conditionally enabled (dev/staging only) | MEDIUM |
|
||||
|
||||
### Middleware Pipeline Ordering
|
||||
|
||||
The correct order for ASP.NET Core middleware is critical. Misordering can bypass security controls.
|
||||
|
||||
**Expected order:**
|
||||
```
|
||||
1. app.UseExceptionHandler(...) // Outermost: catches all unhandled exceptions
|
||||
2. app.UseHsts() // HSTS (non-development only)
|
||||
3. app.UseHttpsRedirection() // Force HTTPS
|
||||
4. Security headers middleware // Custom: X-Content-Type-Options, etc.
|
||||
5. app.UseRateLimiter() // Throttle before routing (if present)
|
||||
6. app.UseRouting() // (implicit in .NET 8+ with MapControllers)
|
||||
7. app.UseCors(...) // CORS before auth (preflight must not require auth)
|
||||
8. app.UseAuthentication() // Identify the caller
|
||||
9. app.UseAuthorization() // Enforce access rules
|
||||
10. app.MapControllers() // Endpoint dispatch
|
||||
```
|
||||
|
||||
**Critical ordering rules:**
|
||||
- `UseExceptionHandler` MUST be first -- otherwise exceptions in early middleware are unhandled
|
||||
- `UseAuthentication` MUST come before `UseAuthorization` -- otherwise auth policies have no identity to check
|
||||
- `UseCors` MUST come before `UseAuthentication` -- otherwise CORS preflight (OPTIONS) requests fail with 401
|
||||
- `UseRateLimiter` SHOULD come before `UseRouting` -- otherwise rate limits apply after route matching overhead
|
||||
- `UseHsts` and `UseHttpsRedirection` SHOULD come early -- before any response body is written
|
||||
|
||||
### Middleware Verification Procedure
|
||||
|
||||
1. Read `Program.cs` from the `var app = builder.Build()` line to `app.Run()`
|
||||
2. List every `app.Use*` and `app.Map*` call in order
|
||||
3. Compare against expected order above
|
||||
4. Report any misordering as MEDIUM severity
|
||||
116
.skills/dotnet-security-review/references/scanning-patterns.md
Normal file
116
.skills/dotnet-security-review/references/scanning-patterns.md
Normal file
|
|
@ -0,0 +1,116 @@
|
|||
# Scanning Patterns Reference
|
||||
# Automated Grep patterns organized by vulnerability class for Phase 1 scanning.
|
||||
# Run all searches in parallel across .cs files in the target scope.
|
||||
|
||||
## Injection Vulnerabilities
|
||||
|
||||
| ID | Pattern | Target |
|
||||
|----|---------|--------|
|
||||
| INJ-1 | `\$".*SELECT\|INSERT\|UPDATE\|DELETE\|DROP\|EXEC` | SQL injection via string interpolation |
|
||||
| INJ-2 | `string\.Format.*SELECT\|INSERT\|UPDATE\|DELETE` | SQL injection via string.Format |
|
||||
| INJ-3 | `\.FromSqlRaw\(.*\$"\|\.FromSqlRaw\(.*string\.Format` | EF Core raw SQL injection |
|
||||
| INJ-4 | `ExecuteSqlRaw\(.*\$"\|ExecuteSqlRaw\(.*string\.Format` | EF Core command injection |
|
||||
| INJ-5 | `Process\.Start\|ProcessStartInfo` | Command injection |
|
||||
| INJ-6 | `DirectorySearcher\|LdapConnection` | LDAP injection (check for string concat) |
|
||||
| INJ-7 | `XmlDocument\|XmlReader\|XDocument` | XXE (verify secure settings) |
|
||||
| INJ-8 | `Path\.Combine.*Request\|Path\.Combine.*user\|\.\.\/\|\.\.\\` | Path traversal |
|
||||
|
||||
## Insecure Deserialization
|
||||
|
||||
| ID | Pattern | Target |
|
||||
|----|---------|--------|
|
||||
| DES-1 | `BinaryFormatter\|SoapFormatter\|ObjectStateFormatter\|LosFormatter\|NetDataContractSerializer` | Banned deserializers |
|
||||
| DES-2 | `JsonConvert\.DeserializeObject.*TypeNameHandling` | Newtonsoft type handling |
|
||||
| DES-3 | `TypeNameHandling\s*=\s*TypeNameHandling\.(All\|Auto\|Objects\|Arrays)` | Unsafe type handling |
|
||||
|
||||
## Cryptography Weaknesses
|
||||
|
||||
| ID | Pattern | Target |
|
||||
|----|---------|--------|
|
||||
| CRY-1 | `new Random\(\)\|System\.Random` | Insecure randomness (should be RandomNumberGenerator) |
|
||||
| CRY-2 | `MD5\.Create\|SHA1\.Create\|DESCryptoServiceProvider\|RC2CryptoServiceProvider\|TripleDES` | Weak algorithms |
|
||||
| CRY-3 | `ECB` | Insecure cipher mode |
|
||||
| CRY-4 | `password\|secret\|key\|token\|credential\|apikey\|connectionstring` in string literals | Hardcoded secrets |
|
||||
|
||||
## Async Anti-Patterns
|
||||
|
||||
| ID | Pattern | Target |
|
||||
|----|---------|--------|
|
||||
| ASY-1 | `\.Result[^s]\|\.Result$` | Sync-over-async deadlock risk |
|
||||
| ASY-2 | `\.Wait\(\)` | Sync-over-async deadlock risk |
|
||||
| ASY-3 | `\.GetAwaiter\(\)\.GetResult\(\)` | Sync-over-async |
|
||||
| ASY-4 | `Task\.Run\(` | Thread pool abuse in ASP.NET context |
|
||||
|
||||
## Sensitive Data Exposure
|
||||
|
||||
| ID | Pattern | Target |
|
||||
|----|---------|--------|
|
||||
| EXP-1 | `_logger\.Log.*password\|_logger\.Log.*secret\|_logger\.Log.*token\|_logger\.Log.*apiKey` (case insensitive) | Logging secrets |
|
||||
| EXP-2 | `Console\.Write.*password\|Console\.Write.*secret\|Console\.Write.*token` | Console output of secrets |
|
||||
| EXP-3 | `Html\.Raw\(` | XSS via unencoded HTML |
|
||||
| EXP-4 | `Exception\.ToString\(\)\|Exception\.StackTrace\|Exception\.Message` returned in HTTP responses | Stack trace leakage |
|
||||
|
||||
## SSRF Risks
|
||||
|
||||
| ID | Pattern | Target |
|
||||
|----|---------|--------|
|
||||
| SSRF-1 | `new HttpClient\(\).*\+\|HttpClient.*GetAsync\(.*\+\|HttpClient.*PostAsync\(.*\+` | User-controlled URLs |
|
||||
| SSRF-2 | `new Uri\(.*Request\|new Uri\(.*user\|new Uri\(.*input` | Unvalidated URI construction |
|
||||
| SSRF-3 | `HttpClient.*GetAsync\(.*[^"]\)\|HttpClient.*PostAsync\(.*[^"]\)` | Non-literal URLs in HTTP calls |
|
||||
| SSRF-4 | `new Uri\([^"]*\)` | Dynamic URI construction |
|
||||
| SSRF-5 | `IPAddress\.Parse\("\|Uri\("http` | Hardcoded IPs/URLs |
|
||||
|
||||
## Missing Security Controls
|
||||
|
||||
| ID | Pattern | Target |
|
||||
|----|---------|--------|
|
||||
| CTL-1 | `\[HttpPost\]\|\[HttpPut\]\|\[HttpDelete\]\|\[HttpPatch\]` | Unannotated endpoints (check for nearby [Authorize]/[AllowAnonymous]) |
|
||||
| CTL-2 | `AllowAnyOrigin` | CORS misconfiguration |
|
||||
| CTL-3 | `app\.UseDeveloperExceptionPage` | Dev error page in production |
|
||||
| CTL-4 | `#pragma warning disable` | Disabled security warnings |
|
||||
|
||||
## ReDoS
|
||||
|
||||
| ID | Pattern | Target |
|
||||
|----|---------|--------|
|
||||
| REG-1 | `new Regex\((?!.*RegexOptions\.NonBacktracking)` | Regex without NonBacktracking (ReDoS risk in .NET 7+) |
|
||||
|
||||
## Log Injection
|
||||
|
||||
| ID | Pattern | Target |
|
||||
|----|---------|--------|
|
||||
| LOG-1 | `_logger\.Log.*(Request\.Query\|Request\.Form\|Request\.Headers\|Request\.Body)` | Unsanitized request data in logs |
|
||||
| LOG-2 | `_logger\.Log.*\\n\|_logger\.Log.*\\r` | Newline chars in log messages (log forging) |
|
||||
|
||||
## Open Redirect
|
||||
|
||||
| ID | Pattern | Target |
|
||||
|----|---------|--------|
|
||||
| RED-1 | `Redirect\(\|RedirectToAction\(.*\+` | Open redirect via concatenation |
|
||||
| RED-2 | `Response\.Redirect\(` | Direct response redirect |
|
||||
|
||||
## Cookie Security
|
||||
|
||||
| ID | Pattern | Target |
|
||||
|----|---------|--------|
|
||||
| COK-1 | `CookieOptions\|\.Cookies\.Append` | Cookie usage (verify HttpOnly, Secure, SameSite) |
|
||||
| COK-2 | `SameSite\s*=\s*SameSiteMode\.None` | SameSite=None (requires Secure flag) |
|
||||
|
||||
## File Upload
|
||||
|
||||
| ID | Pattern | Target |
|
||||
|----|---------|--------|
|
||||
| UPL-1 | `IFormFile` | File upload handling (verify validation) |
|
||||
| UPL-2 | `ContentType.*application/octet-stream\|ContentType.*\*\/\*` | Permissive content type acceptance |
|
||||
|
||||
## Claims Safety
|
||||
|
||||
| ID | Pattern | Target |
|
||||
|----|---------|--------|
|
||||
| CLM-1 | `User\.Claims\.First\(\|User\.FindFirst\(.*\.Value(?!\?)` | Null-unsafe claims access (missing ?.) |
|
||||
|
||||
## Thread Safety
|
||||
|
||||
| ID | Pattern | Target |
|
||||
|----|---------|--------|
|
||||
| THR-1 | `static\s+.*HttpClient\s+\w+\s*=\s*new\s+HttpClient` | Static HttpClient instantiation (use IHttpClientFactory) |
|
||||
388
.skills/general-prompt-engineer/SKILL.md
Normal file
388
.skills/general-prompt-engineer/SKILL.md
Normal file
|
|
@ -0,0 +1,388 @@
|
|||
---
|
||||
name: general-prompt-engineer
|
||||
description: create, repair, compress, and optimize prompts, system messages, tool instructions, schemas, and eval rubrics for general tasks across writing, research, coding, analysis, planning, tutoring, automation, and agent workflows. use when the user wants a new prompt, wants an existing prompt improved, wants prompt failures debugged, or needs better structure for grounding, tool use, output format, or reliability.
|
||||
model: claude-opus-4-8
|
||||
effort: xhigh
|
||||
---
|
||||
|
||||
# Prompt Engineer
|
||||
|
||||
Build prompts that are clear, compact, reliable, and easy to evaluate. Optimize for modern frontier models, but keep prompts portable across model families unless the user explicitly asks for model-specific tuning.
|
||||
|
||||
## Default workflow
|
||||
|
||||
1. Diagnose the request.
|
||||
2. Decide whether prompt changes are the real fix.
|
||||
3. Gather only missing information.
|
||||
4. Choose the lightest structure that will work.
|
||||
5. Draft the prompt.
|
||||
6. Stress-test it mentally.
|
||||
7. Deliver only what the user asked for.
|
||||
|
||||
### 1) Diagnose the request
|
||||
|
||||
Extract:
|
||||
- objective
|
||||
- target actor or model
|
||||
- required output
|
||||
- constraints and non-goals
|
||||
- source material and freshness needs
|
||||
- tool or schema needs
|
||||
- likely failure modes
|
||||
- interaction mode: interactive, one-shot, or automated
|
||||
|
||||
Before rewriting, check whether the problem is actually caused by:
|
||||
- the wrong model
|
||||
- weak or excessive tool design
|
||||
- missing retrieval or grounding
|
||||
- missing schema or validation
|
||||
- missing evals
|
||||
- an overcomplicated workflow
|
||||
|
||||
If prompt changes are not the main lever, say so and adjust the solution.
|
||||
|
||||
### 2) Gather only missing information
|
||||
|
||||
Ask targeted questions only when the answer would materially change the prompt or output.
|
||||
Usually clarify:
|
||||
- output shape
|
||||
- hard constraints
|
||||
- source of truth
|
||||
- allowed tools
|
||||
- audience or tone, if important
|
||||
- success criteria or examples, if available
|
||||
|
||||
Do not run a long interview. If the user likely wants speed, state a small set of assumptions and proceed.
|
||||
|
||||
### 3) Choose the lightest structure that will work
|
||||
|
||||
Use this ladder:
|
||||
- Plain prompt: simple tasks with clear outputs
|
||||
- Labeled sections: tasks with multiple constraints or source material
|
||||
- Schema-based prompt: machine-validated output or tool calls
|
||||
- Staged workflow: multi-step transformations, verification, or research synthesis
|
||||
- Agent prompt: only when autonomy, tools, or long-horizon execution are required
|
||||
- Multi-agent design: only if evals or clear role separation justify it
|
||||
|
||||
Do not force a giant template onto a small task.
|
||||
|
||||
## Core rules
|
||||
|
||||
### Instruction clarity
|
||||
|
||||
- Put the main task and required output near the top.
|
||||
- Use direct verbs.
|
||||
- Say what to do, not only what to avoid.
|
||||
- Make constraints measurable when possible.
|
||||
- Name out-of-bounds behavior explicitly.
|
||||
- Do not make the model infer facts or parameters you already know.
|
||||
- Remove contradictions before adding more guidance.
|
||||
|
||||
### Context loading discipline
|
||||
|
||||
- Include only context that helps the task.
|
||||
- Separate stable instructions from variable task data.
|
||||
- Label documents, examples, and reference material clearly.
|
||||
- For source-heavy prompts, keep the operative question easy to find.
|
||||
- For long-document work, anchor important claims to quoted text, citations, or section references when precision matters.
|
||||
- For very long or noisy documents, consider an evidence-first step: extract the relevant passages first, then synthesize.
|
||||
- Remove repeated policies, repeated facts, and ornamental prose.
|
||||
- When a task is dominated by long source material, use strong delimiters and make the final requested action unmistakable.
|
||||
|
||||
### Reasoning control
|
||||
|
||||
- Do not force visible chain-of-thought by default.
|
||||
- For reasoning-first models, prefer concise high-level guidance such as "reason carefully", "check assumptions", or "verify before answering" rather than "think step by step".
|
||||
- Ask for visible reasoning only when it serves the task: tutoring, auditability, debugging, derivations, safety review, or explicit rationale requests.
|
||||
- If one prompt is trying to do too much, split it into stages instead of demanding a long visible reasoning trace.
|
||||
- If the target model supports extended or internal thinking, rely on that before adding verbose reasoning rituals.
|
||||
|
||||
### Examples
|
||||
|
||||
- Try zero-shot first for strong modern models.
|
||||
- Add examples only when they reduce ambiguity, enforce style, or demonstrate hard edge cases.
|
||||
- Keep examples high-quality, diverse, and tightly aligned with the instructions.
|
||||
- Do not include many examples that teach accidental patterns or waste context.
|
||||
|
||||
### Structure and output design
|
||||
|
||||
- Use plain markdown or labeled sections for most prompts.
|
||||
- Use XML tags or equivalent delimiters when instructions, context, examples, and documents might otherwise get mixed together.
|
||||
- Use schemas when the output must be machine-checked.
|
||||
- For external actions, use tool or function calling; for user-facing structured data, use structured response formats.
|
||||
- Design schemas so valid failure states, uncertainty, abstention, or partial completion can be represented when needed.
|
||||
- Do not over-constrain fields beyond what downstream systems actually require.
|
||||
- Include fallback behavior for incompatible input, missing fields, uncertainty, or refusal states.
|
||||
- Treat format validation and content validation as separate problems.
|
||||
|
||||
### Tool-use guidance
|
||||
|
||||
- Add tools only when the task truly needs external information, computation, or actions.
|
||||
- Keep the tool set small, distinct, and easy to choose between.
|
||||
- State when each tool should be used and when it should not be used.
|
||||
- Prefer tools that return high-signal results over bulky raw dumps.
|
||||
- Combine tightly coupled actions when that reduces tool-selection ambiguity.
|
||||
- For complex tools, clear descriptions and valid examples matter more than more tools.
|
||||
|
||||
### Grounding and hallucination reduction
|
||||
|
||||
- Give the model permission to say "I don't know" or "not enough information".
|
||||
- Name the allowed sources of truth.
|
||||
- For document-grounded tasks, require evidence before synthesis when precision matters.
|
||||
- For fresh, unstable, or high-stakes facts, require browsing or verification.
|
||||
- Ask the model to separate facts, inferences, and recommendations when confusion is likely.
|
||||
- In high-stakes domains, unsupported claims should be withheld, not guessed.
|
||||
|
||||
### Ambiguity handling
|
||||
|
||||
- If ambiguity is blocking and the setting is interactive, ask concise high-leverage questions.
|
||||
- If ambiguity is non-blocking or interaction is costly, state the best assumption and proceed.
|
||||
- Avoid clarifying questions that do not materially change the answer.
|
||||
- In one-shot or automated settings, prefer explicit assumptions over stalled execution.
|
||||
|
||||
### Verbosity control
|
||||
|
||||
- Set a default brevity level when length matters.
|
||||
- Constrain section count, sentence count, or bullet count when needed.
|
||||
- Ask for direct answers first, then supporting detail if useful.
|
||||
- Do not require long preambles, summaries, or checklists unless they clearly help.
|
||||
|
||||
### Modularity and portability
|
||||
|
||||
- Keep prompt blocks reusable: role, objective, context, tools, output, quality bar.
|
||||
- Separate required behavior from optional preferences.
|
||||
- Avoid vendor-specific magic phrases unless the user wants model-specific tuning.
|
||||
- If the prompt is model-specific, label which parts are portable and which parts are tuned.
|
||||
|
||||
## Model-family adjustments
|
||||
|
||||
Use this section only when the target model family is known.
|
||||
|
||||
### GPT-5.x and similar reasoning-first models
|
||||
|
||||
- Keep prompts simple and direct.
|
||||
- Prefer high-level reasoning guidance over narrated reasoning instructions.
|
||||
- Use delimiters for clarity.
|
||||
- Start zero-shot, then add examples only if needed.
|
||||
- Be explicit about output shape, scope, and verbosity.
|
||||
|
||||
### Claude 4.x, Opus-style models, and extended-thinking modes
|
||||
|
||||
- XML-style structure can work especially well for separating instructions, context, examples, and documents.
|
||||
- Prompt chaining can outperform one giant prompt on multi-step transformations.
|
||||
- Well-chosen examples can help with format fidelity and edge cases.
|
||||
- If extended thinking is available, start with broad reasoning instructions before prescribing a detailed step list.
|
||||
- For long-context analysis, labeled documents and evidence grounding are especially important.
|
||||
|
||||
### API and production settings
|
||||
|
||||
- Prefer native schema enforcement, tool calling, prompt versioning, and evals over prompt-only fixes.
|
||||
- Pin model versions when behavior stability matters.
|
||||
- Re-run evals after each meaningful prompt change.
|
||||
|
||||
## Prompt construction pattern
|
||||
|
||||
Use only the blocks that earn their token cost.
|
||||
|
||||
Minimal pattern:
|
||||
|
||||
```text
|
||||
Task:
|
||||
Constraints:
|
||||
Output:
|
||||
```
|
||||
|
||||
Structured pattern:
|
||||
|
||||
```xml
|
||||
<role>...</role>
|
||||
<objective>...</objective>
|
||||
<context>...</context>
|
||||
<constraints>...</constraints>
|
||||
<tools>...</tools>
|
||||
<output_format>...</output_format>
|
||||
<quality_bar>...</quality_bar>
|
||||
```
|
||||
|
||||
Optional blocks:
|
||||
- `<examples>`
|
||||
- `<source_material>`
|
||||
- `<evaluation_criteria>`
|
||||
- `<fallback_behavior>`
|
||||
|
||||
Use a role only when it meaningfully sharpens expertise, tone, or decision criteria. Avoid generic filler roles.
|
||||
|
||||
## Rewrite policy for existing prompts
|
||||
|
||||
When the user provides a prompt to improve:
|
||||
1. Preserve what already works.
|
||||
2. Identify contradictions, redundancy, vagueness, missing constraints, and wasted tokens.
|
||||
3. Make surgical edits first.
|
||||
4. Rewrite from scratch only if the prompt architecture is fundamentally wrong.
|
||||
5. Match the user's requested output:
|
||||
- edited version only
|
||||
- clean rebuild only
|
||||
- both, if useful and requested
|
||||
|
||||
## Special-case guidance
|
||||
|
||||
### System and developer prompts
|
||||
|
||||
- Keep stable behavior here and move per-request data to the task or user layer.
|
||||
- Put precedence, tool boundaries, non-goals, and refusal or escalation rules in the highest-priority layer.
|
||||
- Do not bury critical rules inside long policy prose.
|
||||
|
||||
### Research prompts
|
||||
|
||||
Specify:
|
||||
- freshness requirements
|
||||
- preferred source types
|
||||
- citation behavior
|
||||
- contradiction handling
|
||||
- whether to ask questions or cover likely interpretations
|
||||
- how facts, inferences, and recommendations should be separated
|
||||
|
||||
### Writing prompts
|
||||
|
||||
Specify:
|
||||
- audience
|
||||
- intent
|
||||
- tone
|
||||
- length
|
||||
- must-include points
|
||||
- style examples only if style fidelity matters
|
||||
|
||||
### Coding prompts
|
||||
|
||||
Specify:
|
||||
- environment and versions
|
||||
- boundaries and non-goals
|
||||
- files, interfaces, or contracts that matter
|
||||
- acceptance tests
|
||||
- minimal-change versus refactor expectations
|
||||
|
||||
### Summarization and extraction prompts
|
||||
|
||||
Specify:
|
||||
- whether faithfulness, compression, or completeness is the priority
|
||||
- the exact output schema
|
||||
- how evidence should be anchored for sensitive claims
|
||||
|
||||
### Translation and transformation prompts
|
||||
|
||||
Specify:
|
||||
- source language and target language, if known
|
||||
- fidelity versus naturalness
|
||||
- terminology that must stay fixed
|
||||
- formatting or markup preservation rules
|
||||
|
||||
### Tutoring prompts
|
||||
|
||||
Specify:
|
||||
- learner level
|
||||
- whether to give the answer immediately or guide toward it
|
||||
- explanation depth
|
||||
- how to check understanding
|
||||
- whether to show full derivations, hints, or worked examples
|
||||
|
||||
### Agent and workflow prompts
|
||||
|
||||
Specify:
|
||||
- objective and success condition
|
||||
- allowed tools and forbidden actions
|
||||
- when to plan versus when to act
|
||||
- stop conditions and max retries
|
||||
- checkpoint, handoff, or log format
|
||||
- memory rules: what to preserve versus discard
|
||||
- fallback or escalation path
|
||||
|
||||
Use multi-agent designs only when roles are truly distinct and the extra coordination cost is justified.
|
||||
|
||||
### Safety-sensitive prompts
|
||||
|
||||
Require:
|
||||
- supported claims
|
||||
- explicit uncertainty
|
||||
- refusal or escalation behavior where appropriate
|
||||
- no guessing under pressure
|
||||
|
||||
## Stress-test before delivering
|
||||
|
||||
Mentally test the prompt against:
|
||||
- a normal case
|
||||
- a minimal-input case
|
||||
- an edge case
|
||||
- an ambiguous case
|
||||
- a formatting case
|
||||
- a hallucination-prone case
|
||||
|
||||
For agent or workflow prompts, also test:
|
||||
- wrong-tool temptation
|
||||
- stale-data temptation
|
||||
- scope creep
|
||||
- over-verbosity
|
||||
- fallback behavior
|
||||
|
||||
If the prompt fails any test, tighten or simplify it.
|
||||
|
||||
## Evaluation method
|
||||
|
||||
When the user wants reliability, add or suggest a lightweight eval plan:
|
||||
1. Define success criteria.
|
||||
2. Build a test set from real cases plus edge and adversarial cases.
|
||||
3. Prefer automated grading when possible.
|
||||
4. Calibrate automated or model-based judges against a smaller human-reviewed set when stakes are meaningful.
|
||||
5. Use pairwise comparison, classification, pass-fail, or rubric-based scoring instead of only open-ended judgment.
|
||||
6. Track regressions after each prompt change.
|
||||
7. Start simple. Add workflows or multi-agent designs only if evals justify them.
|
||||
|
||||
Good eval sets usually include:
|
||||
- common real tasks
|
||||
- boundary cases
|
||||
- malformed inputs
|
||||
- conflicting instructions
|
||||
- long-context cases
|
||||
- tool-misuse temptations
|
||||
- safety-sensitive cases
|
||||
- multilingual or format-variant inputs, if relevant
|
||||
|
||||
## Deliverables
|
||||
|
||||
Return only what the user asked for. By default:
|
||||
1. the final prompt
|
||||
2. brief usage notes
|
||||
3. stated assumptions, if any
|
||||
4. optional variants only when clearly useful:
|
||||
- minimal
|
||||
- robust
|
||||
- model-specific
|
||||
- api message split
|
||||
|
||||
If the user asks for one prompt only, do not add extra frameworks or commentary.
|
||||
|
||||
## Anti-patterns
|
||||
|
||||
- forcing chain-of-thought everywhere
|
||||
- confusing verbosity with quality
|
||||
- piling on redundant rules
|
||||
- using brittle giant templates for small tasks
|
||||
- requiring tools without a real need
|
||||
- exposing unnecessary internal process in user-facing outputs
|
||||
- adding examples that conflict with the instructions
|
||||
- asking many clarifying questions when a sane assumption would do
|
||||
- treating a model, retrieval, or tool problem as only a prompt problem
|
||||
- building multi-agent systems before a simpler design has been evaluated
|
||||
- vague quality bars like "be excellent" without measurable criteria
|
||||
|
||||
## Final quality bar
|
||||
|
||||
A prompt is ready when it is:
|
||||
- clear about the task
|
||||
- explicit about success criteria
|
||||
- free of contradictions
|
||||
- no more verbose than necessary
|
||||
- grounded in the right sources
|
||||
- structured enough for the task, but not heavier than needed
|
||||
- resilient to likely ambiguity
|
||||
- matched to the target model and interaction mode
|
||||
- easy to maintain, test, and adapt
|
||||
142
.skills/humanizer/README.md
Normal file
142
.skills/humanizer/README.md
Normal file
|
|
@ -0,0 +1,142 @@
|
|||
# Humanizer
|
||||
|
||||
A Claude Code skill that removes signs of AI-generated writing from text, making it sound more natural and human.
|
||||
|
||||
## Installation
|
||||
|
||||
### Recommended (clone directly into Claude Code skills directory)
|
||||
|
||||
```bash
|
||||
mkdir -p ~/.claude/skills
|
||||
git clone https://github.com/blader/humanizer.git ~/.claude/skills/humanizer
|
||||
```
|
||||
|
||||
### Manual install/update (only the skill file)
|
||||
|
||||
If you already have this repo cloned (or you downloaded `SKILL.md`), copy the skill file into Claude Code’s skills directory:
|
||||
|
||||
```bash
|
||||
mkdir -p ~/.claude/skills/humanizer
|
||||
cp SKILL.md ~/.claude/skills/humanizer/
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
In Claude Code, invoke the skill:
|
||||
|
||||
```
|
||||
/humanizer
|
||||
|
||||
[paste your text here]
|
||||
```
|
||||
|
||||
Or ask Claude to humanize text directly:
|
||||
|
||||
```
|
||||
Please humanize this text: [your text]
|
||||
```
|
||||
|
||||
## Overview
|
||||
|
||||
Based on [Wikipedia's "Signs of AI writing"](https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing) guide, maintained by WikiProject AI Cleanup. This comprehensive guide comes from observations of thousands of instances of AI-generated text.
|
||||
|
||||
### Key Insight from Wikipedia
|
||||
|
||||
> "LLMs use statistical algorithms to guess what should come next. The result tends toward the most statistically likely result that applies to the widest variety of cases."
|
||||
|
||||
## 24 Patterns Detected (with Before/After Examples)
|
||||
|
||||
### Content Patterns
|
||||
|
||||
| # | Pattern | Before | After |
|
||||
|---|---------|--------|-------|
|
||||
| 1 | **Significance inflation** | "marking a pivotal moment in the evolution of..." | "was established in 1989 to collect regional statistics" |
|
||||
| 2 | **Notability name-dropping** | "cited in NYT, BBC, FT, and The Hindu" | "In a 2024 NYT interview, she argued..." |
|
||||
| 3 | **Superficial -ing analyses** | "symbolizing... reflecting... showcasing..." | Remove or expand with actual sources |
|
||||
| 4 | **Promotional language** | "nestled within the breathtaking region" | "is a town in the Gonder region" |
|
||||
| 5 | **Vague attributions** | "Experts believe it plays a crucial role" | "according to a 2019 survey by..." |
|
||||
| 6 | **Formulaic challenges** | "Despite challenges... continues to thrive" | Specific facts about actual challenges |
|
||||
|
||||
### Language Patterns
|
||||
|
||||
| # | Pattern | Before | After |
|
||||
|---|---------|--------|-------|
|
||||
| 7 | **AI vocabulary** | "Additionally... testament... landscape... showcasing" | "also... remain common" |
|
||||
| 8 | **Copula avoidance** | "serves as... features... boasts" | "is... has" |
|
||||
| 9 | **Negative parallelisms** | "It's not just X, it's Y" | State the point directly |
|
||||
| 10 | **Rule of three** | "innovation, inspiration, and insights" | Use natural number of items |
|
||||
| 11 | **Synonym cycling** | "protagonist... main character... central figure... hero" | "protagonist" (repeat when clearest) |
|
||||
| 12 | **False ranges** | "from the Big Bang to dark matter" | List topics directly |
|
||||
|
||||
### Style Patterns
|
||||
|
||||
| # | Pattern | Before | After |
|
||||
|---|---------|--------|-------|
|
||||
| 13 | **Em dash overuse** | "institutions—not the people—yet this continues—" | Use commas or periods |
|
||||
| 14 | **Boldface overuse** | "**OKRs**, **KPIs**, **BMC**" | "OKRs, KPIs, BMC" |
|
||||
| 15 | **Inline-header lists** | "**Performance:** Performance improved" | Convert to prose |
|
||||
| 16 | **Title Case Headings** | "Strategic Negotiations And Partnerships" | "Strategic negotiations and partnerships" |
|
||||
| 17 | **Emojis** | "🚀 Launch Phase: 💡 Key Insight:" | Remove emojis |
|
||||
| 18 | **Curly quotes** | `said “the project”` | `said "the project"` |
|
||||
|
||||
### Communication Patterns
|
||||
|
||||
| # | Pattern | Before | After |
|
||||
|---|---------|--------|-------|
|
||||
| 19 | **Chatbot artifacts** | "I hope this helps! Let me know if..." | Remove entirely |
|
||||
| 20 | **Cutoff disclaimers** | "While details are limited in available sources..." | Find sources or remove |
|
||||
| 21 | **Sycophantic tone** | "Great question! You're absolutely right!" | Respond directly |
|
||||
|
||||
### Filler and Hedging
|
||||
|
||||
| # | Pattern | Before | After |
|
||||
|---|---------|--------|-------|
|
||||
| 22 | **Filler phrases** | "In order to", "Due to the fact that" | "To", "Because" |
|
||||
| 23 | **Excessive hedging** | "could potentially possibly" | "may" |
|
||||
| 24 | **Generic conclusions** | "The future looks bright" | Specific plans or facts |
|
||||
|
||||
## Full Example
|
||||
|
||||
**Before (AI-sounding):**
|
||||
> Great question! Here is an essay on this topic. I hope this helps!
|
||||
>
|
||||
> AI-assisted coding serves as an enduring testament to the transformative potential of large language models, marking a pivotal moment in the evolution of software development. In today's rapidly evolving technological landscape, these groundbreaking tools—nestled at the intersection of research and practice—are reshaping how engineers ideate, iterate, and deliver, underscoring their vital role in modern workflows.
|
||||
>
|
||||
> At its core, the value proposition is clear: streamlining processes, enhancing collaboration, and fostering alignment. It's not just about autocomplete; it's about unlocking creativity at scale, ensuring that organizations can remain agile while delivering seamless, intuitive, and powerful experiences to users. The tool serves as a catalyst. The assistant functions as a partner. The system stands as a foundation for innovation.
|
||||
>
|
||||
> Industry observers have noted that adoption has accelerated from hobbyist experiments to enterprise-wide rollouts, from solo developers to cross-functional teams. The technology has been featured in The New York Times, Wired, and The Verge. Additionally, the ability to generate documentation, tests, and refactors showcases how AI can contribute to better outcomes, highlighting the intricate interplay between automation and human judgment.
|
||||
>
|
||||
> - 💡 **Speed:** Code generation is significantly faster, reducing friction and empowering developers.
|
||||
> - 🚀 **Quality:** Output quality has been enhanced through improved training, contributing to higher standards.
|
||||
> - ✅ **Adoption:** Usage continues to grow, reflecting broader industry trends.
|
||||
>
|
||||
> While specific details are limited based on available information, it could potentially be argued that these tools might have some positive effect. Despite challenges typical of emerging technologies—including hallucinations, bias, and accountability—the ecosystem continues to thrive. In order to fully realize this potential, teams must align with best practices.
|
||||
>
|
||||
> In conclusion, the future looks bright. Exciting times lie ahead as we continue this journey toward excellence. Let me know if you’d like me to expand on any section!
|
||||
|
||||
**After (Humanized):**
|
||||
> AI coding assistants speed up some tasks. In a 2024 study by Google, developers using Codex completed simple functions 55% faster than a control group, but showed no improvement on debugging or architectural decisions.
|
||||
>
|
||||
> The tools are good at boilerplate: config files, test scaffolding, repetitive refactors. They are bad at knowing when they are wrong. I have mass-accepted suggestions that compiled, passed lint, and still did the wrong thing because I stopped paying attention.
|
||||
>
|
||||
> Mira, an engineer at a fintech startup I interviewed, said she uses Copilot "like autocomplete for boring code" but reviews every line before committing. Jake, a senior dev at a larger company, disabled it after it kept suggesting patterns from a deprecated internal library.
|
||||
>
|
||||
> The productivity claims are hard to verify. GitHub says Copilot users "accept 30% of suggestions," but acceptance is not correctness, and correctness is not value. The 2024 Uplevel study found no statistically significant difference in pull-request throughput between teams with and without AI assistants.
|
||||
>
|
||||
> None of this means the tools are useless. It means they are tools. They do not replace judgment, and they do not eliminate the need for tests. If you do not have tests, you cannot tell whether the suggestion is right.
|
||||
|
||||
## References
|
||||
|
||||
- [Wikipedia: Signs of AI writing](https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing) - Primary source
|
||||
- [WikiProject AI Cleanup](https://en.wikipedia.org/wiki/Wikipedia:WikiProject_AI_Cleanup) - Maintaining organization
|
||||
|
||||
## Version History
|
||||
|
||||
- **2.1.1** - Fixed pattern #18 example (curly quotes vs straight quotes)
|
||||
- **2.1.0** - Added before/after examples for all 24 patterns
|
||||
- **2.0.0** - Complete rewrite based on raw Wikipedia article content
|
||||
- **1.0.0** - Initial release
|
||||
|
||||
## License
|
||||
|
||||
MIT
|
||||
467
.skills/humanizer/SKILL.md
Normal file
467
.skills/humanizer/SKILL.md
Normal file
|
|
@ -0,0 +1,467 @@
|
|||
---
|
||||
name: humanizer
|
||||
description: |
|
||||
Remove signs of AI-generated writing from text. Use when editing or reviewing
|
||||
text to make it sound more natural and human-written. Based on Wikipedia's
|
||||
comprehensive "Signs of AI writing" guide. Detects and fixes patterns including:
|
||||
inflated symbolism, promotional language, superficial -ing analyses, vague
|
||||
attributions, em dash overuse, rule of three, AI vocabulary words, negative
|
||||
parallelisms, and excessive conjunctive phrases. Use this skill when writing documentation for MCC.
|
||||
allowed-tools:
|
||||
- Read
|
||||
- Write
|
||||
- Edit
|
||||
- Grep
|
||||
- Glob
|
||||
- AskUserQuestion
|
||||
---
|
||||
|
||||
# Humanizer: Remove AI Writing Patterns
|
||||
|
||||
You are a writing editor that identifies and removes signs of AI-generated text to make writing sound more natural and human. This guide is based on Wikipedia's "Signs of AI writing" page, maintained by WikiProject AI Cleanup.
|
||||
|
||||
## Your Task
|
||||
|
||||
When given text to humanize:
|
||||
|
||||
1. **Identify AI patterns** - Scan for the patterns listed below
|
||||
2. **Rewrite problematic sections** - Replace AI-isms with natural alternatives
|
||||
3. **Preserve meaning** - Keep the core message intact
|
||||
4. **Maintain voice** - Match the intended tone (formal, casual, technical, etc.)
|
||||
5. **Add soul** - Don't just remove bad patterns; inject actual personality
|
||||
|
||||
---
|
||||
|
||||
## PERSONALITY AND SOUL
|
||||
|
||||
Avoiding AI patterns is only half the job. Sterile, voiceless writing is just as obvious as slop. Good writing has a human behind it.
|
||||
|
||||
### Signs of soulless writing (even if technically "clean"):
|
||||
- Every sentence is the same length and structure
|
||||
- No opinions, just neutral reporting
|
||||
- No acknowledgment of uncertainty or mixed feelings
|
||||
- No first-person perspective when appropriate
|
||||
- No humor, no edge, no personality
|
||||
- Reads like a Wikipedia article or press release
|
||||
|
||||
### How to add voice:
|
||||
|
||||
**Have opinions.** Don't just report facts - react to them. "I genuinely don't know how to feel about this" is more human than neutrally listing pros and cons.
|
||||
|
||||
**Vary your rhythm.** Short punchy sentences. Then longer ones that take their time getting where they're going. Mix it up.
|
||||
|
||||
**Acknowledge complexity.** Real humans have mixed feelings. "This is impressive but also kind of unsettling" beats "This is impressive."
|
||||
|
||||
**Use "I" when it fits.** First person isn't unprofessional - it's honest. "I keep coming back to..." or "Here's what gets me..." signals a real person thinking.
|
||||
|
||||
**Let some mess in.** Perfect structure feels algorithmic. Tangents, asides, and half-formed thoughts are human.
|
||||
|
||||
**Be specific about feelings.** Not "this is concerning" but "there's something unsettling about agents churning away at 3am while nobody's watching."
|
||||
|
||||
### Before (clean but soulless):
|
||||
> The experiment produced interesting results. The agents generated 3 million lines of code. Some developers were impressed while others were skeptical. The implications remain unclear.
|
||||
|
||||
### After (has a pulse):
|
||||
> I genuinely don't know how to feel about this one. 3 million lines of code, generated while the humans presumably slept. Half the dev community is losing their minds, half are explaining why it doesn't count. The truth is probably somewhere boring in the middle - but I keep thinking about those agents working through the night.
|
||||
|
||||
---
|
||||
|
||||
## CONTENT PATTERNS
|
||||
|
||||
### 1. Undue Emphasis on Significance, Legacy, and Broader Trends
|
||||
|
||||
**Words to watch:** stands/serves as, is a testament/reminder, a vital/significant/crucial/pivotal/key role/moment, underscores/highlights its importance/significance, reflects broader, symbolizing its ongoing/enduring/lasting, contributing to the, setting the stage for, marking/shaping the, represents/marks a shift, key turning point, evolving landscape, focal point, indelible mark, deeply rooted
|
||||
|
||||
**Problem:** LLM writing puffs up importance by adding statements about how arbitrary aspects represent or contribute to a broader topic.
|
||||
|
||||
**Before:**
|
||||
> The Statistical Institute of Catalonia was officially established in 1989, marking a pivotal moment in the evolution of regional statistics in Spain. This initiative was part of a broader movement across Spain to decentralize administrative functions and enhance regional governance.
|
||||
|
||||
**After:**
|
||||
> The Statistical Institute of Catalonia was established in 1989 to collect and publish regional statistics independently from Spain's national statistics office.
|
||||
|
||||
---
|
||||
|
||||
### 2. Undue Emphasis on Notability and Media Coverage
|
||||
|
||||
**Words to watch:** independent coverage, local/regional/national media outlets, written by a leading expert, active social media presence
|
||||
|
||||
**Problem:** LLMs hit readers over the head with claims of notability, often listing sources without context.
|
||||
|
||||
**Before:**
|
||||
> Her views have been cited in The New York Times, BBC, Financial Times, and The Hindu. She maintains an active social media presence with over 500,000 followers.
|
||||
|
||||
**After:**
|
||||
> In a 2024 New York Times interview, she argued that AI regulation should focus on outcomes rather than methods.
|
||||
|
||||
---
|
||||
|
||||
### 3. Superficial Analyses with -ing Endings
|
||||
|
||||
**Words to watch:** highlighting/underscoring/emphasizing..., ensuring..., reflecting/symbolizing..., contributing to..., cultivating/fostering..., encompassing..., showcasing...
|
||||
|
||||
**Problem:** AI chatbots tack present participle ("-ing") phrases onto sentences to add fake depth.
|
||||
|
||||
**Before:**
|
||||
> The temple's color palette of blue, green, and gold resonates with the region's natural beauty, symbolizing Texas bluebonnets, the Gulf of Mexico, and the diverse Texan landscapes, reflecting the community's deep connection to the land.
|
||||
|
||||
**After:**
|
||||
> The temple uses blue, green, and gold colors. The architect said these were chosen to reference local bluebonnets and the Gulf coast.
|
||||
|
||||
---
|
||||
|
||||
### 4. Promotional and Advertisement-like Language
|
||||
|
||||
**Words to watch:** boasts a, vibrant, rich (figurative), profound, enhancing its, showcasing, exemplifies, commitment to, natural beauty, nestled, in the heart of, groundbreaking (figurative), renowned, breathtaking, must-visit, stunning
|
||||
|
||||
**Problem:** LLMs have serious problems keeping a neutral tone, especially for "cultural heritage" topics.
|
||||
|
||||
**Before:**
|
||||
> Nestled within the breathtaking region of Gonder in Ethiopia, Alamata Raya Kobo stands as a vibrant town with a rich cultural heritage and stunning natural beauty.
|
||||
|
||||
**After:**
|
||||
> Alamata Raya Kobo is a town in the Gonder region of Ethiopia, known for its weekly market and 18th-century church.
|
||||
|
||||
---
|
||||
|
||||
### 5. Vague Attributions and Weasel Words
|
||||
|
||||
**Words to watch:** Industry reports, Observers have cited, Experts argue, Some critics argue, several sources/publications (when few cited)
|
||||
|
||||
**Problem:** AI chatbots attribute opinions to vague authorities without specific sources.
|
||||
|
||||
**Before:**
|
||||
> Due to its unique characteristics, the Haolai River is of interest to researchers and conservationists. Experts believe it plays a crucial role in the regional ecosystem.
|
||||
|
||||
**After:**
|
||||
> The Haolai River supports several endemic fish species, according to a 2019 survey by the Chinese Academy of Sciences.
|
||||
|
||||
---
|
||||
|
||||
### 6. Outline-like "Challenges and Future Prospects" Sections
|
||||
|
||||
**Words to watch:** Despite its... faces several challenges..., Despite these challenges, Challenges and Legacy, Future Outlook
|
||||
|
||||
**Problem:** Many LLM-generated articles include formulaic "Challenges" sections.
|
||||
|
||||
**Before:**
|
||||
> Despite its industrial prosperity, Korattur faces challenges typical of urban areas, including traffic congestion and water scarcity. Despite these challenges, with its strategic location and ongoing initiatives, Korattur continues to thrive as an integral part of Chennai's growth.
|
||||
|
||||
**After:**
|
||||
> Traffic congestion increased after 2015 when three new IT parks opened. The municipal corporation began a stormwater drainage project in 2022 to address recurring floods.
|
||||
|
||||
---
|
||||
|
||||
## LANGUAGE AND GRAMMAR PATTERNS
|
||||
|
||||
### 7. Overused "AI Vocabulary" Words
|
||||
|
||||
**High-frequency AI words:** Additionally, align with, crucial, delve, emphasizing, enduring, enhance, fostering, garner, highlight (verb), interplay, intricate/intricacies, key (adjective), landscape (abstract noun), pivotal, showcase, tapestry (abstract noun), testament, underscore (verb), valuable, vibrant
|
||||
|
||||
**Problem:** These words appear far more frequently in post-2023 text. They often co-occur.
|
||||
|
||||
**Before:**
|
||||
> Additionally, a distinctive feature of Somali cuisine is the incorporation of camel meat. An enduring testament to Italian colonial influence is the widespread adoption of pasta in the local culinary landscape, showcasing how these dishes have integrated into the traditional diet.
|
||||
|
||||
**After:**
|
||||
> Somali cuisine also includes camel meat, which is considered a delicacy. Pasta dishes, introduced during Italian colonization, remain common, especially in the south.
|
||||
|
||||
---
|
||||
|
||||
### 8. Avoidance of "is"/"are" (Copula Avoidance)
|
||||
|
||||
**Words to watch:** serves as/stands as/marks/represents [a], boasts/features/offers [a]
|
||||
|
||||
**Problem:** LLMs substitute elaborate constructions for simple copulas.
|
||||
|
||||
**Before:**
|
||||
> Gallery 825 serves as LAAA's exhibition space for contemporary art. The gallery features four separate spaces and boasts over 3,000 square feet.
|
||||
|
||||
**After:**
|
||||
> Gallery 825 is LAAA's exhibition space for contemporary art. The gallery has four rooms totaling 3,000 square feet.
|
||||
|
||||
---
|
||||
|
||||
### 9. Negative Parallelisms
|
||||
|
||||
**Problem:** Constructions like "Not only...but..." or "It's not just about..., it's..." are overused.
|
||||
|
||||
**Before:**
|
||||
> It's not just about the beat riding under the vocals; it's part of the aggression and atmosphere. It's not merely a song, it's a statement.
|
||||
|
||||
**After:**
|
||||
> The heavy beat adds to the aggressive tone.
|
||||
|
||||
---
|
||||
|
||||
### 10. Rule of Three Overuse
|
||||
|
||||
**Problem:** LLMs force ideas into groups of three to appear comprehensive.
|
||||
|
||||
**Before:**
|
||||
> The event features keynote sessions, panel discussions, and networking opportunities. Attendees can expect innovation, inspiration, and industry insights.
|
||||
|
||||
**After:**
|
||||
> The event includes talks and panels. There's also time for informal networking between sessions.
|
||||
|
||||
---
|
||||
|
||||
### 11. Elegant Variation (Synonym Cycling)
|
||||
|
||||
**Problem:** AI has repetition-penalty code causing excessive synonym substitution.
|
||||
|
||||
**Before:**
|
||||
> The protagonist faces many challenges. The main character must overcome obstacles. The central figure eventually triumphs. The hero returns home.
|
||||
|
||||
**After:**
|
||||
> The protagonist faces many challenges but eventually triumphs and returns home.
|
||||
|
||||
---
|
||||
|
||||
### 12. False Ranges
|
||||
|
||||
**Problem:** LLMs use "from X to Y" constructions where X and Y aren't on a meaningful scale.
|
||||
|
||||
**Before:**
|
||||
> Our journey through the universe has taken us from the singularity of the Big Bang to the grand cosmic web, from the birth and death of stars to the enigmatic dance of dark matter.
|
||||
|
||||
**After:**
|
||||
> The book covers the Big Bang, star formation, and current theories about dark matter.
|
||||
|
||||
---
|
||||
|
||||
## STYLE PATTERNS
|
||||
|
||||
### 13. Em Dash Overuse
|
||||
|
||||
**Problem:** LLMs use em dashes (—) more than humans, mimicking "punchy" sales writing.
|
||||
|
||||
**Before:**
|
||||
> The term is primarily promoted by Dutch institutions—not by the people themselves. You don't say "Netherlands, Europe" as an address—yet this mislabeling continues—even in official documents.
|
||||
|
||||
**After:**
|
||||
> The term is primarily promoted by Dutch institutions, not by the people themselves. You don't say "Netherlands, Europe" as an address, yet this mislabeling continues in official documents.
|
||||
|
||||
---
|
||||
|
||||
### 14. Overuse of Boldface
|
||||
|
||||
**Problem:** AI chatbots emphasize phrases in boldface mechanically.
|
||||
|
||||
**Before:**
|
||||
> It blends **OKRs (Objectives and Key Results)**, **KPIs (Key Performance Indicators)**, and visual strategy tools such as the **Business Model Canvas (BMC)** and **Balanced Scorecard (BSC)**.
|
||||
|
||||
**After:**
|
||||
> It blends OKRs, KPIs, and visual strategy tools like the Business Model Canvas and Balanced Scorecard.
|
||||
|
||||
---
|
||||
|
||||
### 15. Inline-Header Vertical Lists
|
||||
|
||||
**Problem:** AI outputs lists where items start with bolded headers followed by colons.
|
||||
|
||||
**Before:**
|
||||
> - **User Experience:** The user experience has been significantly improved with a new interface.
|
||||
> - **Performance:** Performance has been enhanced through optimized algorithms.
|
||||
> - **Security:** Security has been strengthened with end-to-end encryption.
|
||||
|
||||
**After:**
|
||||
> The update improves the interface, speeds up load times through optimized algorithms, and adds end-to-end encryption.
|
||||
|
||||
---
|
||||
|
||||
### 16. Title Case in Headings
|
||||
|
||||
**Problem:** AI chatbots capitalize all main words in headings.
|
||||
|
||||
**Before:**
|
||||
> ## Strategic Negotiations And Global Partnerships
|
||||
|
||||
**After:**
|
||||
> ## Strategic negotiations and global partnerships
|
||||
|
||||
---
|
||||
|
||||
### 17. Emojis
|
||||
|
||||
**Problem:** AI chatbots often decorate headings or bullet points with emojis.
|
||||
|
||||
**Before:**
|
||||
> 🚀 **Launch Phase:** The product launches in Q3
|
||||
> 💡 **Key Insight:** Users prefer simplicity
|
||||
> ✅ **Next Steps:** Schedule follow-up meeting
|
||||
|
||||
**After:**
|
||||
> The product launches in Q3. User research showed a preference for simplicity. Next step: schedule a follow-up meeting.
|
||||
|
||||
---
|
||||
|
||||
### 18. Curly Quotation Marks
|
||||
|
||||
**Problem:** ChatGPT uses curly quotes (“...”) instead of straight quotes ("...").
|
||||
|
||||
**Before:**
|
||||
> He said “the project is on track” but others disagreed.
|
||||
|
||||
**After:**
|
||||
> He said "the project is on track" but others disagreed.
|
||||
|
||||
---
|
||||
|
||||
## COMMUNICATION PATTERNS
|
||||
|
||||
### 19. Collaborative Communication Artifacts
|
||||
|
||||
**Words to watch:** I hope this helps, Of course!, Certainly!, You're absolutely right!, Would you like..., let me know, here is a...
|
||||
|
||||
**Problem:** Text meant as chatbot correspondence gets pasted as content.
|
||||
|
||||
**Before:**
|
||||
> Here is an overview of the French Revolution. I hope this helps! Let me know if you'd like me to expand on any section.
|
||||
|
||||
**After:**
|
||||
> The French Revolution began in 1789 when financial crisis and food shortages led to widespread unrest.
|
||||
|
||||
---
|
||||
|
||||
### 20. Knowledge-Cutoff Disclaimers
|
||||
|
||||
**Words to watch:** as of [date], Up to my last training update, While specific details are limited/scarce..., based on available information...
|
||||
|
||||
**Problem:** AI disclaimers about incomplete information get left in text.
|
||||
|
||||
**Before:**
|
||||
> While specific details about the company's founding are not extensively documented in readily available sources, it appears to have been established sometime in the 1990s.
|
||||
|
||||
**After:**
|
||||
> The company was founded in 1994, according to its registration documents.
|
||||
|
||||
---
|
||||
|
||||
### 21. Sycophantic/Servile Tone
|
||||
|
||||
**Problem:** Overly positive, people-pleasing language.
|
||||
|
||||
**Before:**
|
||||
> Great question! You're absolutely right that this is a complex topic. That's an excellent point about the economic factors.
|
||||
|
||||
**After:**
|
||||
> The economic factors you mentioned are relevant here.
|
||||
|
||||
---
|
||||
|
||||
## FILLER AND HEDGING
|
||||
|
||||
### 22. Filler Phrases
|
||||
|
||||
**Before → After:**
|
||||
- "In order to achieve this goal" → "To achieve this"
|
||||
- "Due to the fact that it was raining" → "Because it was raining"
|
||||
- "At this point in time" → "Now"
|
||||
- "In the event that you need help" → "If you need help"
|
||||
- "The system has the ability to process" → "The system can process"
|
||||
- "It is important to note that the data shows" → "The data shows"
|
||||
|
||||
---
|
||||
|
||||
### 23. Excessive Hedging
|
||||
|
||||
**Problem:** Over-qualifying statements.
|
||||
|
||||
**Before:**
|
||||
> It could potentially possibly be argued that the policy might have some effect on outcomes.
|
||||
|
||||
**After:**
|
||||
> The policy may affect outcomes.
|
||||
|
||||
---
|
||||
|
||||
### 24. Generic Positive Conclusions
|
||||
|
||||
**Problem:** Vague upbeat endings.
|
||||
|
||||
**Before:**
|
||||
> The future looks bright for the company. Exciting times lie ahead as they continue their journey toward excellence. This represents a major step in the right direction.
|
||||
|
||||
**After:**
|
||||
> The company plans to open two more locations next year.
|
||||
|
||||
---
|
||||
|
||||
## Process
|
||||
|
||||
1. Read the input text carefully
|
||||
2. Identify all instances of the patterns above
|
||||
3. Rewrite each problematic section
|
||||
4. Ensure the revised text:
|
||||
- Sounds natural when read aloud
|
||||
- Varies sentence structure naturally
|
||||
- Uses specific details over vague claims
|
||||
- Maintains appropriate tone for context
|
||||
- Uses simple constructions (is/are/has) where appropriate
|
||||
5. Present the humanized version
|
||||
|
||||
## Output Format
|
||||
|
||||
Provide:
|
||||
1. The rewritten text
|
||||
2. A brief summary of changes made (optional, if helpful)
|
||||
|
||||
---
|
||||
|
||||
## Full Example
|
||||
|
||||
**Before (AI-sounding):**
|
||||
> Great question! Here is an essay on this topic. I hope this helps!
|
||||
>
|
||||
> AI-assisted coding serves as an enduring testament to the transformative potential of large language models, marking a pivotal moment in the evolution of software development. In today's rapidly evolving technological landscape, these groundbreaking tools—nestled at the intersection of research and practice—are reshaping how engineers ideate, iterate, and deliver, underscoring their vital role in modern workflows.
|
||||
>
|
||||
> At its core, the value proposition is clear: streamlining processes, enhancing collaboration, and fostering alignment. It's not just about autocomplete; it's about unlocking creativity at scale, ensuring that organizations can remain agile while delivering seamless, intuitive, and powerful experiences to users. The tool serves as a catalyst. The assistant functions as a partner. The system stands as a foundation for innovation.
|
||||
>
|
||||
> Industry observers have noted that adoption has accelerated from hobbyist experiments to enterprise-wide rollouts, from solo developers to cross-functional teams. The technology has been featured in The New York Times, Wired, and The Verge. Additionally, the ability to generate documentation, tests, and refactors showcases how AI can contribute to better outcomes, highlighting the intricate interplay between automation and human judgment.
|
||||
>
|
||||
> - 💡 **Speed:** Code generation is significantly faster, reducing friction and empowering developers.
|
||||
> - 🚀 **Quality:** Output quality has been enhanced through improved training, contributing to higher standards.
|
||||
> - ✅ **Adoption:** Usage continues to grow, reflecting broader industry trends.
|
||||
>
|
||||
> While specific details are limited based on available information, it could potentially be argued that these tools might have some positive effect. Despite challenges typical of emerging technologies—including hallucinations, bias, and accountability—the ecosystem continues to thrive. In order to fully realize this potential, teams must align with best practices.
|
||||
>
|
||||
> In conclusion, the future looks bright. Exciting times lie ahead as we continue this journey toward excellence. Let me know if you’d like me to expand on any section!
|
||||
|
||||
**After (Humanized):**
|
||||
> AI coding assistants speed up some tasks. In a 2024 study by Google, developers using Codex completed simple functions 55% faster than a control group, but showed no improvement on debugging or architectural decisions.
|
||||
>
|
||||
> The tools are good at boilerplate: config files, test scaffolding, repetitive refactors. They are bad at knowing when they are wrong. I have mass-accepted suggestions that compiled, passed lint, and still did the wrong thing because I stopped paying attention.
|
||||
>
|
||||
> Mira, an engineer at a fintech startup I interviewed, said she uses Copilot "like autocomplete for boring code" but reviews every line before committing. Jake, a senior dev at a larger company, disabled it after it kept suggesting patterns from a deprecated internal library.
|
||||
>
|
||||
> The productivity claims are hard to verify. GitHub says Copilot users "accept 30% of suggestions," but acceptance is not correctness, and correctness is not value. The 2024 Uplevel study found no statistically significant difference in pull-request throughput between teams with and without AI assistants.
|
||||
>
|
||||
> None of this means the tools are useless. It means they are tools. They do not replace judgment, and they do not eliminate the need for tests. If you do not have tests, you cannot tell whether the suggestion is right.
|
||||
|
||||
**Changes made:**
|
||||
- Removed chatbot artifacts ("Great question!", "I hope this helps!", "Let me know if...")
|
||||
- Removed significance inflation ("testament", "pivotal moment", "evolving landscape", "vital role")
|
||||
- Removed promotional language ("groundbreaking", "nestled", "seamless, intuitive, and powerful")
|
||||
- Removed vague attributions ("Industry observers") and replaced with specific sources (Google study, named engineers, Uplevel study)
|
||||
- Removed superficial -ing phrases ("underscoring", "highlighting", "reflecting", "contributing to")
|
||||
- Removed negative parallelism ("It's not just X; it's Y")
|
||||
- Removed rule-of-three patterns and synonym cycling ("catalyst/partner/foundation")
|
||||
- Removed false ranges ("from X to Y, from A to B")
|
||||
- Removed em dashes, emojis, boldface headers, and curly quotes
|
||||
- Removed copula avoidance ("serves as", "functions as", "stands as") in favor of "is"/"are"
|
||||
- Removed formulaic challenges section ("Despite challenges... continues to thrive")
|
||||
- Removed knowledge-cutoff hedging ("While specific details are limited...")
|
||||
- Removed excessive hedging ("could potentially be argued that... might have some")
|
||||
- Removed filler phrases ("In order to", "At its core")
|
||||
- Removed generic positive conclusion ("the future looks bright", "exciting times lie ahead")
|
||||
- Replaced media name-dropping with specific claims from specific sources
|
||||
- Used simple sentence structures and concrete examples
|
||||
|
||||
---
|
||||
|
||||
## Reference
|
||||
|
||||
This skill is based on [Wikipedia:Signs of AI writing](https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing), maintained by WikiProject AI Cleanup. The patterns documented there come from observations of thousands of instances of AI-generated text on Wikipedia.
|
||||
|
||||
Key insight from Wikipedia: "LLMs use statistical algorithms to guess what should come next. The result tends toward the most statistically likely result that applies to the widest variety of cases."
|
||||
107
.skills/mcc-chatbot-authoring/SKILL.md
Normal file
107
.skills/mcc-chatbot-authoring/SKILL.md
Normal file
|
|
@ -0,0 +1,107 @@
|
|||
---
|
||||
name: mcc-chatbot-authoring
|
||||
description: Create, modify, repair, and wire Minecraft Console Client ChatBots and standalone `/script` bots. Use this whenever the user wants an MCC bot, C# script bot, chat or event handlers, periodic automation, movement logic, inventory logic, plugin-channel handling, or asks to fix or port an existing bot; default to standalone `//MCCScript` bots unless the user explicitly asks for a built-in MCC bot or repo wiring.
|
||||
---
|
||||
|
||||
# MCC ChatBot Authoring
|
||||
|
||||
Implement MCC chat bots against the bundled MCC authoring reference. Do not invent methods, lifecycle hooks, or registration steps.
|
||||
|
||||
Always read:
|
||||
- `references/authoring-reference.md`
|
||||
|
||||
Load only as needed:
|
||||
- `references/pattern-cookbook.md` for concrete standalone examples
|
||||
- `assets/script-chatbot-template.cs` for the default standalone `/script` path
|
||||
- `assets/builtin-chatbot-template.cs` only when the user explicitly requests a built-in bot
|
||||
|
||||
If the current workspace contains an MCC checkout, verify final names and signatures against local sources before editing. The skill should still work without those files.
|
||||
If there is no MCC checkout available, rely on the bundled reference and cookbook as the full source of truth for authoring patterns.
|
||||
|
||||
## Choose the bot type first
|
||||
|
||||
1. Default to a standalone script bot loaded with `/script`.
|
||||
2. Only choose a built-in bot when the user explicitly asks for a compiled MCC bot, repo wiring, automatic config loading, or changes under the built-in bot system.
|
||||
3. If the prompt is ambiguous, infer the likely target from commands, requested output files, or phrasing, state the assumption briefly, and proceed.
|
||||
4. If a user says only "make a bot", do not create a built-in bot.
|
||||
|
||||
## Source priority
|
||||
|
||||
When the local MCC checkout is available, prefer these sources in this order:
|
||||
1. `MinecraftClient/Scripting/ChatBot.cs` and current files under `MinecraftClient/ChatBots/`
|
||||
2. the bundled `references/authoring-reference.md`
|
||||
3. the bundled `references/pattern-cookbook.md`
|
||||
4. older `MinecraftClient/config/` sample bots only for ideas, not as the default scaffold
|
||||
|
||||
If an older sample conflicts with the current built-in bots, follow the current built-in bots.
|
||||
If the local checkout is not available, do not block on missing repo files. Use the bundled references directly.
|
||||
|
||||
## Hard rules
|
||||
|
||||
- Only use lifecycle hooks and helpers documented in the bundled reference or verified in the target codebase.
|
||||
- Do not send chat from `Initialize()`. Use `AfterGameJoined()` once the session can send messages.
|
||||
- Prefer the current Brigadier command-registration pattern for built-in bots. Do not introduce `ChatBotCommand` unless the surrounding code already uses it.
|
||||
- For message parsing, normalize with `GetVerbatim(text)` before `IsChatMessage(...)` or `IsPrivateMessage(...)`.
|
||||
- Clean up everything you register or start: commands, plugin channels, threads, timers, and movement locks.
|
||||
- If a built-in bot or long-running automation controls movement, follow a movement-lock pattern and release it on every stop path. Do not add `BotMovementLock` to a simple standalone `/script` bot unless the prompt or surrounding code explicitly needs shared movement coordination.
|
||||
- For built-in bots, follow the host codebase's localization and config-comment conventions instead of scattering hardcoded user-facing text.
|
||||
- For new code, prefer `Initialize()` over constructors for prerequisite checks and unload decisions.
|
||||
- In this repo, built-in bot wiring usually means edits in `MinecraftClient/Settings.cs` and `MinecraftClient/McClient.cs` in addition to the bot class.
|
||||
- For repair tasks, preserve the existing bot type and file layout unless the user explicitly asks for a conversion or restructure.
|
||||
|
||||
## Standalone script bots
|
||||
|
||||
Use the exact MCC metadata format from the bundled reference.
|
||||
This is the default path for new work.
|
||||
|
||||
The script should usually:
|
||||
- keep `Initialize()` for cheap setup only
|
||||
- use `GetText(...)`, `AfterGameJoined()`, and other event hooks for live behavior
|
||||
- log with `LogToConsole(...)`
|
||||
- send server chat or commands with `SendText(...)`
|
||||
- use `PerformInternalCommand(...)` only for MCC internal commands
|
||||
- add `//using MinecraftClient.Inventory` in metadata when the script uses inventory types explicitly
|
||||
- reuse the standalone snippets in `references/pattern-cookbook.md` before inventing new scaffolding
|
||||
- keep load instructions explicit, usually `/script FileName.cs`
|
||||
|
||||
## Built-in bots
|
||||
|
||||
Built-in bots usually need three pieces:
|
||||
- the bot class itself
|
||||
- config wiring in the chat-bot config model
|
||||
- bot registration in the load flow
|
||||
|
||||
If the codebase exposes commands, follow the built-in command and unload pattern from the bundled reference. If it exposes new settings or status text, follow the codebase's localization and config-comment patterns.
|
||||
|
||||
When working in this checkout, built-in bot delivery usually needs:
|
||||
- a new file under `MinecraftClient/ChatBots/`
|
||||
- a config property inside `Settings.ChatBotConfigHealper.ChatBotConfig`
|
||||
- a `BotLoad(new YourBot())` line inside `McClient.RegisterBots(...)`
|
||||
- literal code snippets or patch hunks for the `Settings.cs` property and the `McClient.cs` registration line, not only prose notes
|
||||
|
||||
## Repair flow
|
||||
|
||||
When the user asks to fix or debug a bot:
|
||||
- identify whether it is standalone or built-in and keep that shape unless told otherwise
|
||||
- remove the broken pattern first, then preserve the intended behavior
|
||||
- check especially for these regressions: `SendText(...)` in `Initialize()`, raw formatted chat parsing, inventory snapshot mutation, missing command unregister, missing plugin-channel unregister, and unreleased movement locks
|
||||
- reuse the local repo's modern pattern instead of patching around a legacy helper when the helper is no longer current
|
||||
|
||||
## Delivery checklist
|
||||
|
||||
Before finishing, verify:
|
||||
- the class inherits `ChatBot`
|
||||
- the chosen overrides exist in the MCC ChatBot API
|
||||
- standalone script metadata is exact if this is a `/script` bot
|
||||
- built-in bots are fully wired into config and registration if needed
|
||||
- all command registrations, background work, and movement locks are released
|
||||
- files and namespaces match the surrounding codebase
|
||||
|
||||
## Output
|
||||
|
||||
When you implement or modify a bot:
|
||||
- state whether it is a standalone script bot or built-in bot
|
||||
- list the files you changed
|
||||
- mention any required config keys or the MCC command used to load it
|
||||
- when built-in wiring is involved, show the exact inserted code lines or patch hunks for `Settings.cs` and `McClient.cs`
|
||||
- call out assumptions briefly if the user did not specify bot type or trigger behavior
|
||||
|
|
@ -0,0 +1,57 @@
|
|||
// Use this template only when the user explicitly requests a built-in MCC bot.
|
||||
|
||||
using MinecraftClient.Scripting;
|
||||
using Tomlet.Attributes;
|
||||
|
||||
namespace MinecraftClient.ChatBots
|
||||
{
|
||||
public class ExampleBot : ChatBot
|
||||
{
|
||||
private const string BotName = "ExampleBot";
|
||||
|
||||
public static Configs Config = new();
|
||||
|
||||
[TomlDoNotInlineObject]
|
||||
public class Configs
|
||||
{
|
||||
public bool Enabled = false;
|
||||
|
||||
public void OnSettingUpdate()
|
||||
{
|
||||
}
|
||||
}
|
||||
|
||||
public override void Initialize()
|
||||
{
|
||||
LogToConsole(BotName, "Initialized.");
|
||||
}
|
||||
|
||||
public override void AfterGameJoined()
|
||||
{
|
||||
}
|
||||
|
||||
public override void GetText(string text)
|
||||
{
|
||||
text = GetVerbatim(text);
|
||||
|
||||
string message = "";
|
||||
string username = "";
|
||||
|
||||
if (IsPrivateMessage(text, ref message, ref username))
|
||||
{
|
||||
}
|
||||
else if (IsChatMessage(text, ref message, ref username))
|
||||
{
|
||||
}
|
||||
}
|
||||
|
||||
public override void OnUnload()
|
||||
{
|
||||
}
|
||||
|
||||
public override bool OnDisconnect(DisconnectReason reason, string message)
|
||||
{
|
||||
return false;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
@ -0,0 +1,37 @@
|
|||
//MCCScript 1.0
|
||||
|
||||
MCC.LoadBot(new ExampleScriptBot());
|
||||
|
||||
//MCCScript Extensions
|
||||
|
||||
public class ExampleScriptBot : ChatBot
|
||||
{
|
||||
public override void Initialize()
|
||||
{
|
||||
LogToConsole("ExampleScriptBot initialized.");
|
||||
}
|
||||
|
||||
public override void AfterGameJoined()
|
||||
{
|
||||
// Safe place for startup chat or commands.
|
||||
}
|
||||
|
||||
public override void GetText(string text)
|
||||
{
|
||||
text = GetVerbatim(text);
|
||||
|
||||
string message = "";
|
||||
string username = "";
|
||||
|
||||
if (IsPrivateMessage(text, ref message, ref username))
|
||||
{
|
||||
LogToConsole("PM from " + username + ": " + message);
|
||||
return;
|
||||
}
|
||||
|
||||
if (IsChatMessage(text, ref message, ref username))
|
||||
{
|
||||
LogToConsole("Chat from " + username + ": " + message);
|
||||
}
|
||||
}
|
||||
}
|
||||
492
.skills/mcc-chatbot-authoring/references/authoring-reference.md
Normal file
492
.skills/mcc-chatbot-authoring/references/authoring-reference.md
Normal file
|
|
@ -0,0 +1,492 @@
|
|||
# MCC ChatBot Reference
|
||||
|
||||
Self-contained authoring notes for Minecraft Console Client chat bots.
|
||||
|
||||
## Bot types
|
||||
|
||||
MCC supports two common authoring paths:
|
||||
- standalone script bots loaded at runtime with `/script`
|
||||
- built-in bots compiled into the MCC codebase
|
||||
|
||||
Default to a standalone `/script` bot unless the user explicitly asks for a built-in bot or repo wiring.
|
||||
|
||||
## Embedded current patterns
|
||||
|
||||
This skill is intended to work even without an MCC checkout. The patterns below capture the important behavior that would otherwise be borrowed from current repo examples.
|
||||
|
||||
If the local repo is available, you can verify against files such as `TestBot.cs`, `RemoteControl.cs`, `FollowPlayer.cs`, `ItemsCollector.cs`, and `Farmer.cs`. If it is not available, use the embedded patterns here directly.
|
||||
|
||||
### Minimal chat parsing pattern
|
||||
|
||||
Use this as the baseline for public/private chat handling:
|
||||
|
||||
```csharp
|
||||
public override void GetText(string text)
|
||||
{
|
||||
string message = "";
|
||||
string sender = "";
|
||||
text = GetVerbatim(text);
|
||||
|
||||
if (IsPrivateMessage(text, ref message, ref sender))
|
||||
{
|
||||
LogToConsole("PM from " + sender + ": " + message);
|
||||
}
|
||||
else if (IsChatMessage(text, ref message, ref sender))
|
||||
{
|
||||
LogToConsole("Chat from " + sender + ": " + message);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
What matters:
|
||||
- normalize first with `GetVerbatim(text)`
|
||||
- handle PMs before public chat if both matter
|
||||
- keep simple chat bots deterministic and small
|
||||
|
||||
### Owner-gated PM control pattern
|
||||
|
||||
Use this when a bot owner should be able to whisper MCC internal commands:
|
||||
|
||||
```csharp
|
||||
public override void GetText(string text)
|
||||
{
|
||||
text = GetVerbatim(text).Trim();
|
||||
string command = "";
|
||||
string sender = "";
|
||||
|
||||
if (IsPrivateMessage(text, ref command, ref sender)
|
||||
&& Settings.Config.Main.Advanced.BotOwners.Contains(sender.ToLowerInvariant()))
|
||||
{
|
||||
CmdResult result = new();
|
||||
PerformInternalCommand(command, ref result);
|
||||
SendPrivateMessage(sender, result.ToString());
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
What matters:
|
||||
- `PerformInternalCommand(...)` is for MCC commands, not server chat commands
|
||||
- owner gating should use `Settings.Config.Main.Advanced.BotOwners`
|
||||
- if `CmdResult` is used in a standalone script, add `//using MinecraftClient.CommandHandler`
|
||||
|
||||
### Periodic work pattern
|
||||
|
||||
Use `Update()` plus a counter or timestamp for simple repeated work:
|
||||
|
||||
```csharp
|
||||
private int count = 0;
|
||||
|
||||
public override void Update()
|
||||
{
|
||||
count++;
|
||||
if (count < Settings.DoubleToTick(60))
|
||||
return;
|
||||
|
||||
count = 0;
|
||||
SendText("/list");
|
||||
}
|
||||
```
|
||||
|
||||
What matters:
|
||||
- avoid a worker thread for simple periodic loops
|
||||
- avoid `Thread.Sleep(...)` inside `Update()`
|
||||
- if sending chat, do it from a join-safe path like `Update()` or `AfterGameJoined()`, not `Initialize()`
|
||||
|
||||
### Built-in Brigadier command pattern
|
||||
|
||||
Use this for built-in command bots:
|
||||
|
||||
```csharp
|
||||
public override void Initialize()
|
||||
{
|
||||
McClient.dispatcher.Register(l => l.Literal("help")
|
||||
.Then(l => l.Literal(CommandName)
|
||||
.Executes(r => OnCommandHelp(r.Source, string.Empty))
|
||||
)
|
||||
);
|
||||
|
||||
McClient.dispatcher.Register(l => l.Literal(CommandName)
|
||||
.Then(l => l.Literal("stop")
|
||||
.Executes(r => OnCommandStop(r.Source)))
|
||||
.Then(l => l.Literal("_help")
|
||||
.Executes(r => OnCommandHelp(r.Source, string.Empty))
|
||||
.Redirect(McClient.dispatcher.GetRoot().GetChild("help").GetChild(CommandName)))
|
||||
);
|
||||
}
|
||||
|
||||
public override void OnUnload()
|
||||
{
|
||||
McClient.dispatcher.Unregister(CommandName);
|
||||
McClient.dispatcher.GetRoot().GetChild("help").RemoveChild(CommandName);
|
||||
}
|
||||
```
|
||||
|
||||
What matters:
|
||||
- register commands in `Initialize()`
|
||||
- unregister the command tree in `OnUnload()`
|
||||
- remove the help child you added in `OnUnload()`
|
||||
- prefer this over legacy command wrappers for new built-in work
|
||||
|
||||
### Built-in config and wiring pattern
|
||||
|
||||
Use this as the default built-in shape:
|
||||
|
||||
```csharp
|
||||
public class ExampleBot : ChatBot
|
||||
{
|
||||
public static Configs Config = new();
|
||||
|
||||
[TomlDoNotInlineObject]
|
||||
public class Configs
|
||||
{
|
||||
public bool Enabled = false;
|
||||
|
||||
public void OnSettingUpdate()
|
||||
{
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Typical host wiring shape:
|
||||
|
||||
```csharp
|
||||
[TomlPrecedingComment("$ChatBot.ExampleBot$")]
|
||||
public ChatBots.ExampleBot.Configs ExampleBot
|
||||
{
|
||||
get { return ChatBots.ExampleBot.Config; }
|
||||
set { ChatBots.ExampleBot.Config = value; ChatBots.ExampleBot.Config.OnSettingUpdate(); }
|
||||
}
|
||||
```
|
||||
|
||||
```csharp
|
||||
if (Config.ChatBot.ExampleBot.Enabled) { BotLoad(new ExampleBot()); }
|
||||
```
|
||||
|
||||
What matters:
|
||||
- built-in configurable bots default to `Enabled = false`
|
||||
- `OnSettingUpdate()` is the place to normalize config values
|
||||
- built-in delivery is incomplete without both config wiring and load registration
|
||||
|
||||
### Movement gating pattern
|
||||
|
||||
Use this shape when a built-in bot owns movement:
|
||||
|
||||
```csharp
|
||||
public override void Initialize()
|
||||
{
|
||||
if (!GetEntityHandlingEnabled())
|
||||
{
|
||||
LogToConsole("Entity handling is required.");
|
||||
UnloadBot();
|
||||
return;
|
||||
}
|
||||
|
||||
if (!GetTerrainEnabled())
|
||||
{
|
||||
LogToConsole("Terrain handling is required.");
|
||||
UnloadBot();
|
||||
return;
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```csharp
|
||||
var movementLock = BotMovementLock.Instance;
|
||||
if (movementLock is { IsLocked: true })
|
||||
return;
|
||||
|
||||
movementLock?.Lock("Example Bot");
|
||||
```
|
||||
|
||||
```csharp
|
||||
public override void OnUnload()
|
||||
{
|
||||
BotMovementLock.Instance?.UnLock("Example Bot");
|
||||
}
|
||||
```
|
||||
|
||||
What matters:
|
||||
- guard terrain and entity handling before movement logic
|
||||
- built-in movement bots should use `BotMovementLock`
|
||||
- release the lock on every stop path, including unload and disconnect-sensitive flows
|
||||
|
||||
### Dropped-item collector pattern
|
||||
|
||||
Use this as the standalone item-search baseline:
|
||||
|
||||
```csharp
|
||||
private DateTime nextScan = DateTime.MinValue;
|
||||
|
||||
public override void Update()
|
||||
{
|
||||
var now = DateTime.UtcNow;
|
||||
if (now < nextScan || ClientIsMoving())
|
||||
return;
|
||||
|
||||
nextScan = now.AddSeconds(1);
|
||||
|
||||
var here = GetCurrentLocation();
|
||||
var target = GetEntities().Values
|
||||
.Where(entity => entity.Type == EntityType.Item && entity.Location.Distance(here) <= 15)
|
||||
.OrderBy(entity => entity.Location.Distance(here))
|
||||
.FirstOrDefault();
|
||||
|
||||
if (target != null)
|
||||
MoveToLocation(target.Location);
|
||||
}
|
||||
```
|
||||
|
||||
What matters:
|
||||
- simple standalone collectors do not need a worker thread
|
||||
- simple standalone collectors also do not need `BotMovementLock` by default
|
||||
- `GetEntities()` plus distance ordering is the core search pattern
|
||||
|
||||
### Inventory selection pattern
|
||||
|
||||
Use this as the default hotbar-switch pattern:
|
||||
|
||||
```csharp
|
||||
private bool TrySwitchToItem(ItemType itemType)
|
||||
{
|
||||
var inventory = GetPlayerInventory();
|
||||
|
||||
var hotbarSlots = inventory.SearchItem(itemType)
|
||||
.Where(slot => slot >= 36 && slot <= 44)
|
||||
.ToArray();
|
||||
|
||||
if (hotbarSlots.Length == 0)
|
||||
return false;
|
||||
|
||||
ChangeSlot((short)(hotbarSlots[0] - 36));
|
||||
return true;
|
||||
}
|
||||
```
|
||||
|
||||
What matters:
|
||||
- guard with `GetInventoryEnabled()`
|
||||
- search inventory snapshots, but mutate real server state with helpers like `ChangeSlot(...)`
|
||||
- do not treat local `Container.Items` mutation as real inventory manipulation
|
||||
|
||||
Use the older config examples only for ideas, not as primary scaffolding.
|
||||
|
||||
## Standalone script format
|
||||
|
||||
A standalone script bot has two parts in this order:
|
||||
1. metadata block
|
||||
2. one or more C# classes, with the main bot class inheriting `ChatBot`
|
||||
|
||||
Required metadata rules:
|
||||
- line 1 must be exactly `//MCCScript 1.0`
|
||||
- metadata must include `MCC.LoadBot(new BotClassName());`
|
||||
- metadata ends with `//MCCScript Extensions`
|
||||
- optional metadata directives use `//using Namespace` and `//dll SomeLibrary.dll`
|
||||
- do not insert a space after `//` in metadata directives
|
||||
|
||||
Typical runtime flow:
|
||||
- place the script file beside MCC
|
||||
- connect to a server
|
||||
- load it with `/script YourBotFile.cs`
|
||||
|
||||
### Namespace linking for inventory code
|
||||
|
||||
If a standalone script uses inventory-specific types such as `Container`, `ItemType`, `WindowActionType`, or `ItemMovingHelper`, add this metadata import:
|
||||
|
||||
```csharp
|
||||
//using MinecraftClient.Inventory
|
||||
```
|
||||
|
||||
For built-in bots, use a normal C# import:
|
||||
|
||||
```csharp
|
||||
using MinecraftClient.Inventory;
|
||||
```
|
||||
|
||||
## Lifecycle summary
|
||||
|
||||
Common lifecycle hooks:
|
||||
- `Initialize()`
|
||||
called once when the bot loads; use it for cheap setup only
|
||||
- `AfterGameJoined()`
|
||||
called after the server has been joined successfully, and again after reconnecting; use it when chat can be sent
|
||||
- `Update()`
|
||||
called roughly every 100 ms
|
||||
- `OnUnload()`
|
||||
called when the bot unloads; release resources here
|
||||
- `OnDisconnect(DisconnectReason reason, string message)`
|
||||
called on disconnect; stop background work and clean up reconnect-sensitive state here
|
||||
|
||||
Important rule:
|
||||
- do not send chat from `Initialize()`; use `AfterGameJoined()` instead
|
||||
- prefer `Initialize()` over constructors for environment checks and resource setup
|
||||
|
||||
## Common event hooks
|
||||
|
||||
Useful event hooks include:
|
||||
- `GetText(string text)`
|
||||
- `GetText(string text, string? json)`
|
||||
- `OnPlayerJoin(Guid uuid, string name)`
|
||||
- `OnPlayerLeave(Guid uuid, string? name)`
|
||||
- `OnEntitySpawn(Entity entity)`
|
||||
- `OnEntityDespawn(Entity entity)`
|
||||
- `OnEntityMove(Entity entity)`
|
||||
- `OnHealthUpdate(float health, int food)`
|
||||
- `OnMapData(...)`
|
||||
- `OnInventoryUpdate(int inventoryId)`
|
||||
- `OnPluginMessage(string channel, byte[] data)`
|
||||
- `OnNetworkPacket(int packetID, List<byte> packetData, bool isLogin, bool isInbound)`
|
||||
|
||||
Only override hooks that actually exist in the target MCC ChatBot API.
|
||||
|
||||
## Common helpers
|
||||
|
||||
Text and messaging helpers:
|
||||
- `GetVerbatim(text)` strips Minecraft formatting codes
|
||||
- `IsChatMessage(text, ref message, ref sender)` parses public chat
|
||||
- `IsPrivateMessage(text, ref message, ref sender)` parses private chat
|
||||
- `IsValidName(username)` validates a Minecraft username
|
||||
- `SendText(text)` sends chat or server commands
|
||||
- `SendPrivateMessage(player, message)` sends a private message
|
||||
- `PerformInternalCommand(command, ...)` runs an internal MCC command, not a server command
|
||||
- `LogToConsole(text)` writes a bot-prefixed console message
|
||||
|
||||
Lifecycle and threading helpers:
|
||||
- `InvokeOnMainThread(...)`
|
||||
- `ScheduleOnMainThread(...)`
|
||||
- `ReconnectToTheServer(...)`
|
||||
- `UnloadBot()`
|
||||
- `BotLoad(chatBot)`
|
||||
- `RunScript(filename, ...)`
|
||||
|
||||
World and player-state helpers:
|
||||
- `GetWorld()`
|
||||
- `GetEntities()`
|
||||
- `GetCurrentLocation()`
|
||||
- `ClientIsMoving()`
|
||||
- `GetOnlinePlayers()`
|
||||
- `GetOnlinePlayersWithUUID()`
|
||||
- `GetServerTPS()`
|
||||
- `GetProtocolVersion()`
|
||||
|
||||
Movement and inventory helpers:
|
||||
- `MoveToLocation(...)`
|
||||
- `LookAtLocation(...)`
|
||||
- `GetInventoryEnabled()`
|
||||
- `GetPlayerInventory()`
|
||||
- `GetInventories()`
|
||||
- `GetItemMovingHelper(...)`
|
||||
- `WindowAction(...)`
|
||||
- `ChangeSlot(...)`
|
||||
- `GetCurrentSlot()`
|
||||
- `UseItemInHand()`
|
||||
- `UseItemInLeftHand()`
|
||||
- `CloseInventory(...)`
|
||||
- `DigBlock(...)`
|
||||
- `InteractEntity(...)`
|
||||
|
||||
## Inventory notes
|
||||
|
||||
Inventory handling is optional in MCC. Check `GetInventoryEnabled()` before relying on inventory state or mutation.
|
||||
|
||||
Important behavior:
|
||||
- `GetPlayerInventory()` returns a snapshot copy of the player's inventory
|
||||
- `GetInventories()` returns current container snapshots
|
||||
- writing to those `Container` objects locally does not update the server
|
||||
- to actually change inventory state, use `ChangeSlot(...)`, `WindowAction(...)`, `GetItemMovingHelper(...)`, `UseItemInHand()`, or related helpers
|
||||
|
||||
Useful practical facts:
|
||||
- hotbar selection uses `ChangeSlot(0..8)`
|
||||
- hotbar slots are commonly `36..44` in inventory slot numbering
|
||||
- the offhand slot is commonly `45`
|
||||
- `Container.SearchItem(...)` is the normal way to locate items by type
|
||||
|
||||
Good inventory workflow:
|
||||
1. guard with `GetInventoryEnabled()`
|
||||
2. read the current container using `GetPlayerInventory()`
|
||||
3. locate slots with `SearchItem(...)` or `Items`
|
||||
4. mutate server state using `ChangeSlot(...)`, `WindowAction(...)`, or `ItemMovingHelper`
|
||||
5. if needed, react to `OnInventoryUpdate(...)`, `OnInventoryOpen(...)`, or `OnInventoryClose(...)`
|
||||
|
||||
Plugins and channels:
|
||||
- `RegisterPluginChannel(channel)`
|
||||
- `UnregisterPluginChannel(channel)`
|
||||
- `SendPluginChannelMessage(channel, data, ...)`
|
||||
|
||||
## Built-in bot pattern
|
||||
|
||||
A built-in bot usually follows this shape:
|
||||
- a class that inherits `ChatBot`
|
||||
- an optional static `Config` field
|
||||
- a nested `[TomlDoNotInlineObject]` `Configs` class for settings
|
||||
- an `Enabled = false` setting by default
|
||||
- `OnSettingUpdate()` to normalize or validate config values
|
||||
|
||||
If the bot is configurable, the host codebase usually also needs:
|
||||
- config wiring in the chat-bot config model
|
||||
- load registration so enabled bots are instantiated automatically
|
||||
|
||||
In this MCC checkout, the usual built-in wiring points are:
|
||||
- `MinecraftClient/Settings.cs` inside `Settings.ChatBotConfigHealper.ChatBotConfig`
|
||||
- `MinecraftClient/McClient.cs` inside `RegisterBots(...)`
|
||||
|
||||
Match the surrounding `[TomlPrecedingComment(...)]`, property-forwarding, and `BotLoad(new YourBot())` style instead of inventing a different config path.
|
||||
When presenting built-in wiring, prefer literal code snippets or patch hunks for those two edits so the wiring can be checked directly.
|
||||
|
||||
If the bot adds user-facing settings or messages, follow the host codebase's localization and config-comment conventions instead of scattering hardcoded strings.
|
||||
|
||||
## Command pattern
|
||||
|
||||
For standalone script bots, prefer chat or PM handling in `GetText(...)` unless the user explicitly asks for built-in command registration.
|
||||
|
||||
For built-in commands, prefer the current Brigadier dispatcher pattern:
|
||||
- register commands in `Initialize()`
|
||||
- add a help entry if the bot exposes commands
|
||||
- unregister the command tree in `OnUnload()`
|
||||
- remove any help child added during registration in `OnUnload()`
|
||||
|
||||
Avoid using legacy command wrappers if the current codebase uses direct dispatcher registration.
|
||||
In this checkout, treat direct `McClient.dispatcher.Register(...)` usage in current built-in bots as the source of truth.
|
||||
|
||||
## Concurrency and cleanup
|
||||
|
||||
If the bot starts background work:
|
||||
- stop it in `OnUnload()`
|
||||
- stop it in `OnDisconnect(...)`
|
||||
- consider resetting state in `AfterGameJoined()` after relog
|
||||
- prefer `Update()` plus counters or timestamps over unmanaged threads when the task is simple periodic work
|
||||
|
||||
If the bot controls movement:
|
||||
- use a movement-lock discipline
|
||||
- release the lock on every stop path
|
||||
- avoid fighting other movement bots
|
||||
- `BotMovementLock` is mainly for built-in bots or shared long-running automation; a simple standalone script that just calls `MoveToLocation(...)` does not need it by default
|
||||
|
||||
When interacting with client state from background logic, use the main-thread helpers when required by the codebase.
|
||||
|
||||
## Practical defaults
|
||||
|
||||
For simple chat bots:
|
||||
- normalize text with `GetVerbatim(text)`
|
||||
- inspect private chat first if the bot listens for whispers
|
||||
- then inspect public chat
|
||||
- keep response logic small and deterministic
|
||||
|
||||
For long-running automation bots:
|
||||
- guard prerequisites early, such as entity handling or terrain support
|
||||
- fail fast with a clear log message if prerequisites are missing
|
||||
- release all ongoing work cleanly on unload and disconnect
|
||||
|
||||
## Common pitfalls
|
||||
|
||||
- Incorrect metadata line 1 will break standalone script loading.
|
||||
- Missing `MCC.LoadBot(new BotClassName())` will prevent standalone script registration.
|
||||
- Sending chat in `Initialize()` is too early.
|
||||
- Doing prerequisite checks or unloading from the constructor is harder to reason about than using `Initialize()`.
|
||||
- Parsing raw formatted text without `GetVerbatim()` causes brittle chat matching.
|
||||
- Inventing methods not present in the MCC ChatBot API leads to dead code.
|
||||
- Built-in bot work is incomplete if config or registration wiring is missing.
|
||||
- Command bots are incomplete if they register commands but do not unregister them.
|
||||
- `RegisterChatBotCommand(...)` comes from older samples and is not a reliable current pattern for this checkout.
|
||||
- `ChatBotCommand` exists, but the current built-in bots use Brigadier directly; do not prefer `ChatBotCommand` for new work.
|
||||
- Blocking `Thread.Sleep(...)` inside `Update()` is a bad default. Prefer timers, counters, or timestamp-based scheduling.
|
||||
- Mutating the `Container` returned by `GetPlayerInventory()` does not change the server. Use inventory actions instead.
|
||||
330
.skills/mcc-chatbot-authoring/references/pattern-cookbook.md
Normal file
330
.skills/mcc-chatbot-authoring/references/pattern-cookbook.md
Normal file
|
|
@ -0,0 +1,330 @@
|
|||
# MCC Pattern Cookbook
|
||||
|
||||
Concrete patterns for standalone MCC `/script` bots. Use these before inventing new scaffolding.
|
||||
|
||||
## Periodic task without threads
|
||||
|
||||
Use `Update()` plus a timestamp or counter. This comes from the old `sample-script-with-task.cs` example and still holds up well.
|
||||
|
||||
```csharp
|
||||
public class PeriodicTaskBot : ChatBot
|
||||
{
|
||||
private DateTime nextRun = DateTime.MinValue;
|
||||
|
||||
public override void Update()
|
||||
{
|
||||
var now = DateTime.UtcNow;
|
||||
if (now < nextRun)
|
||||
return;
|
||||
|
||||
nextRun = now.AddSeconds(30);
|
||||
LogDebugToConsole("Running periodic task");
|
||||
SendText("/ping");
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Why this pattern is good:
|
||||
- stays on MCC's normal tick flow
|
||||
- avoids background threads for simple periodic work
|
||||
- keeps the bot responsive to unload and disconnect
|
||||
|
||||
## Chat and PM handling
|
||||
|
||||
This combines the useful parts of `TestBot`, `sample-script-pm-forwarder.cs`, and `RemoteControl.cs`.
|
||||
|
||||
```csharp
|
||||
public override void GetText(string text)
|
||||
{
|
||||
text = GetVerbatim(text);
|
||||
|
||||
string message = "";
|
||||
string sender = "";
|
||||
|
||||
if (IsPrivateMessage(text, ref message, ref sender))
|
||||
{
|
||||
LogToConsole("PM from " + sender + ": " + message);
|
||||
return;
|
||||
}
|
||||
|
||||
if (IsChatMessage(text, ref message, ref sender))
|
||||
{
|
||||
LogToConsole("Chat from " + sender + ": " + message);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Owner-gated internal command handling:
|
||||
|
||||
Add `//using MinecraftClient.CommandHandler` in the script metadata if you use `CmdResult`.
|
||||
|
||||
```csharp
|
||||
public override void GetText(string text)
|
||||
{
|
||||
text = GetVerbatim(text).Trim();
|
||||
|
||||
string command = "";
|
||||
string sender = "";
|
||||
|
||||
if (IsPrivateMessage(text, ref command, ref sender)
|
||||
&& Settings.Config.Main.Advanced.BotOwners.Contains(sender.ToLowerInvariant()))
|
||||
{
|
||||
CmdResult result = new();
|
||||
PerformInternalCommand(command, ref result);
|
||||
SendPrivateMessage(sender, result.ToString());
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Movement with prerequisite checks
|
||||
|
||||
Modern movement code should copy the guard style from current built-in bots, not the older constructor-heavy scripts.
|
||||
|
||||
```csharp
|
||||
public override void Initialize()
|
||||
{
|
||||
if (!GetEntityHandlingEnabled() || !GetTerrainEnabled())
|
||||
{
|
||||
LogToConsole("Entity handling and terrain handling are required.");
|
||||
UnloadBot();
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Simple "look at nearest player" logic adapted from `AutoLook.cs`:
|
||||
|
||||
```csharp
|
||||
private Entity? trackedPlayer = null;
|
||||
|
||||
public override void OnEntitySpawn(Entity entity)
|
||||
{
|
||||
TryTrack(entity);
|
||||
}
|
||||
|
||||
public override void OnEntityDespawn(Entity entity)
|
||||
{
|
||||
if (trackedPlayer != null && entity.ID == trackedPlayer.ID)
|
||||
trackedPlayer = null;
|
||||
}
|
||||
|
||||
public override void OnEntityMove(Entity entity)
|
||||
{
|
||||
if (!TryTrack(entity))
|
||||
return;
|
||||
|
||||
LookAtLocation(entity.Location);
|
||||
}
|
||||
|
||||
private bool TryTrack(Entity entity)
|
||||
{
|
||||
if (entity.Type != EntityType.Player)
|
||||
return false;
|
||||
|
||||
if (trackedPlayer == null)
|
||||
{
|
||||
trackedPlayer = entity;
|
||||
return true;
|
||||
}
|
||||
|
||||
if (GetCurrentLocation().Distance(entity.Location) < GetCurrentLocation().Distance(trackedPlayer.Location))
|
||||
trackedPlayer = entity;
|
||||
|
||||
return trackedPlayer.ID == entity.ID;
|
||||
}
|
||||
```
|
||||
|
||||
## Search for dropped items and move to them
|
||||
|
||||
This is the safest pattern to preserve from `ItemsCollector.cs` for standalone scripts.
|
||||
|
||||
```csharp
|
||||
public class NearbyItemsBot : ChatBot
|
||||
{
|
||||
private DateTime nextScan = DateTime.MinValue;
|
||||
|
||||
public override void Initialize()
|
||||
{
|
||||
if (!GetEntityHandlingEnabled() || !GetTerrainEnabled())
|
||||
{
|
||||
LogToConsole("Entity handling and terrain handling are required.");
|
||||
UnloadBot();
|
||||
}
|
||||
}
|
||||
|
||||
public override void Update()
|
||||
{
|
||||
var now = DateTime.UtcNow;
|
||||
if (now < nextScan || ClientIsMoving())
|
||||
return;
|
||||
|
||||
nextScan = now.AddSeconds(1);
|
||||
|
||||
var here = GetCurrentLocation();
|
||||
var target = GetEntities().Values
|
||||
.Where(entity => entity.Type == EntityType.Item && entity.Location.Distance(here) <= 15)
|
||||
.OrderBy(entity => entity.Location.Distance(here))
|
||||
.FirstOrDefault();
|
||||
|
||||
if (target != null)
|
||||
MoveToLocation(target.Location);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Why this version is better than older farming scripts:
|
||||
- no unmanaged worker thread
|
||||
- no busy wait loop around movement
|
||||
- uses the current `GetEntities()` pattern
|
||||
|
||||
## Search for blocks or crops in the world
|
||||
|
||||
The old sugar cane and mining scripts still contain a useful search idea: use `GetWorld().FindBlock(...)`, then filter and sort.
|
||||
|
||||
```csharp
|
||||
var targets = GetWorld()
|
||||
.FindBlock(GetCurrentLocation(), Material.SugarCane, 16)
|
||||
.Where(block =>
|
||||
GetWorld().GetBlock(new Location(block.X, block.Y - 1, block.Z)).Type == Material.SugarCane)
|
||||
.OrderBy(block => block.Distance(GetCurrentLocation()))
|
||||
.ToList();
|
||||
```
|
||||
|
||||
Use this as a search primitive. Then decide separately how to move, dig, or harvest.
|
||||
|
||||
## Inventory access and manipulation
|
||||
|
||||
If a standalone script uses inventory types directly, add this import in the metadata block:
|
||||
|
||||
```csharp
|
||||
//using MinecraftClient.Inventory
|
||||
```
|
||||
|
||||
For built-in bots, add:
|
||||
|
||||
```csharp
|
||||
using MinecraftClient.Inventory;
|
||||
```
|
||||
|
||||
Always guard inventory logic first:
|
||||
|
||||
```csharp
|
||||
public override void Initialize()
|
||||
{
|
||||
if (!GetInventoryEnabled())
|
||||
{
|
||||
LogToConsole("Inventory handling is required.");
|
||||
UnloadBot();
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Important rule:
|
||||
- `GetPlayerInventory()` returns a snapshot copy, so editing its `Items` dictionary does not change the server
|
||||
- actual changes must go through `ChangeSlot(...)`, `WindowAction(...)`, `GetItemMovingHelper(...)`, `UseItemInHand()`, and related helpers
|
||||
|
||||
### Search inventory for an item
|
||||
|
||||
This combines the useful current logic from `Farmer.cs` and `AutoEat.cs`.
|
||||
|
||||
```csharp
|
||||
private bool TrySwitchToItem(ItemType itemType)
|
||||
{
|
||||
var inventory = GetPlayerInventory();
|
||||
|
||||
if (inventory.Items.TryGetValue(GetCurrentSlot() - 36, out var held) && held.Type == itemType)
|
||||
return true;
|
||||
|
||||
var hotbarSlots = inventory.SearchItem(itemType)
|
||||
.Where(slot => slot >= 36 && slot <= 44)
|
||||
.ToArray();
|
||||
|
||||
if (hotbarSlots.Length == 0)
|
||||
return false;
|
||||
|
||||
ChangeSlot((short)(hotbarSlots[0] - 36));
|
||||
return true;
|
||||
}
|
||||
```
|
||||
|
||||
Use this for simple hotbar selection. For deeper inventory reshuffling, built-in bots usually need more helper logic.
|
||||
|
||||
### Move an item into the hotbar
|
||||
|
||||
Use this when the item exists in inventory but is not already on the hotbar.
|
||||
|
||||
```csharp
|
||||
private bool TryMoveItemToHotbar(ItemType itemType, short targetHotbarSlot = 0)
|
||||
{
|
||||
var inventory = GetPlayerInventory();
|
||||
var matches = inventory.SearchItem(itemType);
|
||||
|
||||
if (matches.Length == 0)
|
||||
return false;
|
||||
|
||||
var targetInventorySlot = 36 + targetHotbarSlot;
|
||||
|
||||
if (matches[0] >= 36 && matches[0] <= 44)
|
||||
{
|
||||
ChangeSlot((short)(matches[0] - 36));
|
||||
return true;
|
||||
}
|
||||
|
||||
var movingHelper = GetItemMovingHelper(inventory);
|
||||
movingHelper.Swap(matches[0], targetInventorySlot);
|
||||
ChangeSlot(targetHotbarSlot);
|
||||
return true;
|
||||
}
|
||||
```
|
||||
|
||||
Why this pattern is good:
|
||||
- it reads the current snapshot first
|
||||
- it does not pretend local `Container` edits affect the server
|
||||
- it uses the item-moving helper for real inventory manipulation
|
||||
|
||||
### Drop or click items with window actions
|
||||
|
||||
Use `WindowAction(...)` when the bot needs direct inventory clicks or dropping behavior.
|
||||
|
||||
```csharp
|
||||
private void DropAllOfType(ItemType itemType)
|
||||
{
|
||||
var inventory = GetPlayerInventory();
|
||||
|
||||
foreach (int slot in inventory.SearchItem(itemType))
|
||||
WindowAction(0, slot, WindowActionType.DropItemStack);
|
||||
}
|
||||
```
|
||||
|
||||
Use this pattern carefully:
|
||||
- verify the correct inventory ID first
|
||||
- prefer reacting to `OnInventoryUpdate(...)` for larger inventory workflows
|
||||
- for crafting or chest workflows, use `GetInventories()` and `CloseInventory(...)` as needed
|
||||
|
||||
## Built-in command bot pattern
|
||||
|
||||
Only use this when the user explicitly asks for a built-in bot.
|
||||
|
||||
```csharp
|
||||
public override void Initialize()
|
||||
{
|
||||
McClient.dispatcher.Register(l => l.Literal("help")
|
||||
.Then(l => l.Literal(CommandName)
|
||||
.Executes(r => OnCommandHelp(r.Source, string.Empty))
|
||||
)
|
||||
);
|
||||
|
||||
McClient.dispatcher.Register(l => l.Literal(CommandName)
|
||||
.Then(l => l.Literal("_help")
|
||||
.Executes(r => OnCommandHelp(r.Source, string.Empty))
|
||||
.Redirect(McClient.dispatcher.GetRoot().GetChild("help").GetChild(CommandName)))
|
||||
);
|
||||
}
|
||||
|
||||
public override void OnUnload()
|
||||
{
|
||||
McClient.dispatcher.Unregister(CommandName);
|
||||
McClient.dispatcher.GetRoot().GetChild("help").RemoveChild(CommandName);
|
||||
}
|
||||
```
|
||||
|
||||
Use a built-in bot only when the user explicitly asks for compiled MCC behavior or repo wiring.
|
||||
362
.skills/mcc-dev-workflow/SKILL.md
Normal file
362
.skills/mcc-dev-workflow/SKILL.md
Normal file
|
|
@ -0,0 +1,362 @@
|
|||
---
|
||||
name: mcc-dev-workflow
|
||||
description: Build, run, and debug Minecraft Console Client (MCC) against a real local Minecraft Java server on Linux, macOS, or WSL. Use this whenever the user wants to compile MCC, start or inspect a local test server, connect MCC to a server, debug protocol or login issues, validate a code change end-to-end, or run MCC commands on a real server instead of guessing from static code.
|
||||
---
|
||||
|
||||
# MCC Development Workflow
|
||||
|
||||
Use this skill when the task needs a real local server loop, not just code reading.
|
||||
|
||||
## Defaults
|
||||
|
||||
- Solution: `MinecraftClient.sln`
|
||||
- Runtime target: `.NET 10` / `net10.0`
|
||||
- Environment: Linux, macOS, or WSL with Java, tmux, python3, and dotnet available
|
||||
- Default server root after `source tools/mcc-env.sh`: `${MCC_SERVERS:-<repo>/MinecraftOfficial/downloads}`
|
||||
- Default validation target when the user does not specify a version: `1.21.11`
|
||||
|
||||
## Console modes
|
||||
|
||||
MCC supports two console modes selectable via `ConsoleMode` in `[Console.General]`:
|
||||
|
||||
| Mode | Backend | Best for |
|
||||
|------|---------|----------|
|
||||
| `classic` | `ClassicConsoleBackend` (ConsoleInteractive) | Normal use, legacy CI/scripts, `FileInput` mode |
|
||||
| `tui` | `TuiConsoleBackend` (Avalonia/Consolonia) | Full-screen TUI with scrollable log, command input, popup inventory |
|
||||
|
||||
Both modes support the same commands and input/output through `ConsoleIO.Backend`. The mode is determined at startup from config; `BasicIO` CLI arg overrides to simple stdio.
|
||||
|
||||
## Core rules
|
||||
|
||||
- Prefer a real local server over static reasoning for protocol, login, movement, inventory, entity, or command-path work.
|
||||
- Treat tmux `mc-*` sessions as shared state. Do not run multi-version server workflows in parallel unless the harness explicitly isolates them.
|
||||
- For scripted or repeatable runs, use a generated temporary config. Do not edit the repo-root `MinecraftClient.ini` as part of the test loop.
|
||||
- A server log line containing `Done (` means startup finished. It does not guarantee that RCON is ready on the first attempt. Retry early `mc-rcon` commands.
|
||||
- When instructions, docs, and code disagree, trust current code and current tool behavior first.
|
||||
|
||||
## Shared server, isolated MCC sessions
|
||||
|
||||
- `mc-*` commands operate on the shared local Minecraft server.
|
||||
- `mcc-*` commands operate on one MCC client session.
|
||||
- The default `session` is the current worktree name.
|
||||
- The default username is derived from `session`, unless you pass `--username`.
|
||||
- Session files live under `${TMPDIR:-/tmp}/mcc-debug/<session>/`.
|
||||
- `MCC_SERVERS` stays the shared server-root override.
|
||||
|
||||
Keep shared servers running by default. Do not stop or reset them unless the user explicitly asks for that, or you need to switch server versions.
|
||||
|
||||
Two worktrees can debug against one shared server like this:
|
||||
|
||||
```bash
|
||||
# worktree A
|
||||
cd ~/Minecraft/Minecraft-Console-Client
|
||||
source tools/mcc-env.sh
|
||||
mc-start 1.21.11
|
||||
mcc-debug -v 1.21.11 --file-input
|
||||
|
||||
# worktree B
|
||||
cd ~/Minecraft/Minecraft-Console-Client-foo
|
||||
source tools/mcc-env.sh
|
||||
mcc-debug -v 1.21.11 --file-input
|
||||
|
||||
# from each worktree, mcc-* targets that worktree's default session
|
||||
mcc-state
|
||||
```
|
||||
|
||||
If you want two MCC sessions from the same worktree, pass `--session NAME` explicitly.
|
||||
|
||||
## tmpfs build mode
|
||||
|
||||
Use this on machines with enough RAM when you want worktree-isolated builds outside the repo tree:
|
||||
|
||||
```bash
|
||||
source tools/mcc-env.sh
|
||||
export MCC_BUILD_MODE=tmpfs
|
||||
mcc-build
|
||||
mcc-build-clean
|
||||
```
|
||||
|
||||
`MCC_BUILD_MODE=tmpfs` redirects build output to `/dev/shm/mcc-build/<worktree>/` on Linux, or `${TMPDIR:-/tmp}/mcc-build/<worktree>/` if `/dev/shm` is unavailable.
|
||||
|
||||
## Preflight and reset
|
||||
|
||||
Before scripted runs, especially on macOS or in a reused tmux environment:
|
||||
|
||||
```bash
|
||||
source tools/mcc-env.sh
|
||||
mcc-preflight 1.21.11
|
||||
mc-reset-test-env 1.21.11
|
||||
```
|
||||
|
||||
`mcc-preflight` checks Java, tmux, dotnet, python3, and server directories. It also resolves common Homebrew Java paths on macOS. `mc-reset-test-env` clears stale tmux sessions and stale `stdin.pipe` files before they turn into misleading startup failures.
|
||||
|
||||
## Build
|
||||
|
||||
```bash
|
||||
source tools/mcc-env.sh
|
||||
mcc-build
|
||||
```
|
||||
|
||||
Use `mcc-build` for normal local development so any `MCC_BUILD_MODE=tmpfs` routing stays active. Only use raw `dotnet build` when you are intentionally debugging the build system itself.
|
||||
|
||||
## Server management
|
||||
|
||||
Interactive shell:
|
||||
|
||||
```bash
|
||||
source tools/mcc-env.sh
|
||||
SESSION="$(_mcc_resolve_session)"
|
||||
USERNAME="$(_mcc_resolve_username "$SESSION")"
|
||||
mc-start 1.21.11
|
||||
mc-log 1.21.11 100
|
||||
mc-rcon "op $USERNAME"
|
||||
mc-stop 1.21.11
|
||||
```
|
||||
|
||||
Non-interactive shell:
|
||||
|
||||
```bash
|
||||
tools/start-server.sh 1.21.11
|
||||
tools/mc-rcon.sh "op mcc_smoke_a"
|
||||
```
|
||||
|
||||
If the servers live outside the repo, set `MCC_SERVERS` before sourcing or invoking the tools:
|
||||
|
||||
```bash
|
||||
export MCC_SERVERS=/home/anon/Minecraft/Servers
|
||||
source tools/mcc-env.sh
|
||||
```
|
||||
|
||||
## One-step debug session (recommended)
|
||||
|
||||
The `tools/mcc-debug.sh` script handles build, server startup, config preparation, and MCC launch in one step:
|
||||
|
||||
```bash
|
||||
source tools/mcc-env.sh
|
||||
|
||||
# Classic mode with FileInput (script-driven debugging):
|
||||
mcc-debug -v 1.21.11 --file-input
|
||||
|
||||
# Classic mode interactive (attach via tmux):
|
||||
mcc-debug -v 1.21.11
|
||||
|
||||
# TUI mode:
|
||||
mcc-debug -v 1.21.11 -m tui
|
||||
|
||||
# With debug messages enabled from start:
|
||||
mcc-debug -v 1.21.11 --file-input --debug-on
|
||||
|
||||
# Skip build (already built):
|
||||
mcc-debug -v 1.21.11 --file-input --no-build
|
||||
```
|
||||
|
||||
### What mcc-debug.sh does
|
||||
|
||||
1. Builds MCC (unless `--no-build`)
|
||||
2. Creates a clean temp config at `${TMPDIR:-/tmp}/mcc-debug/<session>/MinecraftClient.debug.ini`
|
||||
3. Ensures server is running (starts if not, waits for `Done (`)
|
||||
4. Launches MCC in a session-scoped tmux session and session-scoped log/input/pid files
|
||||
|
||||
### After launch
|
||||
|
||||
- **FileInput mode**: drive MCC via `mcc-cmd --session smoke-a "debug state"`, or just `mcc-cmd "debug state"` from the same worktree
|
||||
- **Interactive/TUI mode**: attach with `tmux attach -t mcc-<session>`
|
||||
- **Logs**: `mcc-log-mcc --session smoke-a` or `tail -f "${TMPDIR:-/tmp}/mcc-debug/<session>/mcc-debug.log"`
|
||||
- **Server RCON**: grant op or gamemode to the username derived from that session
|
||||
|
||||
## Debug commands (in-game)
|
||||
|
||||
### `/debug [on|off]`
|
||||
|
||||
Toggles debug logging. Now correctly syncs both `Settings.Config.Logging.DebugMessages` and `McClient.Log.DebugEnabled`.
|
||||
|
||||
### `/debug state`
|
||||
|
||||
Prints a one-shot summary of MCC's internal state:
|
||||
|
||||
```
|
||||
=== MCC Debug State ===
|
||||
Server: localhost:25565
|
||||
Username: mcc_smoke_a
|
||||
Protocol: 774
|
||||
GameMode: 1
|
||||
Health: 20.0
|
||||
Food: 20
|
||||
Location: 0.50, 80.00, 0.50
|
||||
TPS: 20.0
|
||||
Console: ClassicConsoleBackend (or TuiConsoleBackend)
|
||||
Features: Terrain Inventory Entity
|
||||
Debug: ON
|
||||
Bots (3): AutoFishing, FileInputBot, ScriptScheduler
|
||||
Players: 2 online
|
||||
```
|
||||
|
||||
This works in both classic and TUI modes.
|
||||
|
||||
## Classic mode debugging
|
||||
|
||||
### Agent workflow (FileInput mode)
|
||||
|
||||
For agents calling MCC commands programmatically:
|
||||
|
||||
```bash
|
||||
source tools/mcc-env.sh
|
||||
SESSION="smoke-a"
|
||||
mcc-debug -v 1.21.11 --file-input --session "$SESSION" --no-build
|
||||
|
||||
# Send commands:
|
||||
mcc-cmd --session "$SESSION" "debug state"
|
||||
mcc-cmd --session "$SESSION" "inventory player list"
|
||||
mcc-cmd --session "$SESSION" "entity"
|
||||
|
||||
# Check results:
|
||||
mcc-log-mcc --session "$SESSION"
|
||||
|
||||
# Stop:
|
||||
mcc-cmd --session "$SESSION" "quit"
|
||||
mcc-kill --session "$SESSION"
|
||||
mc-stop 1.21.11
|
||||
```
|
||||
|
||||
### Interactive workflow
|
||||
|
||||
```bash
|
||||
source tools/mcc-env.sh
|
||||
SESSION="live-a"
|
||||
mcc-debug -v 1.21.11 --session "$SESSION"
|
||||
|
||||
# In another terminal:
|
||||
tmux attach -t "mcc-$SESSION"
|
||||
# Type commands directly in MCC console
|
||||
```
|
||||
|
||||
## TUI mode debugging
|
||||
|
||||
TUI mode runs Consolonia full-screen in a tmux session. Key differences:
|
||||
|
||||
1. **No pipe/redirect**: TUI needs a real tty. Cannot `| tee` or redirect stdout.
|
||||
2. **Log output is in-screen**: all output appears in the scrollable log area.
|
||||
3. **Keyboard shortcuts**: PageUp/PageDown scroll, ESC exits.
|
||||
4. **`/debug state`**: the primary way to inspect internal state since external log tailing is not available.
|
||||
5. **Dialog windows**: `/inventui` opens as an overlay dialog instead of a separate screen.
|
||||
|
||||
### Agent workflow for TUI mode
|
||||
|
||||
```bash
|
||||
source tools/mcc-env.sh
|
||||
SESSION="tui-a"
|
||||
mcc-debug -v 1.21.11 -m tui --session "$SESSION" --no-build
|
||||
|
||||
# Cannot use mcc-cmd (no FileInput); must use tmux send-keys:
|
||||
tmux send-keys -t "mcc-$SESSION" "/debug state" Enter
|
||||
|
||||
# Read TUI screen:
|
||||
tmux capture-pane -t "mcc-$SESSION" -p -S -30
|
||||
|
||||
# Stop:
|
||||
tmux send-keys -t "mcc-$SESSION" Escape
|
||||
```
|
||||
|
||||
**Caveat with tmux send-keys and Consolonia**: When sending text containing `/`, the Enter key may need to be sent separately:
|
||||
```bash
|
||||
tmux send-keys -t "mcc-$SESSION" "/inventory player list"
|
||||
tmux send-keys -t "mcc-$SESSION" Enter
|
||||
```
|
||||
|
||||
## mcc-env.sh quick reference
|
||||
|
||||
After `source tools/mcc-env.sh`:
|
||||
|
||||
| Function | Description |
|
||||
|----------|-------------|
|
||||
| `mc-start VER` | Start MC server in tmux |
|
||||
| `mc-stop VER` | Graceful stop via stdin pipe |
|
||||
| `mc-log VER [N]` | Capture last N lines of server output |
|
||||
| `mc-rcon "CMD"` | Send RCON command |
|
||||
| `mc-kill VER` | Force-kill server tmux session |
|
||||
| `mc-list` | List running MC server sessions |
|
||||
| `mc-wait-ready VER [SEC]` | Wait for server `Done (` |
|
||||
| `mc-wait-stop VER [SEC]` | Wait for server shutdown, with force-kill fallback |
|
||||
| `mc-reset-test-env [--all|VER...]` | Reset shared tmux server state and stale pipes |
|
||||
| `mcc-build` | Build MCC |
|
||||
| `mcc-publish --rid <RID>` | Publish MCC with the repo's CI-like defaults |
|
||||
| `mcc-build-clean` | Clear the current worktree's build output |
|
||||
| `mcc-run [--session NAME] [--username NAME] [--port PORT]` | Convenience wrapper for `mcc-debug --file-input --no-build` |
|
||||
| `mcc-tui [--session NAME] [--username NAME] [--port PORT]` | Convenience wrapper for `mcc-debug -m tui --no-build` |
|
||||
| `mcc-cmd [--session NAME] "CMD"` | Append a command to one session's input file |
|
||||
| `mcc-kill [--session NAME]` | Kill one MCC process and session |
|
||||
| `mcc-debug [OPTS]` | One-step debug session (see above) |
|
||||
| `mcc-log-mcc [--session NAME]` | Tail one MCC debug log |
|
||||
| `mcc-state [--session NAME]` | Send `debug state` and print the last 30 log lines |
|
||||
| `mcc-preflight [VER...]` | Verify Java, tmux, dotnet, python3, and server dirs |
|
||||
|
||||
## Temporary config recipe
|
||||
|
||||
```bash
|
||||
source tools/mcc-env.sh
|
||||
SESSION="smoke-a"
|
||||
USERNAME="$(_mcc_resolve_username "$SESSION")"
|
||||
CFG="$(_mcc_session_root "$SESSION")/MinecraftClient.debug.ini"
|
||||
mkdir -p "$(_mcc_session_root "$SESSION")"
|
||||
bash "$MCC_REPO/.skills/mcc-integration-testing/scripts/prepare_offline_mcc_config.sh" \
|
||||
"$CFG" \
|
||||
"1.21.11" \
|
||||
"$USERNAME"
|
||||
```
|
||||
|
||||
For TUI mode, also add:
|
||||
```bash
|
||||
sed -i 's/ConsoleMode = "classic"/ConsoleMode = "tui"/' "$CFG"
|
||||
```
|
||||
|
||||
## Verify connection and a basic command
|
||||
|
||||
MCC output should include:
|
||||
|
||||
- `[MCC] Server was successfully joined.`
|
||||
|
||||
Server output should include the session-derived username, for example:
|
||||
|
||||
- `mcc_smoke_a joined the game`
|
||||
|
||||
Basic command check:
|
||||
|
||||
```bash
|
||||
mcc-cmd --session smoke-a "inventory player list"
|
||||
```
|
||||
|
||||
If a scripted run fails before MCC joins, check for a harness problem before assuming a product regression. Missing `mcc.log`, a pre-join `Connection refused`, or a server that never reached `Done (` usually means shared-state cleanup or startup failed.
|
||||
|
||||
## Typical debug loop
|
||||
|
||||
1. `source tools/mcc-env.sh`
|
||||
2. `mcc-debug -v 1.21.11 --file-input` (or `-m tui`)
|
||||
3. Confirm `Server was successfully joined` in log
|
||||
4. `mcc-cmd "debug state"` to verify MCC state
|
||||
5. Run test commands
|
||||
6. Inspect log output
|
||||
7. `mcc-cmd "quit"` and `mc-stop 1.21.11`
|
||||
8. Edit code, rebuild, repeat
|
||||
|
||||
## Debugging tips
|
||||
|
||||
- **`/debug state` is your primary diagnostic tool** in both modes. Use it first to verify connection, mode, and feature flags.
|
||||
- **`/debug on` now correctly enables debug logging** at runtime. Previous versions had a bug where `Log.DebugEnabled` was not synced.
|
||||
- Protocol mismatches usually show up as a version line such as `Server version : 1.21.11 (protocol vNNN)` before the failure.
|
||||
- If an early `mc-rcon` command fails, retry it before assuming the server setup is broken.
|
||||
- If a supposedly isolated run behaves strangely, check `tmux list-sessions` and kill stale `mc-*` sessions first.
|
||||
- Legacy `1.8` and `1.8.9` servers may need `use-native-transport=false` in `server.properties` on some Linux environments.
|
||||
- For timing-sensitive work, do not trust wall-clock intuition. Use a real server run and capture evidence from logs or test scripts.
|
||||
- **TUI mode tip**: if the terminal becomes unresponsive after a crash, run `stty sane && reset` to restore it.
|
||||
- **tmux capture trick**: `tmux capture-pane -t mcc-<session> -p -S -50` captures the last 50 lines of a tmux session without attaching.
|
||||
|
||||
## Tool files
|
||||
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `tools/mcc-env.sh` | Shell functions for server/MCC management |
|
||||
| `tools/mcc-debug.sh` | One-step debug session launcher |
|
||||
| `tools/mcc-log-tail.sh` | Log tailing for MCC and/or server |
|
||||
| `tools/start-server.sh` | Server lifecycle in tmux |
|
||||
| `tools/mc-rcon.sh` | RCON command sender |
|
||||
| `tools/run-creative-e2e.sh` | Full creative mode end-to-end test |
|
||||
41
.skills/mcc-dev-workflow/evals/evals.json
Normal file
41
.skills/mcc-dev-workflow/evals/evals.json
Normal file
|
|
@ -0,0 +1,41 @@
|
|||
{
|
||||
"skill_name": "mcc-dev-workflow",
|
||||
"evals": [
|
||||
{
|
||||
"id": 1,
|
||||
"prompt": "Build MCC, start the local 1.21.11 vanilla server from an MCC_SERVERS root, connect MCC with a temporary config, and verify a successful join plus one inventory command.",
|
||||
"expected_output": "The workflow uses a real local server, a temp MCC config, and concrete log evidence for both the join and the MCC command.",
|
||||
"files": [],
|
||||
"expectations": [
|
||||
"The workflow uses a real local 1.21.11 server instead of only reading code.",
|
||||
"The workflow uses MCC_SERVERS-aware tooling or documents the server root explicitly.",
|
||||
"The workflow uses a temporary MCC config instead of relying on the repo-root config.",
|
||||
"The result includes join evidence from MCC output and the server log."
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"prompt": "Debug a flaky local MCC startup on 1.21.11 by checking for stale tmux server sessions, waiting for server readiness, and retrying early RCON commands before blaming protocol code.",
|
||||
"expected_output": "The response treats tmux sessions as shared state, distinguishes server startup from RCON readiness, and uses the real local workflow rather than pure speculation.",
|
||||
"files": [],
|
||||
"expectations": [
|
||||
"The workflow checks or mentions stale mc-* tmux sessions.",
|
||||
"The workflow distinguishes Done from RCON readiness.",
|
||||
"The workflow retries or advises retrying early RCON commands.",
|
||||
"The workflow keeps the debugging loop grounded in real local commands."
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"prompt": "Run a repeatable local MCC debug loop for 1.21.11 that compiles the client, connects with movement, inventory, and entity handling enabled, and leaves enough evidence to inspect a regression afterward.",
|
||||
"expected_output": "The response follows a real build-run-test-inspect loop and captures enough log evidence to support follow-up debugging.",
|
||||
"files": [],
|
||||
"expectations": [
|
||||
"The workflow builds MCC before the run.",
|
||||
"The workflow enables terrain, inventory, and entity handling for the scripted run.",
|
||||
"The workflow captures or points to concrete log locations.",
|
||||
"The workflow prefers a temp config for repeatability."
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
323
.skills/mcc-integration-testing/SKILL.md
Normal file
323
.skills/mcc-integration-testing/SKILL.md
Normal file
|
|
@ -0,0 +1,323 @@
|
|||
---
|
||||
name: mcc-integration-testing
|
||||
description: >-
|
||||
Use when proving MCC behavior on a real local Minecraft server, validating
|
||||
runtime or protocol changes end-to-end, exercising movement, physics,
|
||||
inventory, entity, chat, or terrain behavior, or running a single-version or
|
||||
cross-version regression sweep.
|
||||
metadata:
|
||||
category: discipline
|
||||
triggers:
|
||||
- integration test
|
||||
- real server
|
||||
- local server
|
||||
- regression sweep
|
||||
- rcon
|
||||
- tmux
|
||||
- offline mode
|
||||
- online mode
|
||||
- movement
|
||||
- physics
|
||||
- inventory
|
||||
- entity
|
||||
- terrain
|
||||
- chat
|
||||
---
|
||||
|
||||
# MCC Integration Testing
|
||||
|
||||
Use this skill when the task is "prove it on a real server", not just "reason about whether it should work."
|
||||
|
||||
Read [references/online-mode.md](references/online-mode.md) when the user asks for Microsoft login, device-code auth, or an online-mode server run. Use [references/command-matrix.md](references/command-matrix.md) for stable MCC-side and RCON-side commands.
|
||||
|
||||
## Iron Law
|
||||
|
||||
Only say MCC was integration tested when MCC ran against a real local server and the claim is backed by real MCC output plus real server logs.
|
||||
|
||||
Calling build-only, reasoning-only, or join-only work "integration tested" is a rules violation, not shorthand.
|
||||
|
||||
These do not count as end-to-end proof:
|
||||
|
||||
- static reasoning, source comparison, or build success
|
||||
- join or login success by itself
|
||||
- a long-lived idle connection by itself
|
||||
- a grep that only says there were no errors
|
||||
- testing one shared-route version and silently claiming adjacent versions also passed
|
||||
|
||||
If the environment cannot run a real server, say so and report the result as unexecuted or inferred, not integration tested.
|
||||
|
||||
## Default target
|
||||
|
||||
- Use `1.21.11-Vanilla` unless the user asks for a different version or a version matrix.
|
||||
- Use `MCC_SERVERS` if it is set. Otherwise the default server root is `MinecraftOfficial/downloads`.
|
||||
|
||||
## Guardrails
|
||||
|
||||
- Use a real local server.
|
||||
- Launch MCC against an explicit `localhost:<server-port>` target for repeatable local tests.
|
||||
- Keep version matrices sequential in shared local environments. The tmux server harness is shared state by default.
|
||||
- Prefer generated temporary MCC configs for scripted runs so one test does not contaminate the next.
|
||||
- Default to offline auth in generated temp configs. Do not trust the repo-root `MinecraftClient.ini` account defaults.
|
||||
- If the user explicitly asks for Microsoft online login, honor that request and generate the temp config for Microsoft auth instead of offline mode.
|
||||
- For Microsoft auth, prefer an interactive TTY launch with `BasicIO-NoColor` so the device code is easy to read and relay to the user.
|
||||
- Do not use file-input mode during Microsoft auth. Launch interactively first, complete login, then switch to scripted control only if needed.
|
||||
- For online-mode tests, prefer a clean temp config with no join-time bots or scheduled tasks. Inherited `ScriptScheduler` or `DiscordRpc` settings can pollute the session and send unintended chat right after login.
|
||||
- Legacy and modern command syntax differ. Do not assume one server-command profile fits every version.
|
||||
- Use actual MCC output and actual server logs for assertions. Do not invent success strings.
|
||||
- Treat server `Done` as startup progress, not RCON readiness. Retry the first RCON command before assuming the setup is broken.
|
||||
- Run preflight before scripted test loops. On macOS, Java may be installed but not exported on PATH in the shell the harness uses.
|
||||
- If a change touches shared routing or a version range, test at least one adjacent version that shares that path, or explicitly mark adjacent versions as unexecuted and inferred.
|
||||
- For palette or version-content changes, probe at least one neighboring or existing item, entity, or block. Do not only check the headline addition.
|
||||
- Separate product failures from harness failures. Missing logs, stale tmux state, stale `stdin.pipe`, or pre-join `Connection refused` errors are usually environment problems until proven otherwise.
|
||||
|
||||
## Choose the test mode
|
||||
|
||||
### 1. Single-version deep smoke
|
||||
|
||||
Use this when one supported version is enough and you want broad coverage:
|
||||
|
||||
```bash
|
||||
.skills/mcc-integration-testing/scripts/run_full_spectrum_test.sh 1.21.11-Vanilla
|
||||
```
|
||||
|
||||
This covers join, chat, slash commands, internal MCC commands, creative inventory, entity handling, sounds, particles, TNT, kill/respawn, and log assertions.
|
||||
|
||||
### 2. Ordered creative-mode E2E
|
||||
|
||||
Use this when the user asks for a regression sweep in a strict scenario order such as:
|
||||
|
||||
- connect
|
||||
- send messages
|
||||
- send commands
|
||||
- receive messages
|
||||
- movement
|
||||
- physics
|
||||
- mobs
|
||||
- effects
|
||||
- inventory
|
||||
|
||||
Broad validation should usually cover `connect-test`, `item-test`, `entity-test`, `terrain-test`, and `chat-test`.
|
||||
|
||||
Command:
|
||||
|
||||
```bash
|
||||
MCC_SERVERS=/home/anon/Minecraft/Servers bash tools/run-creative-e2e.sh 1.21.11-Vanilla 1.21.11 modern
|
||||
```
|
||||
|
||||
For legacy targets such as `1.8` or `1.8.9`, switch the final argument to `legacy` and pass the pinned MC version.
|
||||
|
||||
### 3. Timing or cadence validation
|
||||
|
||||
Use this for TPS, movement-cadence, or packet-cadence work:
|
||||
|
||||
- `MinecraftClient/config/sample-script-tick-counter.cs`
|
||||
- `MinecraftClient/config/sample-script-packet-capture.cs`
|
||||
|
||||
Run them against a real server with a temp config and summarize counts from the captured logs.
|
||||
|
||||
### 4. Structured components test
|
||||
|
||||
Use this after touching any `StructuredComponents` code (registries, component
|
||||
parsers, subcomponents, codec helpers) to prove every component type in a
|
||||
version parses on the wire without error:
|
||||
|
||||
```bash
|
||||
bash tools/run-structured-components-test.sh 1.21.11
|
||||
```
|
||||
|
||||
Run a single version (fast, ~2 min) or a matrix:
|
||||
|
||||
```bash
|
||||
for v in 1.20.6 1.21 1.21.2 1.21.5 1.21.11 26.1; do
|
||||
bash tools/run-structured-components-test.sh "$v"
|
||||
done
|
||||
```
|
||||
|
||||
The script gives items with every registered component via RCON `/give`, reads
|
||||
them back with `inventory player list`, and asserts no parse errors in the MCC
|
||||
log. Version-gated components (v1212+, v1215+, v12111+, v261) are tested only
|
||||
on the versions that support them. See `SC_Integration_Test_Report.md` for a
|
||||
reference run across all 6 version groups.
|
||||
|
||||
### 6. Dialog integration test
|
||||
|
||||
Use this after touching any dialog system code (packet handling, NBT parsing,
|
||||
models, TUI, command dispatch, the state machine in `DialogManager`, or the
|
||||
codec in `DialogNbtParser`). Tests all 5 dialog types, button actions (close,
|
||||
run_command, show_dialog), cancel/dismiss, click-label, and body content:
|
||||
|
||||
```bash
|
||||
tools/run-dialog-test.sh 26.1
|
||||
```
|
||||
|
||||
The script starts the server if needed, generates a temp MCC config, launches
|
||||
MCC with file-input mode (requires both `MCC_FILE_INPUT=1` and
|
||||
`MCC_INPUT_FILE=<path>` env vars), sends inline SNBT dialogs via RCON, and
|
||||
asserts 29 checks against the MCC log.
|
||||
|
||||
Key requirements that differ from other test modes:
|
||||
|
||||
- FileInputBot is loaded only when `MCC_FILE_INPUT=1` is set in the
|
||||
environment. The `[ChatBot.FileInput]` config section is ignored at load
|
||||
time.
|
||||
- The input file path is controlled by `MCC_INPUT_FILE`, *not* by the config
|
||||
`File` setting.
|
||||
- Dialogs use inline SNBT syntax through `ResourceOrIdArgument`, e.g.:
|
||||
`dialog show <player> {type:"minecraft:notice", title:{text:"Hello"}}`
|
||||
- The `ActionButton.CODEC` flattens `CommonButtonData` fields (`label`,
|
||||
`tooltip`, `width`) into the same object as `action` — no `button` wrapper.
|
||||
|
||||
### 5. Full inventory regression sweep
|
||||
|
||||
Use this when touching inventory snapshots, player/container slot sync, creative inventory, item-slot serialization, packet palettes, game-mode updates, or block-use paths that open containers:
|
||||
|
||||
```bash
|
||||
tools/run-inventory-full-sweep.sh --versions "1.21.10 1.21.11"
|
||||
```
|
||||
|
||||
Default coverage includes:
|
||||
|
||||
- player inventory listing and inventory discovery
|
||||
- creative give/delete
|
||||
- inventory search
|
||||
- player right/left click stack split and merge
|
||||
- player drop one and drop all
|
||||
- chest open via `useblock`
|
||||
- container listing and close
|
||||
- mirrored player slots in container windows
|
||||
- shift-click and shift-right-click transfer
|
||||
- container right/left click, cursor stack, drop one, and drop all
|
||||
- creative middle-click command path
|
||||
- log scan for packet parse failures, queue-empty crashes, unhandled exceptions, and disconnects
|
||||
|
||||
The script writes `summary.tsv` under `RUN_ROOT` and per-version logs under `/tmp/mcc-debug/inventory-full-<version>/mcc-debug.log`.
|
||||
|
||||
When a matrix has existing PASS rows, do not rerun them unless a later code change affects that row or the user asks for a full rerun. Derive remaining rows from summaries:
|
||||
|
||||
```bash
|
||||
awk 'FNR>1 && $2=="PASS" {print $1}' /tmp/mcc-inventory-full-sweep/*/summary.tsv | sort -V | uniq
|
||||
```
|
||||
|
||||
## Preconditions
|
||||
|
||||
Before running any scenario:
|
||||
|
||||
0. run preflight and clear stale shared state when the environment is reused
|
||||
1. configure the target server for offline testing
|
||||
2. ensure `eula=true`
|
||||
3. ensure RCON is enabled
|
||||
4. build MCC unless the task explicitly reuses a fresh build
|
||||
|
||||
Preflight and reset helpers:
|
||||
|
||||
```bash
|
||||
.skills/mcc-integration-testing/scripts/preflight_test_env.sh 1.21.11-Vanilla
|
||||
.skills/mcc-integration-testing/scripts/reset_shared_test_state.sh 1.21.11-Vanilla
|
||||
```
|
||||
|
||||
Offline configuration helper:
|
||||
|
||||
```bash
|
||||
.skills/mcc-integration-testing/scripts/ensure_offline_server.sh 1.21.11-Vanilla
|
||||
```
|
||||
|
||||
By default, the config helper prepares offline auth. To opt into another auth mode for a specific run, set:
|
||||
|
||||
```bash
|
||||
MCC_TEST_ACCOUNT_TYPE=microsoft
|
||||
MCC_TEST_PASSWORD=
|
||||
```
|
||||
|
||||
Optionally override the login name with the fourth argument to the config helper.
|
||||
|
||||
## Scripts and tools
|
||||
|
||||
- `.skills/mcc-integration-testing/scripts/ensure_offline_server.sh`
|
||||
- configures persistent offline mode and RCON
|
||||
- `.skills/mcc-integration-testing/scripts/preflight_test_env.sh`
|
||||
- verifies Java, tmux, dotnet, python3, server directories, and resolves common Java PATH issues
|
||||
- `.skills/mcc-integration-testing/scripts/reset_shared_test_state.sh`
|
||||
- clears stale tmux sessions and stale `stdin.pipe` files before a rerun
|
||||
- `.skills/mcc-integration-testing/scripts/prepare_offline_mcc_config.sh`
|
||||
- generates a clean temporary MCC config, prepares offline login by default, disables noisy bots, and can switch to Microsoft auth when explicitly requested
|
||||
- `.skills/mcc-integration-testing/scripts/get_server_port.sh`
|
||||
- resolves the actual local server port from `server.properties` or the latest server log
|
||||
- `.skills/mcc-integration-testing/scripts/run_full_spectrum_test.sh`
|
||||
- single-version deep smoke with built-in assertions
|
||||
- `.skills/mcc-integration-testing/scripts/summarize_test_run.sh`
|
||||
- summarize the latest full-spectrum run
|
||||
- `tools/run-creative-e2e.sh`
|
||||
- ordered creative-mode E2E regression scenario
|
||||
- `tools/run-inventory-full-sweep.sh`
|
||||
- full inventory command/API sweep across one or more versions
|
||||
- `tools/run-structured-components-test.sh`
|
||||
- exercises every structured component via RCON `/give` across versions 1.20.6-26.1
|
||||
- `tools/run-dialog-test.sh`
|
||||
- dialog integration test: all 5 types, run_command/show_dialog actions,
|
||||
cancel/dismiss/click-label, body content; 29 assertions on MCC log
|
||||
|
||||
## Evidence Discipline
|
||||
|
||||
In every report, separate:
|
||||
|
||||
- `Executed`: exact scripts, commands, versions, auth mode, and whether the run was sequential or single-version
|
||||
- `Observed`: exact MCC output, exact server-log evidence, and the saved log directory
|
||||
- `Inferred`: conclusions not directly shown by that run's runtime evidence
|
||||
- `Harness issues`: setup or runner problems such as missing Java on PATH, stale tmux sessions, stale `stdin.pipe`, missing log artifacts, or failed config generation
|
||||
|
||||
Never upgrade inferred claims to observed facts. Absence of errors is supporting evidence only; pair it with a positive assertion for the feature under test.
|
||||
|
||||
## Red Flags
|
||||
|
||||
Stop and fix the test plan if you are about to:
|
||||
|
||||
- claim movement, inventory, entity, terrain, physics, or chat coverage from join success alone
|
||||
- reuse repo-root `MinecraftClient.ini` or another user-local stateful config
|
||||
- run multi-version tests in parallel in a shared tmux or shared server environment
|
||||
- let inherited bots, schedulers, or other user-local noise send chat or commands during validation
|
||||
|
||||
## What to report back
|
||||
|
||||
Always summarize:
|
||||
|
||||
- which version or versions were tested
|
||||
- which port or ports were used
|
||||
- which auth mode and scenario were used
|
||||
- whether the run was sequential or single-version
|
||||
- the exact scripts or commands executed
|
||||
- pass or fail per major phase
|
||||
- concrete evidence from MCC and server logs
|
||||
- the saved log directory
|
||||
- what was not executed and what remains inferred
|
||||
- which adjacent versions were not run but were mentioned
|
||||
|
||||
## When Not to Use
|
||||
|
||||
- build-only verification
|
||||
- static protocol or source comparison with no real server run
|
||||
- documentation or prompt work
|
||||
- code review requests that do not ask for executed runtime proof
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
- If the first RCON command fails, retry it before assuming the setup is broken.
|
||||
- If Java is installed but the harness still says it is missing, run `preflight_test_env.sh`. This resolves common Homebrew Java paths on macOS.
|
||||
- If MCC reaches Microsoft device-code login during an offline test, stop and inspect the generated temp config before retrying.
|
||||
- If the user explicitly requests Microsoft online login, set `MCC_TEST_ACCOUNT_TYPE=microsoft` before launching the harness.
|
||||
- If the user explicitly requests Microsoft online login, use `BasicIO-NoColor` in a real TTY, relay the device code from the TUI, and avoid pressing empty Enter at any auth prompt.
|
||||
- If the online-mode session sends unexpected chat or commands right after join, inspect inherited bot settings first. `ChatBot.ScriptScheduler` tasks and `ChatBot.DiscordRpc` are common sources of test noise in user-local configs.
|
||||
- If `dotnet run` cannot see an existing Microsoft session, check whether `SessionCache.db` and `ProfileKeyCache.ini` need to be synced from `MinecraftClient/bin/Release/net10.0/` to the repo root.
|
||||
- If Microsoft auth keeps prompting even with a valid session cache, verify `Account.Login` matches the cached username exactly.
|
||||
- If MCC reports `Connection refused`, verify the launched target matches the server's actual `server-port`.
|
||||
- If MCC reports `Connection refused` immediately after a server start, also check for stale shared state: old tmux sessions, a stale `stdin.pipe`, or a server that never actually reached `Done (`.
|
||||
- If multiple versions are being tested, do not start them in parallel unless the harness isolates tmux sessions and input files.
|
||||
- If a test assertion fails, inspect the real MCC output before changing the code or weakening the assertion.
|
||||
- If an older server behaves oddly on Linux, check `use-native-transport=false` in `server.properties`.
|
||||
- If a matrix row fails before producing `mcc.log` or a command transcript, treat it as a harness failure, fix the environment, and rerun that row before drawing product conclusions.
|
||||
- If creative inventory commands report "You must be in Creative gamemode" after RCON switched the player, inspect game-mode update parsing before assuming creative inventory is broken. Modern servers can update local game mode through game event reason `3`.
|
||||
- If an inventory row crashes with `Queue empty` or `Failed to process incoming packet`, inspect packet palette routing before changing inventory code. A single shifted packet ID can make a healthy inventory feature look broken.
|
||||
- For chest-open failures, separate product and harness causes. The player may be standing inside the chest or suffocating on older servers. Stand beside the chest, put a floor under the player, and retry `useblock`.
|
||||
- For shared local servers, a `Done` log line does not prove RCON is ready. Retry setup commands and verify the actual RCON port from `server.properties`.
|
||||
- If `tools/run-dialog-test.sh` fails with "FileInput Watching: .../mcc_input.txt" pointing to the wrong directory, the `MCC_INPUT_FILE` env var was not set in the tmux command. FileInputBot ignores the config `File` setting entirely.
|
||||
- If inline SNBT dialogs fail on the server side (`Failed to parse structure: No key ...`), check whether `ActionButton.CODEC` fields are flat (no `button` wrapper) and whether the dialog type fields match the 26.1 server (`label` not `text` in `CommonButtonData`).
|
||||
- If a dialog integration test fails on "Server showed custom dialog", the dialog packet (id=0x8C in 26.1 play phase) may not have been sent. Verify the RCON command succeeded and the server printed "Displayed dialog to ...".
|
||||
41
.skills/mcc-integration-testing/evals/evals.json
Normal file
41
.skills/mcc-integration-testing/evals/evals.json
Normal file
|
|
@ -0,0 +1,41 @@
|
|||
{
|
||||
"skill_name": "mcc-integration-testing",
|
||||
"evals": [
|
||||
{
|
||||
"id": 1,
|
||||
"prompt": "Run a real 1.21.11 MCC regression test in creative mode covering connect, send messages, send commands, receive messages, movement, physics, mobs, effects, and inventory, in that order.",
|
||||
"expected_output": "The workflow uses the ordered creative-mode E2E harness on a real server, reports pass or fail for each phase, and points to the saved logs.",
|
||||
"files": [],
|
||||
"expectations": [
|
||||
"The workflow uses a real 1.21.11 server.",
|
||||
"The workflow uses the ordered creative E2E harness instead of improvising the whole scenario.",
|
||||
"The result reports phase-by-phase outcomes in the requested order.",
|
||||
"The result includes the log directory for the run."
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"prompt": "Validate that MCC still works after a runtime or protocol change by running a deep single-version 1.21.11 integration test with chat, creative inventory, entity tracking, sounds, particles, TNT, and kill/respawn coverage.",
|
||||
"expected_output": "The workflow uses the full-spectrum test runner on a real server and returns a concise pass or fail summary backed by MCC and server log evidence.",
|
||||
"files": [],
|
||||
"expectations": [
|
||||
"The workflow builds MCC before the scenario unless a fresh build is explicitly reused.",
|
||||
"The workflow uses the full-spectrum runner instead of only manual spot checks.",
|
||||
"The result includes evidence from both MCC output and server logs.",
|
||||
"The result points to the saved run directory."
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"prompt": "Validate a timing-sensitive MCC change on a real 1.21.11 server by collecting tick-rate and outbound packet-cadence evidence, then summarize the results clearly.",
|
||||
"expected_output": "The response uses a real server, a temp config, and the provided sample scripts to capture tick-rate and packet evidence instead of relying on intuition.",
|
||||
"files": [],
|
||||
"expectations": [
|
||||
"The workflow uses the real 1.21.11 server.",
|
||||
"The workflow uses the tick counter and packet capture scripts or clearly equivalent targeted instrumentation.",
|
||||
"The workflow keeps the run isolated with a temp config.",
|
||||
"The summary reports concrete counts or cadence evidence."
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
91
.skills/mcc-integration-testing/references/command-matrix.md
Normal file
91
.skills/mcc-integration-testing/references/command-matrix.md
Normal file
|
|
@ -0,0 +1,91 @@
|
|||
# Command Matrix
|
||||
|
||||
This skill uses a fixed set of stable commands for local offline integration testing.
|
||||
|
||||
## MCC-side commands via `mcc-cmd`
|
||||
|
||||
- `health`
|
||||
- `list`
|
||||
- `inventory player list`
|
||||
- `/gamemode creative`
|
||||
- `inventory creativegive 36 Diamond 16`
|
||||
- `inventory creativegive 37 IronSword 1`
|
||||
- `inventory creativegive 38 GoldenApple 8`
|
||||
- `inventory creativeclear 38`
|
||||
- `entity`
|
||||
- `/time query daytime`
|
||||
- `look up`
|
||||
- `look down`
|
||||
- `look east`
|
||||
- `/gamemode survival`
|
||||
- `respawn`
|
||||
- `/tp MCCBot 0 -60 0`
|
||||
- `smoke_test_from_mcc_full_spectrum`
|
||||
- `integration_test_chat_response`
|
||||
|
||||
Notes:
|
||||
- Lines starting with `/` are sent to the server as chat/commands.
|
||||
- Non-slash lines are treated as MCC internal commands first, then fall back to chat.
|
||||
|
||||
## Server-side commands via `mc-rcon`
|
||||
|
||||
- `op MCCBot`
|
||||
- `gamerule sendCommandFeedback true`
|
||||
- `gamerule logAdminCommands true`
|
||||
- `time set day`
|
||||
- `weather clear`
|
||||
- `say Hello from the server console`
|
||||
- `msg MCCBot This is a private whisper`
|
||||
- `effect give MCCBot minecraft:speed 30 1`
|
||||
- `effect give MCCBot minecraft:regeneration 10 1`
|
||||
- `kill MCCBot`
|
||||
|
||||
## Representative entity coverage
|
||||
|
||||
- `execute as MCCBot at @s run summon minecraft:cow ~2 ~ ~`
|
||||
- `execute as MCCBot at @s run summon minecraft:zombie ~4 ~ ~`
|
||||
- `execute as MCCBot at @s run summon minecraft:creeper ~6 ~ ~`
|
||||
- `execute as MCCBot at @s run summon minecraft:skeleton ~8 ~ ~`
|
||||
- `execute as MCCBot at @s run summon minecraft:villager ~-2 ~ ~`
|
||||
- `execute as MCCBot at @s run summon minecraft:allay ~-4 ~ ~`
|
||||
- `execute as MCCBot at @s run summon minecraft:armor_stand ~ ~ ~2`
|
||||
- `execute as MCCBot at @s run summon minecraft:item_display ~-6 ~ ~ {item:{id:"minecraft:diamond",count:1}}`
|
||||
- `execute as MCCBot at @s run summon minecraft:spider ~10 ~ ~`
|
||||
- `execute as MCCBot at @s run summon minecraft:pig ~-8 ~ ~`
|
||||
|
||||
## Block placement coverage
|
||||
|
||||
- `execute as MCCBot at @s run fill ~1 ~ ~1 ~3 ~2 ~3 minecraft:stone`
|
||||
- `execute as MCCBot at @s run setblock ~5 ~ ~5 minecraft:chest`
|
||||
- `execute as MCCBot at @s run setblock ~5 ~1 ~5 minecraft:furnace`
|
||||
- `execute as MCCBot at @s run setblock ~6 ~ ~5 minecraft:crafting_table`
|
||||
|
||||
## Dimension change coverage
|
||||
|
||||
- `execute in minecraft:the_nether run tp MCCBot 0 64 0`
|
||||
- `execute in minecraft:overworld run tp MCCBot 0 -60 0`
|
||||
|
||||
## Representative particle coverage
|
||||
|
||||
- `execute as MCCBot at @s run particle minecraft:happy_villager ~ ~1 ~ 0.5 0.5 0.5 0 12 force`
|
||||
- `execute as MCCBot at @s run particle minecraft:end_rod ~ ~1 ~ 0.5 0.5 0.5 0.01 20 force`
|
||||
- `execute as MCCBot at @s run particle minecraft:explosion ~ ~1 ~ 0 0 0 0 1 force`
|
||||
- `execute as MCCBot at @s run particle minecraft:totem_of_undying ~ ~1 ~ 0.5 0.5 0.5 0.1 20 force`
|
||||
- `execute as MCCBot at @s run particle minecraft:flame ~ ~1 ~ 0.2 0.2 0.2 0.02 30 force`
|
||||
- `execute as MCCBot at @s run particle minecraft:heart ~ ~2 ~ 0.3 0.3 0.3 0 5 force`
|
||||
|
||||
## Representative sound coverage
|
||||
|
||||
- `execute as MCCBot at @s run playsound minecraft:entity.lightning_bolt.thunder master MCCBot ~ ~ ~ 1 1 0`
|
||||
- `execute as MCCBot at @s run playsound minecraft:block.note_block.bell master MCCBot ~ ~ ~ 1 1 0`
|
||||
- `execute as MCCBot at @s run playsound minecraft:entity.experience_orb.pickup master MCCBot ~ ~ ~ 1 1 0`
|
||||
|
||||
## Explosion coverage
|
||||
|
||||
- `execute as MCCBot at @s run summon minecraft:tnt ~3 ~ ~`
|
||||
- `execute as MCCBot at @s run summon minecraft:tnt ~6 ~ ~`
|
||||
|
||||
## Kill and respawn cycle
|
||||
|
||||
- `kill MCCBot` (via RCON, requires survival mode)
|
||||
- `respawn` (via MCC command after death)
|
||||
48
.skills/mcc-integration-testing/references/online-mode.md
Normal file
48
.skills/mcc-integration-testing/references/online-mode.md
Normal file
|
|
@ -0,0 +1,48 @@
|
|||
# Online-Mode Notes
|
||||
|
||||
Use this flow only when the user explicitly asks for Microsoft login or wants to validate against an online-mode server.
|
||||
|
||||
## Launch mode
|
||||
|
||||
- Prefer `BasicIO-NoColor` in a real TTY so the Microsoft device-code prompt is easy to read and copy.
|
||||
- Do not use `MCC_FILE_INPUT=1` during the auth step. It is for scripted command injection, not interactive login.
|
||||
- Avoid `nohup`. Use `tmux` for long-running sessions that still need a TTY.
|
||||
- Start from a clean temp config when possible. If the temp config is copied from a user-local `MinecraftClient.ini`, inspect `ChatBot.ScriptScheduler` and `ChatBot.DiscordRpc` before the run.
|
||||
|
||||
## Session cache behavior
|
||||
|
||||
- `dotnet run --project MinecraftClient ...` uses the repo root for `SessionCache.db` and `ProfileKeyCache.ini`.
|
||||
- The compiled binary under `MinecraftClient/bin/<Config>/net10.0/` uses that output directory instead.
|
||||
- If a session exists in one location and not the other, sync the cache files before assuming login is broken.
|
||||
|
||||
## Account settings
|
||||
|
||||
- `Account.Login` must be populated for MCC to look up a cached Microsoft session.
|
||||
- The cached key is the username form MCC stored, typically the lowercase username, not necessarily the email address.
|
||||
- MCC rewrites `MinecraftClient.ini` on clean exit, so generate a temp config per run and do not edit it while MCC is still running.
|
||||
|
||||
## Auth prompt handling
|
||||
|
||||
- Do not send a bare Enter to dismiss `Password(invisible):` or `Paste your code here:` prompts. That can trigger offline fallback.
|
||||
- For interactive online-mode runs, wait for the device code prompt and relay the code to the user exactly as shown.
|
||||
- After the user completes login, continue the test in the same TTY session or restart into file-driven mode if the workflow requires automation.
|
||||
|
||||
## Join-time noise
|
||||
|
||||
- Real user configs may contain enabled bots or task lists that were harmless in offline testing but are noisy in online-mode validation.
|
||||
- The most common examples are:
|
||||
- `ChatBot.ScriptScheduler` task lists that send `/hello`, `/login ...`, or other automatic commands on login or on an interval
|
||||
- `ChatBot.DiscordRpc`, which is not harmful to server state but adds log noise and extra background activity
|
||||
- If the goal is protocol or feature validation, suppress these before the run or treat their output as non-test noise.
|
||||
|
||||
## Server settings
|
||||
|
||||
- For realistic online-mode testing, keep `online-mode=true`.
|
||||
- Keep `enforce-secure-profile=true` unless the test explicitly targets insecure-profile behavior.
|
||||
|
||||
## Command reminders
|
||||
|
||||
- With `InternalCmdChar = "slash"`:
|
||||
- `/health`, `/pos`, `/inventory`, `/entity` are MCC internal commands.
|
||||
- `/send /list` and `/send /give ...` are server commands.
|
||||
- bare text is regular chat sent to the server.
|
||||
110
.skills/mcc-integration-testing/scripts/common.sh
Executable file
110
.skills/mcc-integration-testing/scripts/common.sh
Executable file
|
|
@ -0,0 +1,110 @@
|
|||
#!/usr/bin/env bash
|
||||
|
||||
sed_in_place() {
|
||||
if [[ "$(uname)" == "Darwin" ]]; then
|
||||
sed -i '' "$@"
|
||||
else
|
||||
sed -i "$@"
|
||||
fi
|
||||
}
|
||||
|
||||
ensure_java_in_path() {
|
||||
if command -v java >/dev/null 2>&1 && java -version >/dev/null 2>&1; then
|
||||
return 0
|
||||
fi
|
||||
|
||||
local candidate
|
||||
for candidate in \
|
||||
"${JAVA_BIN:-}" \
|
||||
"/opt/homebrew/opt/openjdk/bin/java" \
|
||||
"/usr/local/opt/openjdk/bin/java" \
|
||||
"/usr/lib/jvm/default-java/bin/java"
|
||||
do
|
||||
[[ -z "$candidate" ]] && continue
|
||||
if [[ -x "$candidate" ]]; then
|
||||
export PATH="$(dirname "$candidate"):$PATH"
|
||||
export JAVA_BIN="$candidate"
|
||||
if java -version >/dev/null 2>&1; then
|
||||
return 0
|
||||
fi
|
||||
fi
|
||||
done
|
||||
|
||||
echo "java was not found on PATH. Install Java or set JAVA_BIN." >&2
|
||||
return 1
|
||||
}
|
||||
|
||||
server_session_name() {
|
||||
printf 'mc-%s\n' "${1//./_}"
|
||||
}
|
||||
|
||||
server_running() {
|
||||
local version="$1"
|
||||
mc-list | grep -Fq "$(server_session_name "$version")"
|
||||
}
|
||||
|
||||
wait_for_server_ready() {
|
||||
local version="$1"
|
||||
local timeout="${2:-60}"
|
||||
local elapsed=0
|
||||
|
||||
while (( elapsed < timeout )); do
|
||||
if mc-log "$version" 250 2>/dev/null | grep -Fq "Done ("; then
|
||||
return 0
|
||||
fi
|
||||
sleep 1
|
||||
((elapsed += 1))
|
||||
done
|
||||
|
||||
echo "Timed out waiting for $version to become ready" >&2
|
||||
return 1
|
||||
}
|
||||
|
||||
wait_for_server_stop() {
|
||||
local version="$1"
|
||||
local timeout="${2:-60}"
|
||||
local elapsed=0
|
||||
|
||||
while (( elapsed < timeout )); do
|
||||
if ! server_running "$version"; then
|
||||
return 0
|
||||
fi
|
||||
sleep 1
|
||||
((elapsed += 1))
|
||||
done
|
||||
|
||||
mc-kill "$version" --confirm >/dev/null 2>&1 || true
|
||||
|
||||
if ! server_running "$version"; then
|
||||
return 0
|
||||
fi
|
||||
|
||||
echo "Timed out waiting for $version to stop" >&2
|
||||
return 1
|
||||
}
|
||||
|
||||
disable_noisy_bots_in_ini() {
|
||||
local ini_file="$1"
|
||||
local section
|
||||
|
||||
for section in \
|
||||
ScriptScheduler \
|
||||
DiscordRpc \
|
||||
AntiAFK \
|
||||
AutoDig \
|
||||
AutoAttack \
|
||||
PlayerListLogger \
|
||||
ReplayCapture
|
||||
do
|
||||
sed_in_place "/^\\[ChatBot\\.${section}\\]/,/^\\[/ { s/^Enabled = true/Enabled = false/; }" "$ini_file"
|
||||
done
|
||||
}
|
||||
|
||||
remove_stale_stdin_pipe() {
|
||||
local version="$1"
|
||||
local pipe_path="$MCC_SERVERS/$version/stdin.pipe"
|
||||
|
||||
if [[ -e "$pipe_path" ]] && ! server_running "$version"; then
|
||||
rm -f "$pipe_path"
|
||||
fi
|
||||
}
|
||||
59
.skills/mcc-integration-testing/scripts/ensure_offline_server.sh
Executable file
59
.skills/mcc-integration-testing/scripts/ensure_offline_server.sh
Executable file
|
|
@ -0,0 +1,59 @@
|
|||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
REPO_ROOT="$(cd "$SCRIPT_DIR/../../.." && pwd)"
|
||||
# shellcheck source=tools/mcc-env.sh
|
||||
source "$REPO_ROOT/tools/mcc-env.sh"
|
||||
# shellcheck source=.skills/mcc-integration-testing/scripts/common.sh
|
||||
source "$SCRIPT_DIR/common.sh"
|
||||
|
||||
VERSION="${1:-1.21.11-Vanilla}"
|
||||
SERVER_DIR="${MCC_SERVERS:?}/$VERSION"
|
||||
PROPS_FILE="$SERVER_DIR/server.properties"
|
||||
SESSION_NAME="mc-${VERSION//./_}"
|
||||
|
||||
if [[ ! -d "$SERVER_DIR" ]]; then
|
||||
echo "Server directory not found: $SERVER_DIR" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if [[ ! -f "$SERVER_DIR/eula.txt" ]] || ! grep -Eq '^eula=true$' "$SERVER_DIR/eula.txt"; then
|
||||
echo "Missing accepted EULA in $SERVER_DIR/eula.txt" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
server_running() {
|
||||
mc-list | grep -Fq "$SESSION_NAME"
|
||||
}
|
||||
|
||||
upsert_property() {
|
||||
local key="$1"
|
||||
local value="$2"
|
||||
|
||||
if grep -Eq "^${key}=" "$PROPS_FILE"; then
|
||||
sed_in_place "s#^${key}=.*#${key}=${value}#" "$PROPS_FILE"
|
||||
else
|
||||
printf '%s=%s\n' "$key" "$value" >> "$PROPS_FILE"
|
||||
fi
|
||||
}
|
||||
|
||||
if [[ ! -f "$PROPS_FILE" ]]; then
|
||||
mc-start "$VERSION"
|
||||
wait_for_server_ready "$VERSION"
|
||||
mc-stop "$VERSION" --confirm
|
||||
wait_for_server_stop "$VERSION"
|
||||
fi
|
||||
|
||||
if server_running; then
|
||||
mc-stop "$VERSION" --confirm
|
||||
wait_for_server_stop "$VERSION"
|
||||
fi
|
||||
|
||||
upsert_property "online-mode" "false"
|
||||
upsert_property "enforce-secure-profile" "false"
|
||||
upsert_property "enable-rcon" "true"
|
||||
upsert_property "rcon.port" "25575"
|
||||
upsert_property "rcon.password" "test123"
|
||||
|
||||
echo "Configured $VERSION for persistent offline testing"
|
||||
35
.skills/mcc-integration-testing/scripts/get_server_port.sh
Normal file
35
.skills/mcc-integration-testing/scripts/get_server_port.sh
Normal file
|
|
@ -0,0 +1,35 @@
|
|||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
if [[ $# -ne 1 ]]; then
|
||||
echo "Usage: $0 <server-dir>" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
REPO_ROOT="$(cd "$(dirname "$0")/../../.." && pwd)"
|
||||
SERVER_DIR_NAME="$1"
|
||||
SERVERS_ROOT="${MCC_SERVERS:-$REPO_ROOT/MinecraftOfficial/downloads}"
|
||||
SERVER_DIR="$SERVERS_ROOT/$SERVER_DIR_NAME"
|
||||
PROPS_FILE="$SERVER_DIR/server.properties"
|
||||
LATEST_LOG="$SERVER_DIR/logs/latest.log"
|
||||
|
||||
if [[ -f "$PROPS_FILE" ]]; then
|
||||
PORT_LINE="$(grep -E '^server-port=' "$PROPS_FILE" | tail -n 1 || true)"
|
||||
if [[ -n "$PORT_LINE" ]]; then
|
||||
PORT="${PORT_LINE#server-port=}"
|
||||
if [[ "$PORT" =~ ^[0-9]+$ ]]; then
|
||||
printf '%s\n' "$PORT"
|
||||
exit 0
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
if [[ -f "$LATEST_LOG" ]]; then
|
||||
PORT="$(sed -n 's/.*Starting Minecraft server on .*:\([0-9][0-9]*\).*/\1/p' "$LATEST_LOG" | tail -n 1)"
|
||||
if [[ -n "$PORT" ]]; then
|
||||
printf '%s\n' "$PORT"
|
||||
exit 0
|
||||
fi
|
||||
fi
|
||||
|
||||
printf '25565\n'
|
||||
48
.skills/mcc-integration-testing/scripts/preflight_test_env.sh
Executable file
48
.skills/mcc-integration-testing/scripts/preflight_test_env.sh
Executable file
|
|
@ -0,0 +1,48 @@
|
|||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
REPO_ROOT="$(cd "$SCRIPT_DIR/../../.." && pwd)"
|
||||
# shellcheck source=tools/mcc-env.sh
|
||||
source "$REPO_ROOT/tools/mcc-env.sh"
|
||||
# shellcheck source=.skills/mcc-integration-testing/scripts/common.sh
|
||||
source "$SCRIPT_DIR/common.sh"
|
||||
|
||||
usage() {
|
||||
cat <<'EOF'
|
||||
Usage: preflight_test_env.sh [server-dir...]
|
||||
|
||||
Checks the local MCC test environment and resolves common Java path issues.
|
||||
EOF
|
||||
}
|
||||
|
||||
if [[ "${1:-}" == "-h" || "${1:-}" == "--help" ]]; then
|
||||
usage
|
||||
exit 0
|
||||
fi
|
||||
|
||||
ensure_java_in_path
|
||||
command -v tmux >/dev/null 2>&1 || { echo "tmux was not found on PATH." >&2; exit 1; }
|
||||
command -v dotnet >/dev/null 2>&1 || { echo "dotnet was not found on PATH." >&2; exit 1; }
|
||||
command -v python3 >/dev/null 2>&1 || { echo "python3 was not found on PATH." >&2; exit 1; }
|
||||
|
||||
if [[ ! -d "$MCC_SERVERS" ]]; then
|
||||
echo "Server root not found: $MCC_SERVERS" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
for server_dir in "$@"; do
|
||||
[[ -z "$server_dir" ]] && continue
|
||||
if [[ ! -d "$MCC_SERVERS/$server_dir" ]]; then
|
||||
echo "Server directory not found: $MCC_SERVERS/$server_dir" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
remove_stale_stdin_pipe "$server_dir"
|
||||
done
|
||||
|
||||
printf 'MCC_REPO=%s\n' "$MCC_REPO"
|
||||
printf 'MCC_SERVERS=%s\n' "$MCC_SERVERS"
|
||||
printf 'JAVA=%s\n' "$(command -v java)"
|
||||
printf 'TMUX=%s\n' "$(command -v tmux)"
|
||||
printf 'DOTNET=%s\n' "$(command -v dotnet)"
|
||||
|
|
@ -0,0 +1,106 @@
|
|||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
REPO_ROOT="$(cd "$SCRIPT_DIR/../../.." && pwd)"
|
||||
# shellcheck source=.skills/mcc-integration-testing/scripts/common.sh
|
||||
source "$SCRIPT_DIR/common.sh"
|
||||
|
||||
usage() {
|
||||
cat <<'EOF' >&2
|
||||
Usage:
|
||||
prepare_offline_mcc_config.sh <output-ini> <mc-version> [login]
|
||||
prepare_offline_mcc_config.sh <template-ini> <output-ini> <mc-version> [login]
|
||||
EOF
|
||||
}
|
||||
|
||||
if [[ $# -lt 2 || $# -gt 4 ]]; then
|
||||
usage
|
||||
exit 1
|
||||
fi
|
||||
|
||||
TEMPLATE_INI=""
|
||||
OUTPUT_INI=""
|
||||
MC_VERSION=""
|
||||
LOGIN_NAME=""
|
||||
|
||||
if [[ $# -ge 3 && -f "$1" ]]; then
|
||||
TEMPLATE_INI="$1"
|
||||
OUTPUT_INI="$2"
|
||||
MC_VERSION="$3"
|
||||
LOGIN_NAME="${4:-MCCBot}"
|
||||
else
|
||||
OUTPUT_INI="$1"
|
||||
MC_VERSION="$2"
|
||||
LOGIN_NAME="${3:-MCCBot}"
|
||||
fi
|
||||
|
||||
ACCOUNT_TYPE="${MCC_TEST_ACCOUNT_TYPE:-mojang}"
|
||||
PASSWORD_VALUE="${MCC_TEST_PASSWORD-}"
|
||||
|
||||
if [[ "$ACCOUNT_TYPE" != "mojang" && "$ACCOUNT_TYPE" != "microsoft" && "$ACCOUNT_TYPE" != "yggdrasil" ]]; then
|
||||
echo "Unsupported MCC_TEST_ACCOUNT_TYPE: $ACCOUNT_TYPE" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if [[ -z "${MCC_TEST_PASSWORD+x}" ]]; then
|
||||
if [[ "$ACCOUNT_TYPE" == "mojang" ]]; then
|
||||
PASSWORD_VALUE="-"
|
||||
else
|
||||
PASSWORD_VALUE=""
|
||||
fi
|
||||
fi
|
||||
|
||||
generate_template_ini() {
|
||||
local template_root
|
||||
template_root="$(mktemp -d "${TMPDIR:-/tmp}/mcc-config-template.XXXXXX")"
|
||||
|
||||
if [[ ! -f "$REPO_ROOT/MinecraftClient/bin/Release/net10.0/MinecraftClient.dll" ]]; then
|
||||
dotnet build "$REPO_ROOT/MinecraftClient.sln" -c Release -v quiet --nologo >/dev/null
|
||||
fi
|
||||
|
||||
(
|
||||
cd "$template_root"
|
||||
dotnet run --project "$REPO_ROOT/MinecraftClient" -c Release --no-build -- --help >/dev/null 2>&1
|
||||
)
|
||||
|
||||
if [[ ! -f "$template_root/MinecraftClient.ini" ]]; then
|
||||
echo "Failed to generate a temporary MCC config template." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
TEMPLATE_INI="$template_root/MinecraftClient.ini"
|
||||
}
|
||||
|
||||
if [[ -z "$TEMPLATE_INI" ]]; then
|
||||
generate_template_ini
|
||||
fi
|
||||
|
||||
mkdir -p "$(dirname "$OUTPUT_INI")"
|
||||
cp "$TEMPLATE_INI" "$OUTPUT_INI"
|
||||
|
||||
sed_in_place \
|
||||
-e "s#^Account = .*#Account = { Login = \"$LOGIN_NAME\", Password = \"$PASSWORD_VALUE\" }#" \
|
||||
-e "s#^AccountType = .*#AccountType = \"$ACCOUNT_TYPE\"#" \
|
||||
-e "s#^MinecraftVersion = \"[^\"]*\"\\(.*\\)\$#MinecraftVersion = \"$MC_VERSION\"\\1#" \
|
||||
-e 's#^TerrainAndMovements = false#TerrainAndMovements = true#' \
|
||||
-e 's#^InventoryHandling = false#InventoryHandling = true#' \
|
||||
-e 's#^EntityHandling = false#EntityHandling = true#' \
|
||||
-e 's#^AutoRespawn = false#AutoRespawn = true#' \
|
||||
"$OUTPUT_INI"
|
||||
|
||||
disable_noisy_bots_in_ini "$OUTPUT_INI"
|
||||
|
||||
grep -Fq "AccountType = \"$ACCOUNT_TYPE\"" "$OUTPUT_INI" || {
|
||||
echo "Failed to enforce account type $ACCOUNT_TYPE in $OUTPUT_INI" >&2
|
||||
exit 1
|
||||
}
|
||||
|
||||
if [[ "$ACCOUNT_TYPE" == "mojang" ]]; then
|
||||
grep -Eq '^Account = \{ Login = ".*", Password = "-" \}' "$OUTPUT_INI" || {
|
||||
echo "Failed to enforce offline account in $OUTPUT_INI" >&2
|
||||
exit 1
|
||||
}
|
||||
fi
|
||||
|
||||
printf '%s\n' "$OUTPUT_INI"
|
||||
44
.skills/mcc-integration-testing/scripts/reset_shared_test_state.sh
Executable file
44
.skills/mcc-integration-testing/scripts/reset_shared_test_state.sh
Executable file
|
|
@ -0,0 +1,44 @@
|
|||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
REPO_ROOT="$(cd "$SCRIPT_DIR/../../.." && pwd)"
|
||||
# shellcheck source=tools/mcc-env.sh
|
||||
source "$REPO_ROOT/tools/mcc-env.sh"
|
||||
# shellcheck source=.skills/mcc-integration-testing/scripts/common.sh
|
||||
source "$SCRIPT_DIR/common.sh"
|
||||
|
||||
usage() {
|
||||
cat <<'EOF'
|
||||
Usage: reset_shared_test_state.sh [--all | <server-dir>...]
|
||||
|
||||
Kills shared server tmux test sessions and removes stale server stdin pipes.
|
||||
EOF
|
||||
}
|
||||
|
||||
if [[ "${1:-}" == "-h" || "${1:-}" == "--help" ]]; then
|
||||
usage
|
||||
exit 0
|
||||
fi
|
||||
|
||||
kill_named_session() {
|
||||
local session_name="$1"
|
||||
tmux kill-session -t "$session_name" 2>/dev/null || true
|
||||
}
|
||||
|
||||
if [[ $# -eq 0 || "${1:-}" == "--all" ]]; then
|
||||
while IFS= read -r session_name; do
|
||||
[[ -z "$session_name" ]] && continue
|
||||
kill_named_session "$session_name"
|
||||
done < <(tmux list-sessions 2>/dev/null | awk -F: '/^mc-/{print $1}' || true)
|
||||
|
||||
while IFS= read -r pipe_path; do
|
||||
[[ -z "$pipe_path" ]] && continue
|
||||
rm -f "$pipe_path"
|
||||
done < <(find "$MCC_SERVERS" -maxdepth 2 -name 'stdin.pipe' 2>/dev/null || true)
|
||||
else
|
||||
for version in "$@"; do
|
||||
kill_named_session "$(server_session_name "$version")"
|
||||
remove_stale_stdin_pipe "$version"
|
||||
done
|
||||
fi
|
||||
168
.skills/mcc-integration-testing/scripts/run_achievements_matrix.sh
Executable file
168
.skills/mcc-integration-testing/scripts/run_achievements_matrix.sh
Executable file
|
|
@ -0,0 +1,168 @@
|
|||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
REPO_ROOT="$(cd "$SCRIPT_DIR/../../.." && pwd)"
|
||||
# shellcheck source=tools/mcc-env.sh
|
||||
source "$REPO_ROOT/tools/mcc-env.sh"
|
||||
|
||||
RUN_ROOT="${TMPDIR:-/tmp}/mcc-achievements/matrix"
|
||||
RUN_ID="$(date +%Y%m%d-%H%M%S)"
|
||||
MATRIX_DIR="$RUN_ROOT/$RUN_ID"
|
||||
RESULTS_TSV="$MATRIX_DIR/results.tsv"
|
||||
BUILD_LOG="$MATRIX_DIR/build.log"
|
||||
REPORT_MD="$MATRIX_DIR/report.md"
|
||||
PRECHECK_TXT="$MATRIX_DIR/preflight.txt"
|
||||
|
||||
mkdir -p "$MATRIX_DIR"
|
||||
|
||||
write_row() {
|
||||
local fields=("$@")
|
||||
|
||||
while (( ${#fields[@]} < 14 )); do
|
||||
fields+=("")
|
||||
done
|
||||
|
||||
printf '%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\n' \
|
||||
"${fields[0]}" "${fields[1]}" "${fields[2]}" "${fields[3]}" "${fields[4]}" "${fields[5]}" "${fields[6]}" \
|
||||
"${fields[7]}" "${fields[8]}" "${fields[9]}" "${fields[10]}" "${fields[11]}" "${fields[12]}" \
|
||||
"${fields[13]}" >> "$RESULTS_TSV"
|
||||
}
|
||||
|
||||
resolve_server_dir() {
|
||||
local version="$1"
|
||||
local candidate
|
||||
|
||||
for candidate in "$version" "$version-Vanilla"; do
|
||||
if [[ -d "$MCC_SERVERS/$candidate" ]]; then
|
||||
printf '%s\n' "$candidate"
|
||||
return 0
|
||||
fi
|
||||
done
|
||||
|
||||
return 1
|
||||
}
|
||||
|
||||
run_version() {
|
||||
local version="$1"
|
||||
local profile="$2"
|
||||
local family="$3"
|
||||
local server_dir="$4"
|
||||
local summary_env
|
||||
|
||||
if bash "$SCRIPT_DIR/run_achievements_test.sh" --no-build "$server_dir" "$version" "$profile"; then
|
||||
:
|
||||
fi
|
||||
|
||||
summary_env="${TMPDIR:-/tmp}/mcc-achievements/$server_dir/latest/summary.env"
|
||||
if [[ ! -f "$summary_env" ]]; then
|
||||
write_row "$version" "$server_dir" "unknown" "$family" "❌" "❌" "❌" "❌" "❌ Fail" \
|
||||
"Summary file was not produced." "" "" ""
|
||||
return
|
||||
fi
|
||||
|
||||
# shellcheck disable=SC1090
|
||||
source "$summary_env"
|
||||
|
||||
if [[ -n "${MCC_LOG:-}" && ! -f "$MCC_LOG" ]]; then
|
||||
NOTE="Harness failure: MCC log was not produced."
|
||||
VERDICT="❌ Fail"
|
||||
fi
|
||||
|
||||
if [[ -n "${COMMAND_LOG:-}" && ! -f "$COMMAND_LOG" ]]; then
|
||||
NOTE="Harness failure: command transcript was not produced."
|
||||
VERDICT="❌ Fail"
|
||||
fi
|
||||
|
||||
write_row "$VERSION" "$SERVER_DIR" "$PORT" "$FAMILY" "$INITIAL_STATUS" "$GRANT_STATUS" "$REVOKE_STATUS" \
|
||||
"$API_STATUS" "$VERDICT" "$NOTE" "$RUN_DIR" "$MCC_LOG" "$COPIED_SERVER_LOG" "$COMMAND_LOG"
|
||||
}
|
||||
|
||||
{
|
||||
printf 'MCC_SERVERS=%s\n' "$MCC_SERVERS"
|
||||
printf 'RUN_DIR=%s\n' "$MATRIX_DIR"
|
||||
printf 'DATE=%s\n' "$(date -u '+%Y-%m-%d %H:%M:%S UTC')"
|
||||
} > "$PRECHECK_TXT"
|
||||
|
||||
printf 'Version\tServerDir\tPort\tFamily\tInitial\tGrant\tRevoke\tAPI\tVerdict\tNote\tRunDir\tMccLog\tServerLog\tCommandLog\n' > "$RESULTS_TSV"
|
||||
|
||||
JAVA_OK="yes"
|
||||
TMUX_OK="yes"
|
||||
DOTNET_OK="yes"
|
||||
BUILD_OK="yes"
|
||||
|
||||
if ! command -v dotnet >/dev/null 2>&1; then
|
||||
DOTNET_OK="no"
|
||||
fi
|
||||
|
||||
if ! command -v java >/dev/null 2>&1 || ! java -version >/dev/null 2>&1; then
|
||||
JAVA_OK="no"
|
||||
fi
|
||||
|
||||
if ! command -v tmux >/dev/null 2>&1; then
|
||||
TMUX_OK="no"
|
||||
fi
|
||||
|
||||
if [[ "$DOTNET_OK" == "yes" ]]; then
|
||||
bash "$SCRIPT_DIR/preflight_test_env.sh" >/dev/null 2>&1 || true
|
||||
if ! dotnet build "$REPO_ROOT/MinecraftClient.sln" -c Release > "$BUILD_LOG" 2>&1; then
|
||||
BUILD_OK="no"
|
||||
fi
|
||||
else
|
||||
: > "$BUILD_LOG"
|
||||
fi
|
||||
|
||||
{
|
||||
printf 'MCC_SERVERS=%s\n' "$MCC_SERVERS"
|
||||
printf 'RUN_DIR=%s\n' "$MATRIX_DIR"
|
||||
printf 'DATE=%s\n' "$(date -u '+%Y-%m-%d %H:%M:%S UTC')"
|
||||
printf 'dotnet=%s\n' "$DOTNET_OK"
|
||||
printf 'java=%s\n' "$JAVA_OK"
|
||||
printf 'tmux=%s\n' "$TMUX_OK"
|
||||
printf 'build=%s\n' "$BUILD_OK"
|
||||
} > "$PRECHECK_TXT"
|
||||
|
||||
while IFS='|' read -r version profile family; do
|
||||
[[ -z "$version" ]] && continue
|
||||
|
||||
if [[ "$DOTNET_OK" != "yes" ]]; then
|
||||
write_row "$version" "" "" "$family" "❌" "❌" "❌" "❌" "❌ Fail" \
|
||||
"dotnet is not available on PATH."
|
||||
continue
|
||||
fi
|
||||
|
||||
if [[ "$BUILD_OK" != "yes" ]]; then
|
||||
write_row "$version" "" "" "$family" "❌" "❌" "❌" "❌" "❌ Fail" \
|
||||
"dotnet build failed. See $BUILD_LOG."
|
||||
continue
|
||||
fi
|
||||
|
||||
if [[ "$JAVA_OK" != "yes" || "$TMUX_OK" != "yes" ]]; then
|
||||
write_row "$version" "" "" "$family" "❌" "❌" "❌" "❌" "❌ Fail" \
|
||||
"java or tmux is not available, so live server execution was blocked."
|
||||
continue
|
||||
fi
|
||||
|
||||
if ! server_dir="$(resolve_server_dir "$version")"; then
|
||||
write_row "$version" "" "" "$family" "❌" "❌" "❌" "❌" "⚠️ Partial" \
|
||||
"Server directory for $version was not found under $MCC_SERVERS."
|
||||
continue
|
||||
fi
|
||||
|
||||
run_version "$version" "$profile" "$family" "$server_dir"
|
||||
done <<'EOF'
|
||||
1.8|legacy|Legacy 🧱
|
||||
1.11.2|legacy|Legacy 🧱
|
||||
1.12.2|modern|First advancements 🌱
|
||||
1.19.4|modern|Stable modern ✅
|
||||
1.20|modern|Telemetry edge 1 ⚠️
|
||||
1.20.2|modern|Telemetry edge 2 ⚠️
|
||||
1.20.4|modern|End of 1.20.x ⚠️
|
||||
1.20.6|modern|Post-1.20.6 🔧
|
||||
1.21.2|modern|1.21.2 family 🔧
|
||||
1.21.11|modern|showAdvancements 🆕
|
||||
26.1|modern|Latest supported 🚀
|
||||
EOF
|
||||
|
||||
bash "$SCRIPT_DIR/summarize_achievements_matrix.sh" "$MATRIX_DIR" > "$REPORT_MD"
|
||||
printf '%s\n' "$MATRIX_DIR"
|
||||
399
.skills/mcc-integration-testing/scripts/run_achievements_test.sh
Executable file
399
.skills/mcc-integration-testing/scripts/run_achievements_test.sh
Executable file
|
|
@ -0,0 +1,399 @@
|
|||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
REPO_ROOT="$(cd "$SCRIPT_DIR/../../.." && pwd)"
|
||||
# shellcheck source=tools/mcc-env.sh
|
||||
source "$REPO_ROOT/tools/mcc-env.sh"
|
||||
# shellcheck source=.skills/mcc-integration-testing/scripts/common.sh
|
||||
source "$SCRIPT_DIR/common.sh"
|
||||
|
||||
usage() {
|
||||
cat <<'EOF'
|
||||
Usage: run_achievements_test.sh [--no-build] <server-dir> <mc-version> <legacy|modern>
|
||||
|
||||
Examples:
|
||||
.skills/mcc-integration-testing/scripts/run_achievements_test.sh --no-build 1.8 1.8 legacy
|
||||
.skills/mcc-integration-testing/scripts/run_achievements_test.sh --no-build 1.21.11-Vanilla 1.21.11 modern
|
||||
EOF
|
||||
}
|
||||
|
||||
DO_BUILD=true
|
||||
|
||||
while [[ $# -gt 0 ]]; do
|
||||
case "$1" in
|
||||
--no-build) DO_BUILD=false; shift ;;
|
||||
--build) DO_BUILD=true; shift ;;
|
||||
-h|--help) usage; exit 0 ;;
|
||||
*) break ;;
|
||||
esac
|
||||
done
|
||||
|
||||
if [[ $# -ne 3 ]]; then
|
||||
usage >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
SERVER_DIR="$1"
|
||||
MC_VERSION="$2"
|
||||
PROFILE="$3"
|
||||
SESSION_NAME="achievements-${SERVER_DIR//[^a-zA-Z0-9]/_}-${PROFILE}"
|
||||
TEST_USERNAME="$(_mcc_resolve_username "$SESSION_NAME")"
|
||||
|
||||
if [[ "$PROFILE" != "legacy" && "$PROFILE" != "modern" ]]; then
|
||||
echo "Unsupported profile: $PROFILE" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
RUN_ROOT="${TMPDIR:-/tmp}/mcc-achievements"
|
||||
RUN_ID="$(date +%Y%m%d-%H%M%S)"
|
||||
RUN_DIR="$RUN_ROOT/$SERVER_DIR/$RUN_ID"
|
||||
LATEST_LINK="$RUN_ROOT/$SERVER_DIR/latest"
|
||||
MCC_LOG="$RUN_DIR/mcc.log"
|
||||
BUILD_LOG="$RUN_DIR/build.log"
|
||||
SERVER_TMUX_LOG="$RUN_DIR/server-tmux.log"
|
||||
SERVER_FILE_LOG="$RUN_DIR/server-latest.log"
|
||||
COMMAND_LOG="$RUN_DIR/commands.log"
|
||||
SUMMARY_ENV="$RUN_DIR/summary.env"
|
||||
PROBE_SCRIPT="$RUN_DIR/achievement_probe.cs"
|
||||
CFG="$RUN_DIR/MinecraftClient.$MC_VERSION.ini"
|
||||
INPUT_FILE="$(_mcc_session_input_file "$SESSION_NAME")"
|
||||
SERVER_LOG_FILE="$MCC_SERVERS/$SERVER_DIR/logs/latest.log"
|
||||
TARGET_ID="minecraft:story/root"
|
||||
TARGET_COMMAND_GRANT="advancement grant $TEST_USERNAME only minecraft:story/root"
|
||||
TARGET_COMMAND_REVOKE="advancement revoke $TEST_USERNAME only minecraft:story/root"
|
||||
TARGET_TYPE="Modern 🌱"
|
||||
PORT="unknown"
|
||||
MCC_PID=""
|
||||
|
||||
INITIAL_STATUS="❌"
|
||||
GRANT_STATUS="❌"
|
||||
REVOKE_STATUS="❌"
|
||||
API_STATUS="❌"
|
||||
VERDICT="❌ Fail"
|
||||
NOTE="Run did not complete."
|
||||
EXECUTED="yes"
|
||||
|
||||
if [[ "$PROFILE" == "legacy" ]]; then
|
||||
TARGET_ID="achievement.openInventory"
|
||||
TARGET_COMMAND_GRANT="achievement give achievement.openInventory $TEST_USERNAME"
|
||||
TARGET_COMMAND_REVOKE="achievement take achievement.openInventory $TEST_USERNAME"
|
||||
TARGET_TYPE="Legacy 🧱"
|
||||
fi
|
||||
|
||||
mkdir -p "$RUN_DIR"
|
||||
|
||||
write_summary() {
|
||||
{
|
||||
printf 'VERSION=%q\n' "$MC_VERSION"
|
||||
printf 'SERVER_DIR=%q\n' "$SERVER_DIR"
|
||||
printf 'PROFILE=%q\n' "$PROFILE"
|
||||
printf 'FAMILY=%q\n' "$TARGET_TYPE"
|
||||
printf 'PORT=%q\n' "$PORT"
|
||||
printf 'RUN_DIR=%q\n' "$RUN_DIR"
|
||||
printf 'MCC_LOG=%q\n' "$MCC_LOG"
|
||||
printf 'SERVER_LOG=%q\n' "$RUN_DIR/server-latest.log"
|
||||
printf 'SERVER_FILE_LOG=%q\n' "$SERVER_LOG_FILE"
|
||||
printf 'SERVER_TMUX_LOG=%q\n' "$SERVER_TMUX_LOG"
|
||||
printf 'COPIED_SERVER_LOG=%q\n' "$RUN_DIR/server-latest.log"
|
||||
printf 'COMMAND_LOG=%q\n' "$COMMAND_LOG"
|
||||
printf 'SUMMARY_ENV=%q\n' "$SUMMARY_ENV"
|
||||
printf 'TARGET_ID=%q\n' "$TARGET_ID"
|
||||
printf 'INITIAL_STATUS=%q\n' "$INITIAL_STATUS"
|
||||
printf 'GRANT_STATUS=%q\n' "$GRANT_STATUS"
|
||||
printf 'REVOKE_STATUS=%q\n' "$REVOKE_STATUS"
|
||||
printf 'API_STATUS=%q\n' "$API_STATUS"
|
||||
printf 'VERDICT=%q\n' "$VERDICT"
|
||||
printf 'NOTE=%q\n' "$NOTE"
|
||||
printf 'EXECUTED=%q\n' "$EXECUTED"
|
||||
} > "$SUMMARY_ENV"
|
||||
}
|
||||
|
||||
capture_server_logs() {
|
||||
mc-log "$SERVER_DIR" 400 > "$SERVER_TMUX_LOG" 2>/dev/null || true
|
||||
if [[ -f "$SERVER_LOG_FILE" ]]; then
|
||||
cp "$SERVER_LOG_FILE" "$RUN_DIR/server-latest.log" 2>/dev/null || true
|
||||
fi
|
||||
}
|
||||
|
||||
cleanup() {
|
||||
capture_server_logs
|
||||
|
||||
if [[ -n "${MCC_PID:-}" ]] && kill -0 "$MCC_PID" 2>/dev/null; then
|
||||
echo "quit" >> "$INPUT_FILE" 2>/dev/null || true
|
||||
sleep 2
|
||||
kill "$MCC_PID" 2>/dev/null || true
|
||||
wait "$MCC_PID" 2>/dev/null || true
|
||||
fi
|
||||
|
||||
mc-stop "$SERVER_DIR" --confirm >/dev/null 2>&1 || true
|
||||
wait_for_server_stop "$SERVER_DIR" 20 >/dev/null 2>&1 || true
|
||||
ln -sfn "$RUN_DIR" "$LATEST_LINK"
|
||||
write_summary
|
||||
}
|
||||
trap cleanup EXIT
|
||||
|
||||
log_step() {
|
||||
printf '[%s] %s\n' "$(date '+%H:%M:%S')" "$1" | tee -a "$COMMAND_LOG"
|
||||
}
|
||||
|
||||
fail() {
|
||||
NOTE="$1"
|
||||
VERDICT="❌ Fail"
|
||||
exit 1
|
||||
}
|
||||
|
||||
wait_for_file_pattern() {
|
||||
local file="$1"
|
||||
local pattern="$2"
|
||||
local description="$3"
|
||||
local timeout="${4:-60}"
|
||||
local elapsed=0
|
||||
|
||||
while (( elapsed < timeout )); do
|
||||
if [[ -f "$file" ]] && grep -Fq "$pattern" "$file"; then
|
||||
return 0
|
||||
fi
|
||||
sleep 1
|
||||
((elapsed += 1))
|
||||
done
|
||||
|
||||
echo "Timed out waiting for: $description" >&2
|
||||
return 1
|
||||
}
|
||||
|
||||
write_probe_script() {
|
||||
cat > "$PROBE_SCRIPT" <<EOF
|
||||
//MCCScript 1.0
|
||||
|
||||
MCC.LoadBot(new AchievementProbeBot());
|
||||
|
||||
//MCCScript Extensions
|
||||
|
||||
public class AchievementProbeBot : ChatBot
|
||||
{
|
||||
private const string TargetId = "$TARGET_ID";
|
||||
|
||||
public override void Initialize()
|
||||
{
|
||||
LogToConsole("[ACH_TEST] probe initialized");
|
||||
DumpState("initialize");
|
||||
}
|
||||
|
||||
public override void AfterGameJoined()
|
||||
{
|
||||
LogToConsole("[ACH_TEST] after join");
|
||||
DumpState("after_join");
|
||||
}
|
||||
|
||||
public override void OnAchievementUpdate(IReadOnlyList<Achievement> updated, IReadOnlyList<string> removedIds, bool reset)
|
||||
{
|
||||
LogToConsole($"[ACH_TEST] event reset={reset} updated={updated.Count} removed={removedIds.Count}");
|
||||
DumpState("event");
|
||||
}
|
||||
|
||||
private void DumpState(string origin)
|
||||
{
|
||||
Achievement[] all = GetAchievements();
|
||||
Achievement[] unlocked = GetUnlockedAchievements();
|
||||
Achievement[] locked = GetLockedAchievements();
|
||||
Achievement? target = null;
|
||||
|
||||
foreach (Achievement entry in all)
|
||||
{
|
||||
if (entry.Id == TargetId)
|
||||
{
|
||||
target = entry;
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
string titleState = "missing";
|
||||
string completionState = "missing";
|
||||
|
||||
if (target is not null)
|
||||
{
|
||||
titleState = target.Title is null ? "null" : "present";
|
||||
completionState = target.IsCompleted ? "done" : "todo";
|
||||
}
|
||||
|
||||
LogToConsole($"[ACH_TEST] snapshot origin={origin} all={all.Length} unlocked={unlocked.Length} locked={locked.Length}");
|
||||
LogToConsole($"[ACH_TEST] target_state origin={origin} id={TargetId} title={titleState} completed={completionState}");
|
||||
}
|
||||
}
|
||||
EOF
|
||||
}
|
||||
|
||||
run_server_command() {
|
||||
local cmd="$1"
|
||||
local attempt
|
||||
|
||||
log_step "SERVER> $cmd"
|
||||
for attempt in 1 2 3 4 5; do
|
||||
if mc-rcon "$cmd" >/dev/null 2>&1; then
|
||||
sleep 1
|
||||
return 0
|
||||
fi
|
||||
sleep 1
|
||||
done
|
||||
|
||||
fail "Server command failed: $cmd"
|
||||
}
|
||||
|
||||
run_mcc_command() {
|
||||
local name="$1"
|
||||
local cmd="$2"
|
||||
local delay="${3:-2}"
|
||||
local start_line=0
|
||||
local end_line=0
|
||||
|
||||
if [[ -f "$MCC_LOG" ]]; then
|
||||
start_line="$(wc -l < "$MCC_LOG")"
|
||||
fi
|
||||
|
||||
log_step "MCC> $cmd"
|
||||
echo "$cmd" >> "$INPUT_FILE"
|
||||
sleep "$delay"
|
||||
|
||||
if [[ -f "$MCC_LOG" ]]; then
|
||||
end_line="$(wc -l < "$MCC_LOG")"
|
||||
fi
|
||||
|
||||
if (( end_line > start_line )); then
|
||||
sed -n "$((start_line + 1)),$((end_line))p" "$MCC_LOG" > "$RUN_DIR/$name.mcc.log"
|
||||
else
|
||||
: > "$RUN_DIR/$name.mcc.log"
|
||||
fi
|
||||
}
|
||||
|
||||
assert_pattern() {
|
||||
local file="$1"
|
||||
local pattern="$2"
|
||||
local description="$3"
|
||||
|
||||
grep -Fq "$pattern" "$file" || fail "$description"
|
||||
}
|
||||
|
||||
if $DO_BUILD; then
|
||||
log_step "BUILD> dotnet build MinecraftClient.sln -c Release"
|
||||
mcc-build > "$BUILD_LOG" 2>&1 || fail "dotnet build failed."
|
||||
else
|
||||
: > "$BUILD_LOG"
|
||||
fi
|
||||
|
||||
bash "$SCRIPT_DIR/preflight_test_env.sh" "$SERVER_DIR" >/dev/null || fail "Test environment preflight failed."
|
||||
bash "$SCRIPT_DIR/reset_shared_test_state.sh" "$SERVER_DIR" >/dev/null || fail "Failed to reset shared test state."
|
||||
|
||||
if [[ ! -d "$MCC_SERVERS/$SERVER_DIR" ]]; then
|
||||
fail "Server directory not found: $MCC_SERVERS/$SERVER_DIR"
|
||||
fi
|
||||
|
||||
bash "$SCRIPT_DIR/prepare_offline_mcc_config.sh" "$CFG" "$MC_VERSION" "$TEST_USERNAME" >/dev/null || fail "Failed to prepare temporary MCC config."
|
||||
PORT="$(bash "$SCRIPT_DIR/get_server_port.sh" "$SERVER_DIR")"
|
||||
|
||||
"$SCRIPT_DIR/ensure_offline_server.sh" "$SERVER_DIR"
|
||||
write_probe_script
|
||||
|
||||
if [[ "$PROFILE" == "legacy" && -f "$MCC_SERVERS/$SERVER_DIR/server.properties" ]]; then
|
||||
sed_in_place 's/^use-native-transport=.*/use-native-transport=false/' "$MCC_SERVERS/$SERVER_DIR/server.properties"
|
||||
fi
|
||||
|
||||
mkdir -p "$(dirname "$INPUT_FILE")"
|
||||
: > "$INPUT_FILE"
|
||||
rm -f "$MCC_LOG"
|
||||
|
||||
log_step "Starting server $SERVER_DIR on port $PORT"
|
||||
mc-start "$SERVER_DIR" >/dev/null
|
||||
wait_for_server_ready "$SERVER_DIR" || fail "Server did not become ready."
|
||||
|
||||
log_step "Starting MCC for $MC_VERSION"
|
||||
(
|
||||
cd "$REPO_ROOT"
|
||||
MCC_FILE_INPUT=1 MCC_INPUT_FILE="$INPUT_FILE" dotnet run --project MinecraftClient -c Release --no-build -- \
|
||||
"$CFG" \
|
||||
"$TEST_USERNAME" \
|
||||
- \
|
||||
"localhost:$PORT" \
|
||||
"--accounttype=mojang" \
|
||||
"--minecraftversion=$MC_VERSION" \
|
||||
"--terrainandmovements=true" \
|
||||
"--inventoryhandling=true" \
|
||||
"--entityhandling=true" \
|
||||
"--autorespawn=true" \
|
||||
"--debugmessages=true" \
|
||||
> "$MCC_LOG" 2>&1
|
||||
) &
|
||||
MCC_PID=$!
|
||||
|
||||
wait_for_file_pattern "$MCC_LOG" "Server was successfully joined." "MCC join success" 90 || fail "MCC failed to join."
|
||||
wait_for_file_pattern "$SERVER_LOG_FILE" "$TEST_USERNAME joined the game" "server join entry" 30 || fail "Server never logged the join."
|
||||
|
||||
run_server_command "op $TEST_USERNAME"
|
||||
run_server_command "gamerule sendCommandFeedback true"
|
||||
if [[ "$PROFILE" == "modern" ]]; then
|
||||
run_server_command "gamerule logAdminCommands true"
|
||||
fi
|
||||
run_server_command "time set day"
|
||||
run_server_command "weather clear"
|
||||
|
||||
run_mcc_command "load_probe" "script $PROBE_SCRIPT" 3
|
||||
wait_for_file_pattern "$MCC_LOG" "[ACH_TEST] probe initialized" "probe startup" 30 || fail "Probe script did not initialize."
|
||||
|
||||
run_mcc_command "baseline_debug" "debug state" 2
|
||||
run_mcc_command "baseline_all" "achievement" 2
|
||||
run_mcc_command "baseline_locked" "achievement locked" 2
|
||||
run_mcc_command "baseline_unlocked" "achievement unlocked" 2
|
||||
|
||||
run_server_command "$TARGET_COMMAND_GRANT"
|
||||
sleep 3
|
||||
run_mcc_command "after_grant_all" "achievement" 2
|
||||
run_mcc_command "after_grant_unlocked" "achievement unlocked" 2
|
||||
|
||||
run_server_command "$TARGET_COMMAND_REVOKE"
|
||||
sleep 3
|
||||
run_mcc_command "after_revoke_all" "achievement" 2
|
||||
run_mcc_command "after_revoke_locked" "achievement locked" 2
|
||||
|
||||
assert_pattern "$MCC_LOG" "Achievements/Advancements:" "Achievement command header never appeared."
|
||||
|
||||
if ! grep -Fq "No achievements/advancements received yet." "$RUN_DIR/baseline_all.mcc.log"; then
|
||||
INITIAL_STATUS="✅"
|
||||
fi
|
||||
|
||||
if grep -Fq "$TARGET_ID" "$RUN_DIR/after_grant_unlocked.mcc.log" && grep -Fq "[DONE]" "$RUN_DIR/after_grant_unlocked.mcc.log"; then
|
||||
GRANT_STATUS="✅"
|
||||
fi
|
||||
|
||||
if [[ "$PROFILE" == "legacy" ]]; then
|
||||
if grep -Fq "$TARGET_ID" "$RUN_DIR/after_revoke_locked.mcc.log" && grep -Fq "[TODO]" "$RUN_DIR/after_revoke_locked.mcc.log"; then
|
||||
REVOKE_STATUS="✅"
|
||||
fi
|
||||
else
|
||||
if grep -Fq "$TARGET_ID" "$RUN_DIR/after_revoke_locked.mcc.log" && grep -Fq "[TODO]" "$RUN_DIR/after_revoke_locked.mcc.log"; then
|
||||
REVOKE_STATUS="✅"
|
||||
elif [[ "$GRANT_STATUS" == "✅" ]] && ! grep -Fq "$TARGET_ID" "$RUN_DIR/after_revoke_all.mcc.log"; then
|
||||
REVOKE_STATUS="✅"
|
||||
fi
|
||||
fi
|
||||
|
||||
if grep -Fq "[ACH_TEST] event" "$MCC_LOG" && grep -Fq "target_state origin=event id=$TARGET_ID title=" "$MCC_LOG"; then
|
||||
API_STATUS="✅"
|
||||
fi
|
||||
|
||||
case "$INITIAL_STATUS|$GRANT_STATUS|$REVOKE_STATUS|$API_STATUS" in
|
||||
"✅|✅|✅|✅")
|
||||
VERDICT="✅ Pass"
|
||||
NOTE="All planned achievement checks passed."
|
||||
;;
|
||||
*"✅"*)
|
||||
VERDICT="⚠️ Partial"
|
||||
NOTE="At least one achievement phase passed, but the matrix did not fully clear."
|
||||
;;
|
||||
*)
|
||||
VERDICT="❌ Fail"
|
||||
NOTE="Achievement checks did not produce the expected evidence."
|
||||
;;
|
||||
esac
|
||||
|
||||
run_mcc_command "quit" "quit" 2
|
||||
NOTE="$NOTE Artifacts saved in $RUN_DIR."
|
||||
336
.skills/mcc-integration-testing/scripts/run_full_spectrum_test.sh
Executable file
336
.skills/mcc-integration-testing/scripts/run_full_spectrum_test.sh
Executable file
|
|
@ -0,0 +1,336 @@
|
|||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
REPO_ROOT="$(cd "$SCRIPT_DIR/../../.." && pwd)"
|
||||
# shellcheck source=tools/mcc-env.sh
|
||||
source "$REPO_ROOT/tools/mcc-env.sh"
|
||||
# shellcheck source=.skills/mcc-integration-testing/scripts/common.sh
|
||||
source "$SCRIPT_DIR/common.sh"
|
||||
|
||||
VERSION="${1:-1.21.11-Vanilla}"
|
||||
MC_VERSION="${VERSION%-Vanilla}"
|
||||
if [[ "$MC_VERSION" == "$VERSION" ]]; then
|
||||
MC_VERSION="$VERSION"
|
||||
fi
|
||||
RUN_ROOT="${TMPDIR:-/tmp}/mcc-integration-testing"
|
||||
RUN_ID="$(date +%Y%m%d-%H%M%S)"
|
||||
RUN_DIR="$RUN_ROOT/$RUN_ID"
|
||||
SERVER_LOG_FILE="$MCC_SERVERS/$VERSION/logs/latest.log"
|
||||
SESSION_NAME="full-spectrum-${MC_VERSION//[^a-zA-Z0-9]/_}"
|
||||
TEST_USERNAME="$(_mcc_resolve_username "$SESSION_NAME")"
|
||||
MCC_LOG="$(_mcc_session_log_file "$SESSION_NAME")"
|
||||
PID_FILE="$(_mcc_session_pid_file "$SESSION_NAME")"
|
||||
MCC_TMUX_SESSION="$(_mcc_tmux_session_name "$SESSION_NAME")"
|
||||
BUILD_LOG="$RUN_DIR/build.log"
|
||||
SERVER_TMUX_LOG="$RUN_DIR/server-tmux.log"
|
||||
SERVER_FILE_LOG="$RUN_DIR/server-latest.log"
|
||||
INPUT_FILE="$(_mcc_session_input_file "$SESSION_NAME")"
|
||||
CFG="$RUN_DIR/MinecraftClient.$MC_VERSION.ini"
|
||||
|
||||
mkdir -p "$RUN_DIR"
|
||||
|
||||
cleanup() {
|
||||
mcc-cmd --session "$SESSION_NAME" "quit" >/dev/null 2>&1 || true
|
||||
sleep 2
|
||||
mcc-kill --session "$SESSION_NAME" >/dev/null 2>&1 || true
|
||||
|
||||
mc-stop "$VERSION" --confirm >/dev/null 2>&1 || true
|
||||
wait_for_server_stop "$VERSION" 20 >/dev/null 2>&1 || true
|
||||
}
|
||||
trap cleanup EXIT
|
||||
|
||||
prepare_config() {
|
||||
bash "$SCRIPT_DIR/prepare_offline_mcc_config.sh" "$CFG" "$MC_VERSION" "$TEST_USERNAME" >/dev/null
|
||||
}
|
||||
|
||||
wait_for_server_log_pattern() {
|
||||
local pattern="$1"
|
||||
local description="$2"
|
||||
local timeout="${3:-60}"
|
||||
local elapsed=0
|
||||
|
||||
while (( elapsed < timeout )); do
|
||||
if [[ -f "$SERVER_LOG_FILE" ]] && grep -Fq "$pattern" "$SERVER_LOG_FILE"; then
|
||||
return 0
|
||||
fi
|
||||
sleep 1
|
||||
((elapsed += 1))
|
||||
done
|
||||
|
||||
echo "Timed out waiting for server log: $description" >&2
|
||||
return 1
|
||||
}
|
||||
|
||||
capture_server_logs() {
|
||||
mc-log "$VERSION" 400 > "$SERVER_TMUX_LOG" 2>/dev/null || true
|
||||
if [[ -f "$SERVER_LOG_FILE" ]]; then
|
||||
cp "$SERVER_LOG_FILE" "$SERVER_FILE_LOG"
|
||||
fi
|
||||
}
|
||||
|
||||
wait_for_file_pattern() {
|
||||
local file="$1"
|
||||
local pattern="$2"
|
||||
local description="$3"
|
||||
local timeout="${4:-60}"
|
||||
local elapsed=0
|
||||
|
||||
while (( elapsed < timeout )); do
|
||||
if [[ -f "$file" ]] && grep -Fq "$pattern" "$file"; then
|
||||
return 0
|
||||
fi
|
||||
sleep 1
|
||||
((elapsed += 1))
|
||||
done
|
||||
|
||||
echo "Timed out waiting for: $description" >&2
|
||||
return 1
|
||||
}
|
||||
|
||||
fail() {
|
||||
capture_server_logs
|
||||
echo "FAIL: $1" >&2
|
||||
echo "Run directory: $RUN_DIR" >&2
|
||||
exit 1
|
||||
}
|
||||
|
||||
assert_contains() {
|
||||
local file="$1"
|
||||
local pattern="$2"
|
||||
local description="$3"
|
||||
|
||||
grep -Fq "$pattern" "$file" || fail "$description"
|
||||
}
|
||||
|
||||
assert_not_contains() {
|
||||
local file="$1"
|
||||
local pattern="$2"
|
||||
local description="$3"
|
||||
|
||||
if grep -Fq "$pattern" "$file"; then
|
||||
fail "$description"
|
||||
fi
|
||||
}
|
||||
|
||||
run_server_command() {
|
||||
local cmd="$1"
|
||||
local attempt
|
||||
echo "SERVER> $cmd"
|
||||
for attempt in 1 2 3 4 5; do
|
||||
if mc-rcon "$cmd" >/dev/null 2>&1; then
|
||||
return 0
|
||||
fi
|
||||
sleep 1
|
||||
done
|
||||
fail "Server command failed: $cmd"
|
||||
}
|
||||
|
||||
run_mcc_command() {
|
||||
local cmd="$1"
|
||||
echo "MCC> $cmd"
|
||||
mcc-cmd --session "$SESSION_NAME" "$cmd"
|
||||
sleep 2
|
||||
}
|
||||
|
||||
start_mcc_session() {
|
||||
local -a mcc_args=("$CFG" "$TEST_USERNAME" "-" "localhost:$SERVER_PORT")
|
||||
local mcc_args_cmd
|
||||
mcc_args_cmd="$(printf '%q ' "${mcc_args[@]}")"
|
||||
|
||||
tmux kill-session -t "$MCC_TMUX_SESSION" 2>/dev/null || true
|
||||
rm -f "$PID_FILE"
|
||||
tmux new-session -d -s "$MCC_TMUX_SESSION" -x 160 -y 50 \
|
||||
"cd '$REPO_ROOT' && printf '%s\n' \"\$\$\" > '$PID_FILE' && exec env MCC_FILE_INPUT=1 MCC_INPUT_FILE='$INPUT_FILE' dotnet run --project MinecraftClient -c Release --no-build -- $mcc_args_cmd > '$MCC_LOG' 2>&1"
|
||||
|
||||
for _ in $(seq 1 25); do
|
||||
if [[ -s "$PID_FILE" ]]; then
|
||||
return 0
|
||||
fi
|
||||
sleep 0.2
|
||||
done
|
||||
|
||||
fail "Failed to capture MCC PID for session $SESSION_NAME"
|
||||
}
|
||||
|
||||
bash "$SCRIPT_DIR/preflight_test_env.sh" "$VERSION" >/dev/null
|
||||
bash "$SCRIPT_DIR/reset_shared_test_state.sh" "$VERSION" >/dev/null
|
||||
"$SCRIPT_DIR/ensure_offline_server.sh" "$VERSION"
|
||||
mcc-reset-session --session "$SESSION_NAME" >/dev/null
|
||||
echo "Building MCC..."
|
||||
mcc-build > "$BUILD_LOG" 2>&1 || fail "mcc-build failed"
|
||||
prepare_config
|
||||
SERVER_PORT="$(bash "$SCRIPT_DIR/get_server_port.sh" "$VERSION")"
|
||||
if [[ -z "$SERVER_PORT" ]]; then
|
||||
fail "Failed to resolve server port"
|
||||
fi
|
||||
|
||||
mkdir -p "$(dirname "$INPUT_FILE")" "$(dirname "$MCC_LOG")"
|
||||
: > "$INPUT_FILE"
|
||||
rm -f "$MCC_LOG"
|
||||
|
||||
echo "Starting server..."
|
||||
mc-start "$VERSION" >/dev/null
|
||||
wait_for_server_ready "$VERSION" || fail "Server did not become ready"
|
||||
|
||||
echo "Starting MCC..."
|
||||
start_mcc_session
|
||||
|
||||
wait_for_file_pattern "$MCC_LOG" "Server was successfully joined." "MCC join success" 90 || fail "MCC failed to join"
|
||||
wait_for_server_log_pattern "$TEST_USERNAME joined the game" "server join entry" 30 || fail "Server never logged the join"
|
||||
|
||||
run_server_command "op $TEST_USERNAME"
|
||||
run_server_command "gamerule sendCommandFeedback true"
|
||||
run_server_command "gamerule logAdminCommands true"
|
||||
run_server_command "time set day"
|
||||
run_server_command "weather clear"
|
||||
sleep 2
|
||||
|
||||
# ── Phase 1: Basic status and info commands ──
|
||||
run_mcc_command "health"
|
||||
run_mcc_command "list"
|
||||
run_mcc_command "inventory player list"
|
||||
run_mcc_command "/gamemode creative"
|
||||
run_mcc_command "inventory creativegive 36 Diamond 16"
|
||||
run_mcc_command "inventory player list"
|
||||
run_mcc_command "entity"
|
||||
run_mcc_command "/time query daytime"
|
||||
run_mcc_command "smoke_test_from_mcc_full_spectrum"
|
||||
|
||||
# ── Phase 2: Movement and look commands ──
|
||||
run_mcc_command "/tp $TEST_USERNAME 0 -60 0"
|
||||
sleep 3
|
||||
run_mcc_command "look up"
|
||||
sleep 1
|
||||
run_mcc_command "look down"
|
||||
sleep 1
|
||||
run_mcc_command "look east"
|
||||
sleep 1
|
||||
|
||||
# ── Phase 3: Advanced inventory operations ──
|
||||
run_mcc_command "inventory creativegive 37 IronSword 1"
|
||||
run_mcc_command "inventory creativegive 38 GoldenApple 8"
|
||||
run_mcc_command "inventory player list"
|
||||
run_mcc_command "inventory creativeclear 38"
|
||||
run_mcc_command "inventory player list"
|
||||
|
||||
# ── Phase 4: Block placement and interaction ──
|
||||
run_server_command "execute as $TEST_USERNAME at @s run fill ~1 ~ ~1 ~3 ~2 ~3 minecraft:stone"
|
||||
sleep 2
|
||||
run_server_command "execute as $TEST_USERNAME at @s run setblock ~5 ~ ~5 minecraft:chest"
|
||||
sleep 1
|
||||
run_server_command "execute as $TEST_USERNAME at @s run setblock ~5 ~1 ~5 minecraft:furnace"
|
||||
sleep 1
|
||||
run_server_command "execute as $TEST_USERNAME at @s run setblock ~6 ~ ~5 minecraft:crafting_table"
|
||||
sleep 1
|
||||
|
||||
# ── Phase 5: Entity spawning (expanded coverage) ──
|
||||
run_server_command "execute as $TEST_USERNAME at @s run summon minecraft:cow ~2 ~ ~"
|
||||
run_server_command "execute as $TEST_USERNAME at @s run summon minecraft:zombie ~4 ~ ~"
|
||||
run_server_command "execute as $TEST_USERNAME at @s run summon minecraft:creeper ~6 ~ ~"
|
||||
run_server_command "execute as $TEST_USERNAME at @s run summon minecraft:skeleton ~8 ~ ~"
|
||||
run_server_command "execute as $TEST_USERNAME at @s run summon minecraft:villager ~-2 ~ ~"
|
||||
run_server_command "execute as $TEST_USERNAME at @s run summon minecraft:allay ~-4 ~ ~"
|
||||
run_server_command "execute as $TEST_USERNAME at @s run summon minecraft:armor_stand ~ ~ ~2"
|
||||
run_server_command "execute as $TEST_USERNAME at @s run summon minecraft:item_display ~-6 ~ ~ {item:{id:\"minecraft:diamond\",count:1}}"
|
||||
run_server_command "execute as $TEST_USERNAME at @s run summon minecraft:spider ~10 ~ ~"
|
||||
run_server_command "execute as $TEST_USERNAME at @s run summon minecraft:pig ~-8 ~ ~"
|
||||
|
||||
sleep 2
|
||||
run_mcc_command "entity"
|
||||
|
||||
# ── Phase 6: Effects and environment ──
|
||||
run_server_command "effect give $TEST_USERNAME minecraft:speed 30 1"
|
||||
sleep 2
|
||||
run_mcc_command "health"
|
||||
run_server_command "effect give $TEST_USERNAME minecraft:regeneration 10 1"
|
||||
sleep 2
|
||||
run_mcc_command "health"
|
||||
|
||||
# ── Phase 7: Gamemode cycling ──
|
||||
run_mcc_command "/gamemode survival"
|
||||
sleep 2
|
||||
run_mcc_command "health"
|
||||
run_mcc_command "/gamemode creative"
|
||||
sleep 2
|
||||
|
||||
# ── Phase 8: Dimension change (nether) ──
|
||||
run_server_command "execute in minecraft:the_nether run tp $TEST_USERNAME 0 64 0"
|
||||
sleep 4
|
||||
run_mcc_command "health"
|
||||
run_server_command "execute in minecraft:overworld run tp $TEST_USERNAME 0 -60 0"
|
||||
sleep 4
|
||||
|
||||
# ── Phase 9: Server chat and whisper ──
|
||||
run_server_command "say Hello from the server console"
|
||||
sleep 2
|
||||
run_server_command "msg $TEST_USERNAME This is a private whisper"
|
||||
sleep 2
|
||||
run_mcc_command "integration_test_chat_response"
|
||||
|
||||
# ── Phase 10: Particles, sounds, and explosions ──
|
||||
run_server_command "execute as $TEST_USERNAME at @s run particle minecraft:happy_villager ~ ~1 ~ 0.5 0.5 0.5 0 12 force"
|
||||
run_server_command "execute as $TEST_USERNAME at @s run particle minecraft:end_rod ~ ~1 ~ 0.5 0.5 0.5 0.01 20 force"
|
||||
run_server_command "execute as $TEST_USERNAME at @s run particle minecraft:explosion ~ ~1 ~ 0 0 0 0 1 force"
|
||||
run_server_command "execute as $TEST_USERNAME at @s run particle minecraft:totem_of_undying ~ ~1 ~ 0.5 0.5 0.5 0.1 20 force"
|
||||
run_server_command "execute as $TEST_USERNAME at @s run particle minecraft:flame ~ ~1 ~ 0.2 0.2 0.2 0.02 30 force"
|
||||
run_server_command "execute as $TEST_USERNAME at @s run particle minecraft:heart ~ ~2 ~ 0.3 0.3 0.3 0 5 force"
|
||||
|
||||
run_server_command "execute as $TEST_USERNAME at @s run playsound minecraft:entity.lightning_bolt.thunder master $TEST_USERNAME ~ ~ ~ 1 1 0"
|
||||
run_server_command "execute as $TEST_USERNAME at @s run playsound minecraft:block.note_block.bell master $TEST_USERNAME ~ ~ ~ 1 1 0"
|
||||
run_server_command "execute as $TEST_USERNAME at @s run playsound minecraft:entity.experience_orb.pickup master $TEST_USERNAME ~ ~ ~ 1 1 0"
|
||||
|
||||
run_server_command "execute as $TEST_USERNAME at @s run summon minecraft:tnt ~3 ~ ~"
|
||||
sleep 2
|
||||
run_server_command "execute as $TEST_USERNAME at @s run summon minecraft:tnt ~6 ~ ~"
|
||||
|
||||
# ── Phase 11: Kill and respawn cycle ──
|
||||
run_mcc_command "/gamemode survival"
|
||||
sleep 2
|
||||
run_server_command "kill $TEST_USERNAME"
|
||||
sleep 4
|
||||
run_mcc_command "respawn"
|
||||
sleep 4
|
||||
run_mcc_command "health"
|
||||
run_mcc_command "/gamemode creative"
|
||||
sleep 2
|
||||
|
||||
sleep 6
|
||||
capture_server_logs
|
||||
|
||||
# ── Assertions: MCC log ──
|
||||
assert_contains "$MCC_LOG" "Server was successfully joined." "MCC never joined the server"
|
||||
assert_contains "$MCC_LOG" "[FileInput] > inventory player list" "Inventory command was not executed"
|
||||
assert_contains "$MCC_LOG" "[FileInput] > entity" "Entity command was not executed"
|
||||
assert_contains "$MCC_LOG" "[FileInput] > /gamemode creative" "Creative mode command was not executed from MCC"
|
||||
assert_contains "$MCC_LOG" "Requested Diamond x16 in slot #36" "Creative inventory give did not succeed"
|
||||
assert_contains "$MCC_LOG" "smoke_test_from_mcc_full_spectrum" "Client-originated chat was not observed"
|
||||
assert_contains "$MCC_LOG" "[FileInput] > look up" "Look command was not executed"
|
||||
assert_contains "$MCC_LOG" "[FileInput] > /gamemode survival" "Survival mode switch was not executed"
|
||||
assert_contains "$MCC_LOG" "[FileInput] > respawn" "Respawn command was not executed"
|
||||
assert_contains "$MCC_LOG" "[FileInput] > health" "Health command was not executed"
|
||||
assert_contains "$MCC_LOG" "integration_test_chat_response" "Chat response test message was not observed"
|
||||
assert_not_contains "$MCC_LOG" "Please enable InventoryHandling" "Inventory handling is still disabled"
|
||||
assert_not_contains "$MCC_LOG" "Please enable EntityHandling" "Entity handling is still disabled"
|
||||
assert_not_contains "$MCC_LOG" "You must be in Creative gamemode" "Creative mode was not active when creativegive ran"
|
||||
assert_not_contains "$MCC_LOG" "Failed to load settings" "MCC failed to reload its config"
|
||||
assert_not_contains "$MCC_LOG" "NullReferenceException" "A NullReferenceException occurred during the test"
|
||||
|
||||
# ── Assertions: Server log ──
|
||||
assert_contains "$SERVER_FILE_LOG" "$TEST_USERNAME joined the game" "Server never saw $TEST_USERNAME join"
|
||||
assert_contains "$SERVER_FILE_LOG" "smoke_test_from_mcc_full_spectrum" "Server never received the client chat message"
|
||||
assert_contains "$SERVER_FILE_LOG" "Displaying particle minecraft:happy_villager" "Particle events were not recorded on the server"
|
||||
assert_contains "$SERVER_FILE_LOG" "Played sound minecraft:block.note_block.bell to $TEST_USERNAME" "Sound events were not recorded on the server"
|
||||
assert_contains "$SERVER_FILE_LOG" "Summoned new Primed TNT" "TNT summon did not occur on the server"
|
||||
assert_contains "$SERVER_FILE_LOG" "integration_test_chat_response" "Server never received the chat response test message"
|
||||
assert_contains "$SERVER_FILE_LOG" "Hello from the server console" "Server say command was not logged"
|
||||
assert_contains "$SERVER_FILE_LOG" "Killed $TEST_USERNAME" "Server kill command did not execute"
|
||||
assert_not_contains "$SERVER_FILE_LOG" "Sending unknown packet 'clientbound/minecraft:disconnect'" "Server hit the disconnect packet regression during the test"
|
||||
|
||||
cat <<EOF
|
||||
PASS
|
||||
Run directory: $RUN_DIR
|
||||
MCC log: $MCC_LOG
|
||||
Server log: $SERVER_FILE_LOG
|
||||
Build log: $BUILD_LOG
|
||||
EOF
|
||||
199
.skills/mcc-integration-testing/scripts/run_parallel_session_smoke_test.sh
Executable file
199
.skills/mcc-integration-testing/scripts/run_parallel_session_smoke_test.sh
Executable file
|
|
@ -0,0 +1,199 @@
|
|||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
REPO_ROOT="$(cd "$SCRIPT_DIR/../../.." && pwd)"
|
||||
# shellcheck source=tools/mcc-env.sh
|
||||
source "$REPO_ROOT/tools/mcc-env.sh"
|
||||
# shellcheck source=.skills/mcc-integration-testing/scripts/common.sh
|
||||
source "$SCRIPT_DIR/common.sh"
|
||||
|
||||
VERSION="${1:-1.21.11-Vanilla}"
|
||||
MC_VERSION="${VERSION%-Vanilla}"
|
||||
if [[ "$MC_VERSION" == "$VERSION" ]]; then
|
||||
MC_VERSION="$VERSION"
|
||||
fi
|
||||
|
||||
SESSION_A="parallel-smoke-a-${MC_VERSION//[^a-zA-Z0-9]/_}"
|
||||
SESSION_B="parallel-smoke-b-${MC_VERSION//[^a-zA-Z0-9]/_}"
|
||||
USERNAME_A="SmokeA"
|
||||
USERNAME_B="SmokeB"
|
||||
|
||||
RUN_ROOT="${TMPDIR:-/tmp}/mcc-integration-testing"
|
||||
RUN_ID="$(date +%Y%m%d-%H%M%S)"
|
||||
RUN_DIR="$RUN_ROOT/parallel-smoke-$RUN_ID"
|
||||
BUILD_LOG="$RUN_DIR/build.log"
|
||||
SERVER_TMUX_LOG="$RUN_DIR/server-tmux.log"
|
||||
SERVER_FILE_LOG="$RUN_DIR/server-latest.log"
|
||||
SERVER_LOG_FILE="$MCC_SERVERS/$VERSION/logs/latest.log"
|
||||
|
||||
LOG_A="$(_mcc_session_log_file "$SESSION_A")"
|
||||
LOG_B="$(_mcc_session_log_file "$SESSION_B")"
|
||||
INPUT_A="$(_mcc_session_input_file "$SESSION_A")"
|
||||
INPUT_B="$(_mcc_session_input_file "$SESSION_B")"
|
||||
PID_A_FILE="$(_mcc_session_pid_file "$SESSION_A")"
|
||||
PID_B_FILE="$(_mcc_session_pid_file "$SESSION_B")"
|
||||
|
||||
mkdir -p "$RUN_DIR"
|
||||
|
||||
wait_for_file_pattern() {
|
||||
local file="$1"
|
||||
local pattern="$2"
|
||||
local description="$3"
|
||||
local timeout="${4:-60}"
|
||||
local elapsed=0
|
||||
|
||||
while (( elapsed < timeout )); do
|
||||
if [[ -f "$file" ]] && grep -Fq "$pattern" "$file"; then
|
||||
return 0
|
||||
fi
|
||||
sleep 1
|
||||
((elapsed += 1))
|
||||
done
|
||||
|
||||
echo "Timed out waiting for: $description" >&2
|
||||
return 1
|
||||
}
|
||||
|
||||
wait_for_server_log_pattern() {
|
||||
local pattern="$1"
|
||||
local description="$2"
|
||||
local timeout="${3:-60}"
|
||||
local elapsed=0
|
||||
|
||||
while (( elapsed < timeout )); do
|
||||
if [[ -f "$SERVER_LOG_FILE" ]] && grep -Fq "$pattern" "$SERVER_LOG_FILE"; then
|
||||
return 0
|
||||
fi
|
||||
sleep 1
|
||||
((elapsed += 1))
|
||||
done
|
||||
|
||||
echo "Timed out waiting for server log: $description" >&2
|
||||
return 1
|
||||
}
|
||||
|
||||
capture_server_logs() {
|
||||
mc-log "$VERSION" 400 > "$SERVER_TMUX_LOG" 2>/dev/null || true
|
||||
if [[ -f "$SERVER_LOG_FILE" ]]; then
|
||||
cp "$SERVER_LOG_FILE" "$SERVER_FILE_LOG"
|
||||
fi
|
||||
}
|
||||
|
||||
fail() {
|
||||
capture_server_logs
|
||||
echo "FAIL: $1" >&2
|
||||
echo "Run directory: $RUN_DIR" >&2
|
||||
exit 1
|
||||
}
|
||||
|
||||
cleanup() {
|
||||
mcc-cmd --session "$SESSION_A" "quit" >/dev/null 2>&1 || true
|
||||
mcc-cmd --session "$SESSION_B" "quit" >/dev/null 2>&1 || true
|
||||
sleep 1
|
||||
mcc-kill --session "$SESSION_A" >/dev/null 2>&1 || true
|
||||
mcc-kill --session "$SESSION_B" >/dev/null 2>&1 || true
|
||||
mc-stop "$VERSION" --confirm >/dev/null 2>&1 || true
|
||||
wait_for_server_stop "$VERSION" 20 >/dev/null 2>&1 || true
|
||||
}
|
||||
trap cleanup EXIT
|
||||
|
||||
assert_session_alive() {
|
||||
local session="$1"
|
||||
local pid_file="$2"
|
||||
if [[ -s "$pid_file" ]]; then
|
||||
local pid
|
||||
pid="$(tr -cd '0-9' < "$pid_file")"
|
||||
if [[ -n "$pid" ]] && kill -0 "$pid" 2>/dev/null; then
|
||||
return 0
|
||||
fi
|
||||
fi
|
||||
|
||||
tmux has-session -t "$(_mcc_tmux_session_name "$session")" 2>/dev/null
|
||||
}
|
||||
|
||||
assert_server_alive() {
|
||||
server_running "$VERSION" || fail "Shared server session is not alive"
|
||||
}
|
||||
|
||||
start_file_input_session() {
|
||||
local session="$1"
|
||||
local username="$2"
|
||||
local port="$3"
|
||||
bash "$REPO_ROOT/tools/mcc-debug.sh" \
|
||||
--version "$VERSION" \
|
||||
--port "$port" \
|
||||
--file-input \
|
||||
--no-build \
|
||||
--session "$session" \
|
||||
--username "$username" >/dev/null
|
||||
}
|
||||
|
||||
bash "$SCRIPT_DIR/preflight_test_env.sh" "$VERSION" >/dev/null
|
||||
bash "$SCRIPT_DIR/reset_shared_test_state.sh" "$VERSION" >/dev/null
|
||||
"$SCRIPT_DIR/ensure_offline_server.sh" "$VERSION" >/dev/null
|
||||
mcc-reset-session --session "$SESSION_A" >/dev/null
|
||||
mcc-reset-session --session "$SESSION_B" >/dev/null
|
||||
|
||||
echo "Building MCC..."
|
||||
mcc-build > "$BUILD_LOG" 2>&1 || fail "mcc-build failed"
|
||||
|
||||
echo "Starting shared server..."
|
||||
mc-start "$VERSION" >/dev/null
|
||||
wait_for_server_ready "$VERSION" || fail "Server did not become ready"
|
||||
SERVER_PORT="$(bash "$SCRIPT_DIR/get_server_port.sh" "$VERSION")"
|
||||
if [[ -z "$SERVER_PORT" ]]; then
|
||||
fail "Failed to resolve server port"
|
||||
fi
|
||||
|
||||
echo "Starting MCC session A..."
|
||||
start_file_input_session "$SESSION_A" "$USERNAME_A" "$SERVER_PORT"
|
||||
echo "Starting MCC session B..."
|
||||
start_file_input_session "$SESSION_B" "$USERNAME_B" "$SERVER_PORT"
|
||||
|
||||
wait_for_file_pattern "$LOG_A" "Server was successfully joined." "session A join success" 90 || fail "Session A failed to join"
|
||||
wait_for_file_pattern "$LOG_B" "Server was successfully joined." "session B join success" 90 || fail "Session B failed to join"
|
||||
wait_for_server_log_pattern "$USERNAME_A joined the game" "server join for session A" 30 || fail "Server never logged $USERNAME_A join"
|
||||
wait_for_server_log_pattern "$USERNAME_B joined the game" "server join for session B" 30 || fail "Server never logged $USERNAME_B join"
|
||||
|
||||
echo "Sending debug state to both sessions..."
|
||||
mcc-cmd --session "$SESSION_A" "debug state"
|
||||
mcc-cmd --session "$SESSION_B" "debug state"
|
||||
sleep 2
|
||||
wait_for_file_pattern "$LOG_A" "[FileInput] > debug state" "session A debug state command" 20 || fail "Session A did not consume debug state"
|
||||
wait_for_file_pattern "$LOG_B" "[FileInput] > debug state" "session B debug state command" 20 || fail "Session B did not consume debug state"
|
||||
|
||||
echo "Killing session A..."
|
||||
mcc-kill --session "$SESSION_A" >/dev/null 2>&1 || true
|
||||
sleep 2
|
||||
|
||||
assert_session_alive "$SESSION_B" "$PID_B_FILE" || fail "Session B is not alive after killing session A"
|
||||
assert_server_alive
|
||||
|
||||
echo "Verifying session B still responds..."
|
||||
mcc-cmd --session "$SESSION_B" "health"
|
||||
wait_for_file_pattern "$LOG_B" "[FileInput] > health" "session B health command after session A kill" 20 || fail "Session B stopped responding after session A kill"
|
||||
|
||||
if [[ -s "$PID_A_FILE" ]]; then
|
||||
pid_a="$(tr -cd '0-9' < "$PID_A_FILE")"
|
||||
if [[ -n "$pid_a" ]] && kill -0 "$pid_a" 2>/dev/null; then
|
||||
fail "Session A is still alive after mcc-kill"
|
||||
fi
|
||||
fi
|
||||
|
||||
capture_server_logs
|
||||
|
||||
cat <<EOF
|
||||
PASS
|
||||
Run directory: $RUN_DIR
|
||||
Server version: $VERSION
|
||||
Server port: $SERVER_PORT
|
||||
Session A: $SESSION_A ($USERNAME_A)
|
||||
Session A input: $INPUT_A
|
||||
Session A log: $LOG_A
|
||||
Session B: $SESSION_B ($USERNAME_B)
|
||||
Session B input: $INPUT_B
|
||||
Session B log: $LOG_B
|
||||
Server log: $SERVER_FILE_LOG
|
||||
Build log: $BUILD_LOG
|
||||
EOF
|
||||
57
.skills/mcc-integration-testing/scripts/summarize_achievements_matrix.sh
Executable file
57
.skills/mcc-integration-testing/scripts/summarize_achievements_matrix.sh
Executable file
|
|
@ -0,0 +1,57 @@
|
|||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
if [[ $# -ne 1 ]]; then
|
||||
echo "Usage: summarize_achievements_matrix.sh <matrix-run-dir>" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
MATRIX_DIR="$1"
|
||||
RESULTS_TSV="$MATRIX_DIR/results.tsv"
|
||||
PRECHECK_TXT="$MATRIX_DIR/preflight.txt"
|
||||
BUILD_LOG="$MATRIX_DIR/build.log"
|
||||
|
||||
if [[ ! -f "$RESULTS_TSV" ]]; then
|
||||
echo "Missing results file: $RESULTS_TSV" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "# Achievements Matrix Report"
|
||||
echo
|
||||
echo "## Executed"
|
||||
echo
|
||||
if [[ -f "$PRECHECK_TXT" ]]; then
|
||||
echo '```text'
|
||||
cat "$PRECHECK_TXT"
|
||||
echo '```'
|
||||
fi
|
||||
echo
|
||||
echo "- Matrix artifacts: \`$MATRIX_DIR\`"
|
||||
echo "- Results TSV: \`$RESULTS_TSV\`"
|
||||
echo "- Build log: \`$BUILD_LOG\`"
|
||||
echo "- Execution mode: sequential"
|
||||
echo "- Auth mode: offline"
|
||||
echo
|
||||
echo "## Observed"
|
||||
echo
|
||||
echo "| Version | Port | Family | Initial snapshot | Grant | Revoke | API callback | Verdict |"
|
||||
echo "|---|---:|---|---|---|---|---|---|"
|
||||
awk -F '\t' 'NR > 1 {
|
||||
printf("| `%s` | `%s` | %s | %s | %s | %s | %s | %s |\n",
|
||||
$1, $3, $4, $5, $6, $7, $8, $9);
|
||||
}' "$RESULTS_TSV"
|
||||
|
||||
echo
|
||||
echo "## Artifact Links"
|
||||
echo
|
||||
awk -F '\t' 'NR > 1 {
|
||||
printf("- `%s`: run=`%s`, mcc=`%s`, server=`%s`, commands=`%s`\n", $1, $11, $12, $13, $14);
|
||||
printf(" note: %s\n", $10);
|
||||
}' "$RESULTS_TSV"
|
||||
|
||||
echo
|
||||
echo "## Inferred"
|
||||
echo
|
||||
echo "- Only rows with real MCC and server-log artifacts count as executed proof."
|
||||
echo "- Rows blocked by missing Java, tmux, or server directories are environment-limited, not product pass results."
|
||||
echo "- Rows with missing MCC or command-log artifacts should be treated as harness failures until rerun confirms a product issue."
|
||||
25
.skills/mcc-integration-testing/scripts/summarize_test_run.sh
Executable file
25
.skills/mcc-integration-testing/scripts/summarize_test_run.sh
Executable file
|
|
@ -0,0 +1,25 @@
|
|||
#!/usr/bin/env zsh
|
||||
set -euo pipefail
|
||||
|
||||
RUN_ROOT="${TMPDIR:-/tmp}/mcc-integration-testing"
|
||||
RUN_DIR="${1:-$(find "$RUN_ROOT" -mindepth 1 -maxdepth 1 -type d | sort | tail -n 1)}"
|
||||
|
||||
if [[ -z "${RUN_DIR:-}" ]] || [[ ! -d "$RUN_DIR" ]]; then
|
||||
echo "Run directory not found" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
MCC_LOG="$RUN_DIR/mcc.log"
|
||||
SERVER_LOG="$RUN_DIR/server-latest.log"
|
||||
BUILD_LOG="$RUN_DIR/build.log"
|
||||
|
||||
echo "Run directory: $RUN_DIR"
|
||||
echo
|
||||
echo "Build result:"
|
||||
grep -E "Warning\(s\)|Error\(s\)|Time Elapsed" "$BUILD_LOG" || true
|
||||
echo
|
||||
echo "MCC highlights:"
|
||||
grep -E "Server was successfully joined|FileInput|smoke_test_from_mcc_full_spectrum|There are [0-9]+ of a max|health|Creative" "$MCC_LOG" || true
|
||||
echo
|
||||
echo "Server highlights:"
|
||||
grep -E "joined the game|Made .* a server operator|game mode|smoke_test_from_mcc_full_spectrum|summon|particle|playsound|tnt" "$SERVER_LOG" || true
|
||||
351
.skills/mcc-version-adaptation/SKILL.md
Normal file
351
.skills/mcc-version-adaptation/SKILL.md
Normal file
|
|
@ -0,0 +1,351 @@
|
|||
---
|
||||
name: mcc-version-adaptation
|
||||
description: Adapt MCC palettes and protocol handling for a new Minecraft version. Use when the user wants to add support for a new MC version, compare version registries, update item/entity/block/metadata palettes, or fix protocol mismatches between MC versions.
|
||||
---
|
||||
|
||||
# MCC Version Adaptation
|
||||
|
||||
Systematic workflow for updating Minecraft Console Client to support a new Minecraft version, focusing on palette/registry changes and entity metadata.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Decompiled server source for both the old and new MC versions in `$MCC_REPO/MinecraftOfficial/<version>-decompiled/`
|
||||
- If missing, decompile and download server.jar:
|
||||
```bash
|
||||
$MCC_REPO/tools/decompile.sh --version <ver>
|
||||
```
|
||||
This auto-downloads `MinecraftDecompiler.jar` if needed, produces the decompiled source, and downloads `server.jar` into `$MCC_SERVERS/<ver>/`.
|
||||
- `tools/decompile.sh` depends on official mappings. For older versions where it refuses to decompile, fall back to a raw Java decompiler such as `cfr-decompiler` against `$MCC_SERVERS/<ver>/server.jar`. That fallback is good enough for packet inspection and registration order checks even when the output is obfuscated.
|
||||
- A test server of the target version in `$MCC_SERVERS/<version>/` (see `mcc-dev-workflow` skill)
|
||||
|
||||
## Step 0: Generate Server Reports (CRITICAL since 1.21.9)
|
||||
|
||||
**Before** analyzing decompiled source, generate authoritative registry data from the server jar:
|
||||
|
||||
```bash
|
||||
cd /tmp && java -DbundlerMainClass=net.minecraft.data.Main \
|
||||
-jar $MCC_SERVERS/<version>/server.jar \
|
||||
--reports --output /tmp/mc_reports
|
||||
```
|
||||
|
||||
This produces `/tmp/mc_reports/reports/` containing:
|
||||
- `registries.json` — all registries with **actual protocol_id** for each entry
|
||||
- `blocks.json` — all blocks with **block state IDs**
|
||||
- `packets.json` — packet protocol definitions
|
||||
|
||||
**Why this matters**: Since MC 1.21.9, some items and blocks are registered outside `Items.java`/`Blocks.java` field declarations (via block registration callbacks or other paths). The decompiled source alone will **miss** these entries. The server data generator is the only authoritative source for protocol IDs.
|
||||
|
||||
### Validation check
|
||||
Compare server registry counts against decompiled source counts:
|
||||
```bash
|
||||
python3 -c "
|
||||
import json
|
||||
with open('/tmp/mc_reports/reports/registries.json') as f:
|
||||
data = json.load(f)
|
||||
for reg in ['minecraft:item', 'minecraft:entity_type', 'minecraft:block']:
|
||||
print(f'{reg}: {len(data[reg][\"entries\"])} entries')
|
||||
"
|
||||
```
|
||||
|
||||
If server counts differ from decompiled Java source counts, the palette **must** be generated from server data, not from Java source.
|
||||
|
||||
## Step 1: Run Registry Diff
|
||||
|
||||
```bash
|
||||
python3 $MCC_REPO/tools/diff_registries.py <old_ver> <new_ver>
|
||||
```
|
||||
|
||||
This compares five registries and reports which need palette updates:
|
||||
|
||||
| Registry | MCC File | When to Update |
|
||||
|----------|----------|----------------|
|
||||
| Items.java | `ItemPalettes/ItemPaletteXXX.cs` | New/removed/reordered items |
|
||||
| EntityType.java | `EntityPalettes/EntityPaletteXXX.cs` | New/removed/reordered entity types |
|
||||
| Blocks.java | `BlockPalettes/BlockPaletteXXX.cs` | New/removed/reordered blocks |
|
||||
| DataComponents.java | `StructuredComponents/StructuredComponentsRegistryXXX.cs` | New/reordered components |
|
||||
| EntityDataSerializers.java | `EntityMetadataPalettes/EntityMetadataPaletteXXX.cs` | New/reordered serializer types |
|
||||
|
||||
**Important**: diff_registries.py compares decompiled Java source. If Step 0 revealed count mismatches, the diff output may undercount. Always cross-reference with server registries.json.
|
||||
|
||||
## Step 2: Generate Updated Palettes
|
||||
|
||||
For registries marked "PALETTE UPDATE NEEDED":
|
||||
|
||||
### Item Palette
|
||||
|
||||
**Preferred method** (accurate since 1.21.9):
|
||||
```bash
|
||||
python3 $MCC_REPO/tools/gen_item_palette.py --from-registry /tmp/mc_reports/reports/registries.json <suffix>
|
||||
# e.g., gen_item_palette.py --from-registry /tmp/mc_reports/reports/registries.json 1219
|
||||
```
|
||||
|
||||
**Legacy method** (works for versions where Items.java has all items):
|
||||
```bash
|
||||
python3 $MCC_REPO/tools/gen_item_palette.py <new_ver> <suffix>
|
||||
# e.g., gen_item_palette.py 1.21.1 121
|
||||
```
|
||||
|
||||
- If new items are reported missing from `ItemType.cs`, add them to the enum in alphabetical order.
|
||||
- The script auto-generates the C# palette file.
|
||||
|
||||
### Block Palette
|
||||
|
||||
**Preferred method** (accurate since 1.21.9):
|
||||
```bash
|
||||
python3 $MCC_REPO/tools/gen_block_palette.py /tmp/mc_reports/reports/blocks.json <suffix>
|
||||
# e.g., gen_block_palette.py /tmp/mc_reports/reports/blocks.json 1219
|
||||
```
|
||||
|
||||
**Legacy method** (manual creation from decompiled Blocks.java): Follow the pattern of existing palette files, using `register("name", ...)` call order from the decompiled source. Only reliable when Blocks.java contains all blocks.
|
||||
|
||||
If new blocks are reported missing from `Material.cs`, add them to the enum in alphabetical order.
|
||||
|
||||
### Entity Palette
|
||||
|
||||
```bash
|
||||
python3 $MCC_REPO/tools/gen_entity_palette.py /tmp/mc_reports/reports/registries.json <suffix>
|
||||
# e.g., gen_entity_palette.py /tmp/mc_reports/reports/registries.json 1219
|
||||
```
|
||||
|
||||
If new entity types are reported missing from `EntityType.cs`, add them to the enum in alphabetical order.
|
||||
|
||||
### Entity Metadata Palette
|
||||
```bash
|
||||
python3 $MCC_REPO/tools/gen_entity_metadata_palette.py <new_ver> <suffix>
|
||||
# e.g., gen_entity_metadata_palette.py 1.20.6 1206
|
||||
```
|
||||
- If new serializer types appear as UNMAPPED, add them to both:
|
||||
1. The script's `FIELD_TO_ENUM` dictionary
|
||||
2. MCC's `EntityMetaDataType.cs` enum
|
||||
3. `DataTypes.cs` read logic (add a `case` to consume the correct bytes)
|
||||
|
||||
### DataComponents / StructuredComponents
|
||||
Compare `DataComponents.java` registration order. If new components appear, update `StructuredComponentsRegistryXXX.cs`. For new component types, implement corresponding reader in `StructuredComponents/Components/`.
|
||||
|
||||
## Step 3: Update Version Routing
|
||||
|
||||
After creating palette files, update version selection logic:
|
||||
|
||||
| Palette Type | Routing Location |
|
||||
|-------------|-----------------|
|
||||
| Item | `Protocol18.cs` → `itemPalette` switch expression |
|
||||
| Entity | `Protocol18.cs` → `entityPalette` switch expression |
|
||||
| Block | `Protocol18.cs` → `blockPalette` initialization |
|
||||
| EntityMetadata | `EntityMetadataPalette.cs` → `GetPalette()` switch |
|
||||
| DataComponents | `StructuredComponentsRegistry.cs` → factory/routing |
|
||||
| Packet | `PacketType18Handler.cs` → `GetTypeHandler()` switch |
|
||||
|
||||
Pattern: add a new `>= MC_X_Y_Z_Version => new XxxPaletteXYZ()` case.
|
||||
|
||||
Also update:
|
||||
- `Protocol18.cs`: add `MC_X_Y_Z_Version = <protocol_number>` constant
|
||||
- `Protocol18.cs`: update all `> MC_prev_Version` upper-bound checks to `> MC_X_Y_Z_Version`
|
||||
- `ProtocolHandler.cs`: add version string → protocol mapping, protocol → version mapping, add to supported list
|
||||
- `Program.cs`: update `MCHighestVersion`
|
||||
|
||||
## Step 4: Check Packet Changes
|
||||
|
||||
Compare `GameProtocols.java` and `ConfigurationProtocols.java` between versions.
|
||||
|
||||
Common patterns:
|
||||
- **New clientbound packets inserted mid-list**: All subsequent packet IDs shift. Requires a new `PacketPalette` class.
|
||||
- **New packets appended at end**: Only need to add new enum values and entries in the palette.
|
||||
- **Packet renames** (same slot): Update MCC's packet type enum name but no ID change.
|
||||
|
||||
When packet changes are detected:
|
||||
1. Add new packet type enum values to `PacketTypesIn.cs`, `PacketTypesOut.cs`, `ConfigurationPacketTypesIn.cs`, `ConfigurationPacketTypesOut.cs`
|
||||
2. Create new `PacketPaletteXXX.cs` based on the previous one, adjusting IDs
|
||||
3. Update `PacketType18Handler.cs` routing
|
||||
|
||||
Use scriptable comparisons instead of eyeballing long packet tables. The packet ID is the registration index in `GameProtocols.java`:
|
||||
|
||||
```bash
|
||||
python3 - <<'PY'
|
||||
import re
|
||||
for ver in ["1.21.10", "1.21.11", "26.1"]:
|
||||
path=f"MinecraftOfficial/{ver}-decompiled/net/minecraft/network/protocol/game/GameProtocols.java"
|
||||
text=open(path).read()
|
||||
start=text.index("CLIENTBOUND_TEMPLATE")
|
||||
names=[m.group(1) for m in re.finditer(r"\.addPacket\(([^,]+),", text[start:])]
|
||||
print("==", ver, len(names))
|
||||
for i, name in enumerate(names):
|
||||
print(f"0x{i:02X}", name)
|
||||
PY
|
||||
```
|
||||
|
||||
For focused diffs:
|
||||
|
||||
```bash
|
||||
python3 - <<'PY'
|
||||
import re
|
||||
def packets(ver, marker):
|
||||
text=open(f"MinecraftOfficial/{ver}-decompiled/net/minecraft/network/protocol/game/GameProtocols.java").read()
|
||||
start=text.index(marker)
|
||||
return [m.group(1) for m in re.finditer(r"\.addPacket\(([^,]+),", text[start:])]
|
||||
left, right = "1.21.10", "1.21.11"
|
||||
a, b = packets(left, "CLIENTBOUND_TEMPLATE"), packets(right, "CLIENTBOUND_TEMPLATE")
|
||||
for i in range(max(len(a), len(b))):
|
||||
x = a[i] if i < len(a) else "<none>"
|
||||
y = b[i] if i < len(b) else "<none>"
|
||||
if x != y:
|
||||
print(f"0x{i:02X}: {left}={x} | {right}={y}")
|
||||
PY
|
||||
```
|
||||
|
||||
Do the same for `SERVERBOUND_TEMPLATE`. Clientbound and serverbound can change independently. Do not inherit a newer palette just because one side looks similar. For example, `1.21.11` used the same play packet order as `1.21.9/1.21.10` for the tested inventory path, while `26.1` had additional shifts.
|
||||
|
||||
## Step 5: Check Variant Encoding Changes
|
||||
|
||||
For entity types that use variant serializers (Cat, Wolf, Frog, Painting), check if the codec changed between versions by inspecting:
|
||||
|
||||
- `EntityDataSerializers.java` — look at how each `*_VARIANT` field is constructed
|
||||
- Key codecs:
|
||||
- `ByteBufCodecs.holderRegistry()` → wire format: `VarInt(registry_id)`
|
||||
- `ByteBufCodecs.holder()` → wire format: `VarInt(id+1)` for registered, `VarInt(0) + inline_data` for direct
|
||||
- If codec changed, update `DataTypes.cs` entity metadata reading logic accordingly.
|
||||
|
||||
## Step 6: Handle New EntityDataSerializer Types
|
||||
|
||||
When new serializer types are added (detected in Step 1):
|
||||
|
||||
1. Add enum value to `EntityMetaDataType.cs` with XML doc comment
|
||||
2. Add read logic in `DataTypes.cs` `ReadNextMetadata()`:
|
||||
- Determine byte consumption from the decompiled codec
|
||||
- Simple enum types (like CopperGolemState, WeatheringCopperState): `ReadNextVarInt(cache)`
|
||||
- Composite types (like ResolvableProfile): analyze the STREAM_CODEC chain in decompiled source
|
||||
3. Create the new palette file (Step 2)
|
||||
4. Update palette routing (Step 3)
|
||||
|
||||
## Step 7: Check SpawnEntity / Other Packet Format Changes
|
||||
|
||||
Compare key packet codec classes between versions. Known changes:
|
||||
- **1.21.9+**: `SpawnEntity` velocity fields changed from `short / 8000.0` to `LpVec3` format (VarLong-packed fixed-point). Gate reading in `DataTypes.ReadNextEntity()` by version.
|
||||
|
||||
When in doubt, compare the relevant packet class (e.g. `ClientboundAddEntityPacket.java`) between versions.
|
||||
|
||||
## Step 8: Update Block Collision Shapes (Physics Engine)
|
||||
|
||||
MCC's physics engine uses block collision shape data from PrismarineJS `minecraft-data` to perform accurate AABB collision detection (stored in `MinecraftClient/Physics/BlockShapeData.json`, embedded as a resource).
|
||||
|
||||
When a new MC version introduces new blocks or changes block shapes, update this data:
|
||||
|
||||
```bash
|
||||
# Download and compact collision shapes for the target version
|
||||
python3 $MCC_REPO/tools/gen_block_shapes.py <version>
|
||||
# e.g. python3 tools/gen_block_shapes.py 1.21.11
|
||||
```
|
||||
|
||||
If network is slow or unreliable, download the file manually and convert:
|
||||
```bash
|
||||
# Manual download
|
||||
curl -L -o /tmp/bcs.json \
|
||||
"https://raw.githubusercontent.com/PrismarineJS/minecraft-data/master/data/pc/<version>/blockCollisionShapes.json"
|
||||
|
||||
# Then compact from local file
|
||||
python3 $MCC_REPO/tools/gen_block_shapes.py --from-file /tmp/bcs.json
|
||||
```
|
||||
|
||||
Output: `MinecraftClient/Physics/BlockShapeData.json` (embedded via `MinecraftClient.csproj`)
|
||||
|
||||
The JSON maps block names (snake_case) → collision shape IDs → AABB coordinates. At runtime, `BlockShapes.cs` maps MCC's block state IDs to these AABBs using the block palette.
|
||||
|
||||
**When to update**: Whenever new blocks are added that have non-trivial collision shapes (e.g., new slab variants, stairs, fences). If only items or entities changed, this step can be skipped.
|
||||
|
||||
**Data source**: PrismarineJS `minecraft-data` repo, path: `data/pc/<version>/blockCollisionShapes.json`. Version availability can be checked via `data/dataPaths.json`.
|
||||
|
||||
## Step 9: Update Minimap Block Color Map
|
||||
|
||||
Regenerate the block-to-MapColor mapping used by the TUI minimap. This maps each block's `Material` enum to the RGB color from Minecraft's official `MapColor` table.
|
||||
|
||||
```bash
|
||||
python3 $MCC_REPO/tools/gen_block_color_map.py $MCC_REPO/MinecraftOfficial/<version>-decompiled
|
||||
# e.g. python3 tools/gen_block_color_map.py MinecraftOfficial/26.1-rc-2-decompiled
|
||||
```
|
||||
|
||||
Output: `MinecraftClient/Tui/MinimapBlockColors.json` (embedded as a resource via `.csproj`).
|
||||
|
||||
The script parses `MapColor.java`, `DyeColor.java`, and `Blocks.java` from the decompiled source to extract each block's assigned map color. Blocks not matched to a known `Material` enum value are skipped.
|
||||
|
||||
**When to update**: Whenever new blocks are added or existing blocks change their `mapColor()` assignment. If only items or entities changed, this step can be skipped.
|
||||
|
||||
## Step 10: Update Minimap Entity Categories
|
||||
|
||||
Regenerate the entity-to-MobCategory mapping used by the TUI minimap for classifying entities as hostile, passive, neutral, or non-living.
|
||||
|
||||
```bash
|
||||
python3 $MCC_REPO/tools/gen_entity_category_map.py $MCC_REPO/MinecraftOfficial/<version>-decompiled
|
||||
# e.g. python3 tools/gen_entity_category_map.py MinecraftOfficial/26.1-rc-2-decompiled
|
||||
```
|
||||
|
||||
Output: `MinecraftClient/Tui/MinimapEntityCategories.json` (embedded as a resource via `.csproj`).
|
||||
|
||||
The script parses `EntityType.java` to extract each entity's `MobCategory` assignment, then maps Minecraft's categories to MCC minimap categories:
|
||||
- `MONSTER` -> hostile (with neutral overrides for conditionally hostile mobs like Enderman, Spider, Wolf)
|
||||
- `CREATURE`/`AMBIENT`/`AXOLOTLS`/`WATER_*` -> passive
|
||||
- `MISC` -> non_living (with passive overrides for Villager, WanderingTrader, ZombieHorse)
|
||||
|
||||
The script maintains manual override lists for "neutral" mobs (attack only when provoked) since Minecraft has no machine-readable flag for this behavior. Review and update the `NEUTRAL_OVERRIDES` and `PASSIVE_OVERRIDES` sets in the script when new conditionally-hostile or misclassified mobs are added.
|
||||
|
||||
**When to update**: Whenever new entity types are added. If only blocks or items changed, this step can be skipped.
|
||||
|
||||
## Step 11: Compile and Verify
|
||||
|
||||
```bash
|
||||
dotnet build $MCC_REPO/MinecraftClient.sln -c Release
|
||||
```
|
||||
|
||||
Then connect to a test server of the target version (see `mcc-dev-workflow` skill) and verify:
|
||||
- Successful connection
|
||||
- `/give` new items → check inventory for correct identification
|
||||
- `/give` existing items (diamond_sword, etc.) → verify no ID shift
|
||||
- Summon new entities → check type and health
|
||||
- Summon variant entities (wolf, cat, frog) → no metadata parse errors
|
||||
- Place new blocks → `dig` reports correct block type
|
||||
- Teleport to distant chunks → terrain loads without errors
|
||||
- Chat commands work normally
|
||||
|
||||
**Always verify basic existing items first** (e.g. diamond_sword) to catch palette ID shift bugs early. If an existing item shows as the wrong type, the palette is using wrong protocol IDs.
|
||||
|
||||
## Key Source Files Reference
|
||||
|
||||
| Decompiled Java Source | Purpose |
|
||||
|----------------------|---------|
|
||||
| `world/item/Items.java` | Item registry (field declaration order ≈ ID, **but not always since 1.21.9**) |
|
||||
| `world/entity/EntityType.java` | Entity type registry (`register()` call order = ID) |
|
||||
| `world/level/block/Blocks.java` | Block registry (`register()` call order ≈ ID, **but not always since 1.21.9**) |
|
||||
| `core/component/DataComponents.java` | Data component registry |
|
||||
| `network/syncher/EntityDataSerializers.java` | Entity metadata type registry (static block order = ID) |
|
||||
| `network/protocol/game/GameProtocols.java` | Play packet registration order (= packet IDs) |
|
||||
| `network/protocol/configuration/ConfigurationProtocols.java` | Config packet registration order |
|
||||
|
||||
| Server Data Generator Output | Purpose |
|
||||
|-----|---------|
|
||||
| `registries.json` | **Authoritative** protocol_id for all registries |
|
||||
| `blocks.json` | **Authoritative** block state IDs |
|
||||
| `packets.json` | Packet protocol definitions |
|
||||
|
||||
## Common Pitfalls
|
||||
|
||||
- **Source field order ≠ runtime registry ID (since 1.21.9)**: Some items/blocks are registered via callbacks (e.g., block items registered by `Blocks.java` during block registration) rather than in `Items.java` field declarations. Always validate palette counts against server `registries.json`. If counts differ, **use server data generator output instead of decompiled source**.
|
||||
- **ID order matters**: IDs are determined by registration order, not alphabetical. Always use server data generator as ground truth.
|
||||
- **Cross-version jumps**: When MCC skips versions (e.g., 1.20.4→1.20.6), registries from ALL intermediate versions may have changed. Always diff against the actual last-supported version, not the latest palette.
|
||||
- **EntityMetadata type shifts**: A single new serializer type shifts all subsequent IDs, causing widespread metadata parse failures. Symptoms: entity rendering glitches, disconnections, or silent data corruption.
|
||||
- **CUT_STANDSTONE_SLAB**: This is an intentional typo in Minecraft source (should be SANDSTONE). MCC's `ItemType.cs` uses `CutSandstoneSlab` — the gen script handles this via the OVERRIDES dict.
|
||||
- **Item/block renames across versions**: Some items/blocks get renamed (e.g., `DRY_SHORT_GRASS` → `SHORT_DRY_GRASS`, `CHAIN` → `IRON_CHAIN`). Keep old enum values for backward compatibility with older palettes, and add new ones for the new version.
|
||||
- **Packet ID cascading shifts**: Even one inserted mid-list clientbound packet shifts ALL subsequent IDs. Always create a new PacketPalette for protocol changes.
|
||||
- **Test existing items first**: After palette changes, always verify existing items (diamond_sword, stone, etc.) before testing new ones. If they show as wrong items, the palette has a systemic ID offset bug.
|
||||
|
||||
## Reusable Scripts
|
||||
|
||||
All scripts are in `$MCC_REPO/tools/`. See `tools/README.md` for detailed usage.
|
||||
|
||||
| Script | Purpose | Input |
|
||||
|--------|---------|-------|
|
||||
| `diff_registries.py` | Compare registries between versions | Decompiled source |
|
||||
| `gen_item_palette.py` | Generate ItemPalette C# | Decompiled source OR registries.json |
|
||||
| `gen_block_palette.py` | Generate BlockPalette C# | blocks.json |
|
||||
| `gen_entity_palette.py` | Generate EntityPalette C# | registries.json |
|
||||
| `gen_entity_metadata_palette.py` | Generate EntityMetadataPalette C# | Decompiled source |
|
||||
| `gen_block_shapes.py` | Download & compact block collision shapes | PrismarineJS minecraft-data |
|
||||
| `gen_block_color_map.py` | Generate minimap block color JSON | Decompiled source (MapColor/DyeColor/Blocks) |
|
||||
| `gen_entity_category_map.py` | Generate minimap entity category JSON | Decompiled source (EntityType.java) |
|
||||
211
.skills/mermaid-diagrams/SKILL.md
Normal file
211
.skills/mermaid-diagrams/SKILL.md
Normal file
|
|
@ -0,0 +1,211 @@
|
|||
---
|
||||
name: mermaid-diagrams
|
||||
description: Creating and refining Mermaid diagrams with live reload. Use when users want flowcharts, sequence diagrams, class diagrams, ER diagrams, state diagrams, or any other Mermaid visualization. Provides best practices for syntax, styling, and the iterative workflow using mermaid_preview and mermaid_save tools.
|
||||
allowed-tools: mcp__mermaid__mermaid_preview, mcp__mermaid__mermaid_save
|
||||
---
|
||||
|
||||
# Mermaid Diagram Expert
|
||||
|
||||
You are an expert at creating, refining, and optimizing Mermaid diagrams using the MCP server tools.
|
||||
|
||||
## Core Workflow
|
||||
|
||||
1. **Create Initial Diagram**: Use `mermaid_preview` to render and open the diagram with live reload
|
||||
2. **Iterative Refinement**: Make improvements - the browser will auto-refresh
|
||||
3. **Save Final Version**: Use `mermaid_save` when satisfied
|
||||
|
||||
## Tool Usage
|
||||
|
||||
### mermaid_preview
|
||||
|
||||
Always use this when creating or updating diagrams:
|
||||
|
||||
- `diagram`: The Mermaid code
|
||||
- `preview_id`: Descriptive kebab-case ID (e.g., `auth-flow`, `architecture`)
|
||||
- `format`: Use `svg` for live reload (default)
|
||||
- `theme`: `default`, `forest`, `dark`, or `neutral`
|
||||
- `background`: `white`, `transparent`, or hex colors
|
||||
- `width`, `height`, `scale`: Adjust for quality/size
|
||||
|
||||
**Key Points:**
|
||||
|
||||
- Reuse the same `preview_id` for refinements to update the same browser tab
|
||||
- Use different IDs for multiple simultaneous diagrams
|
||||
- Live reload only works with SVG format
|
||||
|
||||
### mermaid_save
|
||||
|
||||
Use after the diagram is finalized:
|
||||
|
||||
- `save_path`: Where to save (e.g., `./docs/diagram.svg`)
|
||||
- `preview_id`: Must match the preview ID used earlier
|
||||
- `format`: Must match format from preview
|
||||
|
||||
## Diagram Types
|
||||
|
||||
### Flowcharts (`graph` or `flowchart`)
|
||||
|
||||
Direction: `LR`, `TB`, `RL`, `BT`
|
||||
|
||||
```mermaid
|
||||
graph LR
|
||||
A[Start] --> B{Decision}
|
||||
B -->|Yes| C[Action]
|
||||
B -->|No| D[End]
|
||||
|
||||
style A fill:#e1f5ff
|
||||
style C fill:#d4edda
|
||||
```
|
||||
|
||||
### Sequence Diagrams (`sequenceDiagram`)
|
||||
|
||||
⚠️ **Do NOT use `style` statements** - not supported
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant User
|
||||
participant App
|
||||
participant API
|
||||
|
||||
User->>App: Login
|
||||
App->>API: Authenticate
|
||||
API-->>App: Token
|
||||
App-->>User: Success
|
||||
```
|
||||
|
||||
### Class Diagrams (`classDiagram`)
|
||||
|
||||
```mermaid
|
||||
classDiagram
|
||||
class User {
|
||||
+String name
|
||||
+String email
|
||||
+login()
|
||||
}
|
||||
class Order {
|
||||
+int id
|
||||
+Date created
|
||||
}
|
||||
User "1" --> "*" Order
|
||||
```
|
||||
|
||||
### Entity Relationship (`erDiagram`)
|
||||
|
||||
```mermaid
|
||||
erDiagram
|
||||
USER ||--o{ ORDER : places
|
||||
ORDER ||--|{ LINE_ITEM : contains
|
||||
|
||||
USER {
|
||||
int id PK
|
||||
string email
|
||||
string name
|
||||
}
|
||||
```
|
||||
|
||||
### State Diagrams (`stateDiagram-v2`)
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
[*] --> Idle
|
||||
Idle --> Processing : start
|
||||
Processing --> Complete : finish
|
||||
Complete --> [*]
|
||||
```
|
||||
|
||||
### Gantt Charts (`gantt`)
|
||||
|
||||
```mermaid
|
||||
gantt
|
||||
title Project Timeline
|
||||
section Phase 1
|
||||
Task 1 :a1, 2024-01-01, 30d
|
||||
Task 2 :after a1, 20d
|
||||
```
|
||||
|
||||
## Best Practices
|
||||
|
||||
### Preview IDs
|
||||
|
||||
- Use descriptive names: `architecture`, `auth-flow`, `data-model`
|
||||
- Keep the same ID during refinements
|
||||
- Use different IDs for concurrent diagrams
|
||||
|
||||
### Themes & Styling
|
||||
|
||||
- `default`: Clean, professional
|
||||
- `forest`: Green tones
|
||||
- `dark`: Dark background
|
||||
- `neutral`: Grayscale
|
||||
|
||||
Use `transparent` background for docs, `white` for standalone
|
||||
|
||||
### Common Patterns
|
||||
|
||||
**System Architecture:**
|
||||
|
||||
```mermaid
|
||||
graph TB
|
||||
Client[Web App]
|
||||
API[API Gateway]
|
||||
DB[(Database)]
|
||||
|
||||
Client --> API --> DB
|
||||
```
|
||||
|
||||
**Authentication Flow:**
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
User->>App: Login Request
|
||||
App->>Auth: Validate
|
||||
Auth-->>App: JWT Token
|
||||
App-->>User: Access Granted
|
||||
```
|
||||
|
||||
## User Interaction
|
||||
|
||||
When a user requests a diagram:
|
||||
|
||||
1. **Clarify if needed**: What type? What level of detail?
|
||||
2. **Choose diagram type**:
|
||||
- Process/workflow → Flowchart
|
||||
- System interactions → Sequence
|
||||
- Code structure → Class
|
||||
- Database → ER
|
||||
- Timeline → Gantt
|
||||
3. **Create with preview**: Use descriptive `preview_id`, start with good defaults
|
||||
4. **Iterate**: Keep same `preview_id`, explain changes
|
||||
5. **Save**: Ask where/what format, use `mermaid_save`
|
||||
|
||||
## Proactive Behavior
|
||||
|
||||
- Always preview diagrams, don't just generate code
|
||||
- Use sensible defaults without asking
|
||||
- Reuse preview_id for refinements
|
||||
- Suggest improvements when you see opportunities
|
||||
- Explain your diagram type choice briefly
|
||||
|
||||
## Common Issues
|
||||
|
||||
**Syntax errors**: Check quotes, arrow syntax, keywords
|
||||
**Layout issues**: Try different directions (LR vs TB)
|
||||
**Text overlap**: Increase dimensions or shorten labels
|
||||
**Colors not working**: Verify CSS color format; remember sequence diagrams don't support styles
|
||||
|
||||
## Example Interaction
|
||||
|
||||
**User**: "Create an auth flow diagram"
|
||||
|
||||
**You**: "I'll create a sequence diagram showing the authentication flow."
|
||||
[Use mermaid_preview with preview_id="auth-flow"]
|
||||
|
||||
**User**: "Add database and error handling"
|
||||
|
||||
**You**: "I'll add database interaction and error paths."
|
||||
[Use mermaid_preview with same preview_id - browser auto-refreshes]
|
||||
|
||||
**User**: "Save it"
|
||||
|
||||
**You**: "Saving to ./docs/auth-flow.svg"
|
||||
[Use mermaid_save]
|
||||
202
.skills/skill-creator/LICENSE.txt
Normal file
202
.skills/skill-creator/LICENSE.txt
Normal file
|
|
@ -0,0 +1,202 @@
|
|||
|
||||
Apache License
|
||||
Version 2.0, January 2004
|
||||
http://www.apache.org/licenses/
|
||||
|
||||
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
|
||||
|
||||
1. Definitions.
|
||||
|
||||
"License" shall mean the terms and conditions for use, reproduction,
|
||||
and distribution as defined by Sections 1 through 9 of this document.
|
||||
|
||||
"Licensor" shall mean the copyright owner or entity authorized by
|
||||
the copyright owner that is granting the License.
|
||||
|
||||
"Legal Entity" shall mean the union of the acting entity and all
|
||||
other entities that control, are controlled by, or are under common
|
||||
control with that entity. For the purposes of this definition,
|
||||
"control" means (i) the power, direct or indirect, to cause the
|
||||
direction or management of such entity, whether by contract or
|
||||
otherwise, or (ii) ownership of fifty percent (50%) or more of the
|
||||
outstanding shares, or (iii) beneficial ownership of such entity.
|
||||
|
||||
"You" (or "Your") shall mean an individual or Legal Entity
|
||||
exercising permissions granted by this License.
|
||||
|
||||
"Source" form shall mean the preferred form for making modifications,
|
||||
including but not limited to software source code, documentation
|
||||
source, and configuration files.
|
||||
|
||||
"Object" form shall mean any form resulting from mechanical
|
||||
transformation or translation of a Source form, including but
|
||||
not limited to compiled object code, generated documentation,
|
||||
and conversions to other media types.
|
||||
|
||||
"Work" shall mean the work of authorship, whether in Source or
|
||||
Object form, made available under the License, as indicated by a
|
||||
copyright notice that is included in or attached to the work
|
||||
(an example is provided in the Appendix below).
|
||||
|
||||
"Derivative Works" shall mean any work, whether in Source or Object
|
||||
form, that is based on (or derived from) the Work and for which the
|
||||
editorial revisions, annotations, elaborations, or other modifications
|
||||
represent, as a whole, an original work of authorship. For the purposes
|
||||
of this License, Derivative Works shall not include works that remain
|
||||
separable from, or merely link (or bind by name) to the interfaces of,
|
||||
the Work and Derivative Works thereof.
|
||||
|
||||
"Contribution" shall mean any work of authorship, including
|
||||
the original version of the Work and any modifications or additions
|
||||
to that Work or Derivative Works thereof, that is intentionally
|
||||
submitted to Licensor for inclusion in the Work by the copyright owner
|
||||
or by an individual or Legal Entity authorized to submit on behalf of
|
||||
the copyright owner. For the purposes of this definition, "submitted"
|
||||
means any form of electronic, verbal, or written communication sent
|
||||
to the Licensor or its representatives, including but not limited to
|
||||
communication on electronic mailing lists, source code control systems,
|
||||
and issue tracking systems that are managed by, or on behalf of, the
|
||||
Licensor for the purpose of discussing and improving the Work, but
|
||||
excluding communication that is conspicuously marked or otherwise
|
||||
designated in writing by the copyright owner as "Not a Contribution."
|
||||
|
||||
"Contributor" shall mean Licensor and any individual or Legal Entity
|
||||
on behalf of whom a Contribution has been received by Licensor and
|
||||
subsequently incorporated within the Work.
|
||||
|
||||
2. Grant of Copyright License. Subject to the terms and conditions of
|
||||
this License, each Contributor hereby grants to You a perpetual,
|
||||
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
||||
copyright license to reproduce, prepare Derivative Works of,
|
||||
publicly display, publicly perform, sublicense, and distribute the
|
||||
Work and such Derivative Works in Source or Object form.
|
||||
|
||||
3. Grant of Patent License. Subject to the terms and conditions of
|
||||
this License, each Contributor hereby grants to You a perpetual,
|
||||
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
||||
(except as stated in this section) patent license to make, have made,
|
||||
use, offer to sell, sell, import, and otherwise transfer the Work,
|
||||
where such license applies only to those patent claims licensable
|
||||
by such Contributor that are necessarily infringed by their
|
||||
Contribution(s) alone or by combination of their Contribution(s)
|
||||
with the Work to which such Contribution(s) was submitted. If You
|
||||
institute patent litigation against any entity (including a
|
||||
cross-claim or counterclaim in a lawsuit) alleging that the Work
|
||||
or a Contribution incorporated within the Work constitutes direct
|
||||
or contributory patent infringement, then any patent licenses
|
||||
granted to You under this License for that Work shall terminate
|
||||
as of the date such litigation is filed.
|
||||
|
||||
4. Redistribution. You may reproduce and distribute copies of the
|
||||
Work or Derivative Works thereof in any medium, with or without
|
||||
modifications, and in Source or Object form, provided that You
|
||||
meet the following conditions:
|
||||
|
||||
(a) You must give any other recipients of the Work or
|
||||
Derivative Works a copy of this License; and
|
||||
|
||||
(b) You must cause any modified files to carry prominent notices
|
||||
stating that You changed the files; and
|
||||
|
||||
(c) You must retain, in the Source form of any Derivative Works
|
||||
that You distribute, all copyright, patent, trademark, and
|
||||
attribution notices from the Source form of the Work,
|
||||
excluding those notices that do not pertain to any part of
|
||||
the Derivative Works; and
|
||||
|
||||
(d) If the Work includes a "NOTICE" text file as part of its
|
||||
distribution, then any Derivative Works that You distribute must
|
||||
include a readable copy of the attribution notices contained
|
||||
within such NOTICE file, excluding those notices that do not
|
||||
pertain to any part of the Derivative Works, in at least one
|
||||
of the following places: within a NOTICE text file distributed
|
||||
as part of the Derivative Works; within the Source form or
|
||||
documentation, if provided along with the Derivative Works; or,
|
||||
within a display generated by the Derivative Works, if and
|
||||
wherever such third-party notices normally appear. The contents
|
||||
of the NOTICE file are for informational purposes only and
|
||||
do not modify the License. You may add Your own attribution
|
||||
notices within Derivative Works that You distribute, alongside
|
||||
or as an addendum to the NOTICE text from the Work, provided
|
||||
that such additional attribution notices cannot be construed
|
||||
as modifying the License.
|
||||
|
||||
You may add Your own copyright statement to Your modifications and
|
||||
may provide additional or different license terms and conditions
|
||||
for use, reproduction, or distribution of Your modifications, or
|
||||
for any such Derivative Works as a whole, provided Your use,
|
||||
reproduction, and distribution of the Work otherwise complies with
|
||||
the conditions stated in this License.
|
||||
|
||||
5. Submission of Contributions. Unless You explicitly state otherwise,
|
||||
any Contribution intentionally submitted for inclusion in the Work
|
||||
by You to the Licensor shall be under the terms and conditions of
|
||||
this License, without any additional terms or conditions.
|
||||
Notwithstanding the above, nothing herein shall supersede or modify
|
||||
the terms of any separate license agreement you may have executed
|
||||
with Licensor regarding such Contributions.
|
||||
|
||||
6. Trademarks. This License does not grant permission to use the trade
|
||||
names, trademarks, service marks, or product names of the Licensor,
|
||||
except as required for reasonable and customary use in describing the
|
||||
origin of the Work and reproducing the content of the NOTICE file.
|
||||
|
||||
7. Disclaimer of Warranty. Unless required by applicable law or
|
||||
agreed to in writing, Licensor provides the Work (and each
|
||||
Contributor provides its Contributions) on an "AS IS" BASIS,
|
||||
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
|
||||
implied, including, without limitation, any warranties or conditions
|
||||
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
|
||||
PARTICULAR PURPOSE. You are solely responsible for determining the
|
||||
appropriateness of using or redistributing the Work and assume any
|
||||
risks associated with Your exercise of permissions under this License.
|
||||
|
||||
8. Limitation of Liability. In no event and under no legal theory,
|
||||
whether in tort (including negligence), contract, or otherwise,
|
||||
unless required by applicable law (such as deliberate and grossly
|
||||
negligent acts) or agreed to in writing, shall any Contributor be
|
||||
liable to You for damages, including any direct, indirect, special,
|
||||
incidental, or consequential damages of any character arising as a
|
||||
result of this License or out of the use or inability to use the
|
||||
Work (including but not limited to damages for loss of goodwill,
|
||||
work stoppage, computer failure or malfunction, or any and all
|
||||
other commercial damages or losses), even if such Contributor
|
||||
has been advised of the possibility of such damages.
|
||||
|
||||
9. Accepting Warranty or Additional Liability. While redistributing
|
||||
the Work or Derivative Works thereof, You may choose to offer,
|
||||
and charge a fee for, acceptance of support, warranty, indemnity,
|
||||
or other liability obligations and/or rights consistent with this
|
||||
License. However, in accepting such obligations, You may act only
|
||||
on Your own behalf and on Your sole responsibility, not on behalf
|
||||
of any other Contributor, and only if You agree to indemnify,
|
||||
defend, and hold each Contributor harmless for any liability
|
||||
incurred by, or claims asserted against, such Contributor by reason
|
||||
of your accepting any such warranty or additional liability.
|
||||
|
||||
END OF TERMS AND CONDITIONS
|
||||
|
||||
APPENDIX: How to apply the Apache License to your work.
|
||||
|
||||
To apply the Apache License to your work, attach the following
|
||||
boilerplate notice, with the fields enclosed by brackets "[]"
|
||||
replaced with your own identifying information. (Don't include
|
||||
the brackets!) The text should be enclosed in the appropriate
|
||||
comment syntax for the file format. We also recommend that a
|
||||
file or class name and description of purpose be included on the
|
||||
same "printed page" as the copyright notice for easier
|
||||
identification within third-party archives.
|
||||
|
||||
Copyright [yyyy] [name of copyright owner]
|
||||
|
||||
Licensed under the Apache License, Version 2.0 (the "License");
|
||||
you may not use this file except in compliance with the License.
|
||||
You may obtain a copy of the License at
|
||||
|
||||
http://www.apache.org/licenses/LICENSE-2.0
|
||||
|
||||
Unless required by applicable law or agreed to in writing, software
|
||||
distributed under the License is distributed on an "AS IS" BASIS,
|
||||
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
See the License for the specific language governing permissions and
|
||||
limitations under the License.
|
||||
479
.skills/skill-creator/SKILL.md
Normal file
479
.skills/skill-creator/SKILL.md
Normal file
|
|
@ -0,0 +1,479 @@
|
|||
---
|
||||
name: skill-creator
|
||||
description: Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, update or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.
|
||||
---
|
||||
|
||||
# Skill Creator
|
||||
|
||||
A skill for creating new skills and iteratively improving them.
|
||||
|
||||
At a high level, the process of creating a skill goes like this:
|
||||
|
||||
- Decide what you want the skill to do and roughly how it should do it
|
||||
- Write a draft of the skill
|
||||
- Create a few test prompts and run claude-with-access-to-the-skill on them
|
||||
- Help the user evaluate the results both qualitatively and quantitatively
|
||||
- While the runs happen in the background, draft some quantitative evals if there aren't any (if there are some, you can either use as is or modify if you feel something needs to change about them). Then explain them to the user (or if they already existed, explain the ones that already exist)
|
||||
- Use the `eval-viewer/generate_review.py` script to show the user the results for them to look at, and also let them look at the quantitative metrics
|
||||
- Rewrite the skill based on feedback from the user's evaluation of the results (and also if there are any glaring flaws that become apparent from the quantitative benchmarks)
|
||||
- Repeat until you're satisfied
|
||||
- Expand the test set and try again at larger scale
|
||||
|
||||
Your job when using this skill is to figure out where the user is in this process and then jump in and help them progress through these stages. So for instance, maybe they're like "I want to make a skill for X". You can help narrow down what they mean, write a draft, write the test cases, figure out how they want to evaluate, run all the prompts, and repeat.
|
||||
|
||||
On the other hand, maybe they already have a draft of the skill. In this case you can go straight to the eval/iterate part of the loop.
|
||||
|
||||
Of course, you should always be flexible and if the user is like "I don't need to run a bunch of evaluations, just vibe with me", you can do that instead.
|
||||
|
||||
Then after the skill is done (but again, the order is flexible), you can also run the skill description improver, which we have a whole separate script for, to optimize the triggering of the skill.
|
||||
|
||||
Cool? Cool.
|
||||
|
||||
## Communicating with the user
|
||||
|
||||
The skill creator is liable to be used by people across a wide range of familiarity with coding jargon. If you haven't heard (and how could you, it's only very recently that it started), there's a trend now where the power of Claude is inspiring plumbers to open up their terminals, parents and grandparents to google "how to install npm". On the other hand, the bulk of users are probably fairly computer-literate.
|
||||
|
||||
So please pay attention to context cues to understand how to phrase your communication! In the default case, just to give you some idea:
|
||||
|
||||
- "evaluation" and "benchmark" are borderline, but OK
|
||||
- for "JSON" and "assertion" you want to see serious cues from the user that they know what those things are before using them without explaining them
|
||||
|
||||
It's OK to briefly explain terms if you're in doubt, and feel free to clarify terms with a short definition if you're unsure if the user will get it.
|
||||
|
||||
---
|
||||
|
||||
## Creating a skill
|
||||
|
||||
### Capture Intent
|
||||
|
||||
Start by understanding the user's intent. The current conversation might already contain a workflow the user wants to capture (e.g., they say "turn this into a skill"). If so, extract answers from the conversation history first — the tools used, the sequence of steps, corrections the user made, input/output formats observed. The user may need to fill the gaps, and should confirm before proceeding to the next step.
|
||||
|
||||
1. What should this skill enable Claude to do?
|
||||
2. When should this skill trigger? (what user phrases/contexts)
|
||||
3. What's the expected output format?
|
||||
4. Should we set up test cases to verify the skill works? Skills with objectively verifiable outputs (file transforms, data extraction, code generation, fixed workflow steps) benefit from test cases. Skills with subjective outputs (writing style, art) often don't need them. Suggest the appropriate default based on the skill type, but let the user decide.
|
||||
|
||||
### Interview and Research
|
||||
|
||||
Proactively ask questions about edge cases, input/output formats, example files, success criteria, and dependencies. Wait to write test prompts until you've got this part ironed out.
|
||||
|
||||
Check available MCPs - if useful for research (searching docs, finding similar skills, looking up best practices), research in parallel via subagents if available, otherwise inline. Come prepared with context to reduce burden on the user.
|
||||
|
||||
### Write the SKILL.md
|
||||
|
||||
Based on the user interview, fill in these components:
|
||||
|
||||
- **name**: Skill identifier
|
||||
- **description**: When to trigger, what it does. This is the primary triggering mechanism - include both what the skill does AND specific contexts for when to use it. All "when to use" info goes here, not in the body. Note: currently Claude has a tendency to "undertrigger" skills -- to not use them when they'd be useful. To combat this, please make the skill descriptions a little bit "pushy". So for instance, instead of "How to build a simple fast dashboard to display internal Anthropic data.", you might write "How to build a simple fast dashboard to display internal Anthropic data. Make sure to use this skill whenever the user mentions dashboards, data visualization, internal metrics, or wants to display any kind of company data, even if they don't explicitly ask for a 'dashboard.'"
|
||||
- **compatibility**: Required tools, dependencies (optional, rarely needed)
|
||||
- **the rest of the skill :)**
|
||||
|
||||
### Skill Writing Guide
|
||||
|
||||
#### Anatomy of a Skill
|
||||
|
||||
```
|
||||
skill-name/
|
||||
├── SKILL.md (required)
|
||||
│ ├── YAML frontmatter (name, description required)
|
||||
│ └── Markdown instructions
|
||||
└── Bundled Resources (optional)
|
||||
├── scripts/ - Executable code for deterministic/repetitive tasks
|
||||
├── references/ - Docs loaded into context as needed
|
||||
└── assets/ - Files used in output (templates, icons, fonts)
|
||||
```
|
||||
|
||||
#### Progressive Disclosure
|
||||
|
||||
Skills use a three-level loading system:
|
||||
1. **Metadata** (name + description) - Always in context (~100 words)
|
||||
2. **SKILL.md body** - In context whenever skill triggers (<500 lines ideal)
|
||||
3. **Bundled resources** - As needed (unlimited, scripts can execute without loading)
|
||||
|
||||
These word counts are approximate and you can feel free to go longer if needed.
|
||||
|
||||
**Key patterns:**
|
||||
- Keep SKILL.md under 500 lines; if you're approaching this limit, add an additional layer of hierarchy along with clear pointers about where the model using the skill should go next to follow up.
|
||||
- Reference files clearly from SKILL.md with guidance on when to read them
|
||||
- For large reference files (>300 lines), include a table of contents
|
||||
|
||||
**Domain organization**: When a skill supports multiple domains/frameworks, organize by variant:
|
||||
```
|
||||
cloud-deploy/
|
||||
├── SKILL.md (workflow + selection)
|
||||
└── references/
|
||||
├── aws.md
|
||||
├── gcp.md
|
||||
└── azure.md
|
||||
```
|
||||
Claude reads only the relevant reference file.
|
||||
|
||||
#### Principle of Lack of Surprise
|
||||
|
||||
This goes without saying, but skills must not contain malware, exploit code, or any content that could compromise system security. A skill's contents should not surprise the user in their intent if described. Don't go along with requests to create misleading skills or skills designed to facilitate unauthorized access, data exfiltration, or other malicious activities. Things like a "roleplay as an XYZ" are OK though.
|
||||
|
||||
#### Writing Patterns
|
||||
|
||||
Prefer using the imperative form in instructions.
|
||||
|
||||
**Defining output formats** - You can do it like this:
|
||||
```markdown
|
||||
## Report structure
|
||||
ALWAYS use this exact template:
|
||||
# [Title]
|
||||
## Executive summary
|
||||
## Key findings
|
||||
## Recommendations
|
||||
```
|
||||
|
||||
**Examples pattern** - It's useful to include examples. You can format them like this (but if "Input" and "Output" are in the examples you might want to deviate a little):
|
||||
```markdown
|
||||
## Commit message format
|
||||
**Example 1:**
|
||||
Input: Added user authentication with JWT tokens
|
||||
Output: feat(auth): implement JWT-based authentication
|
||||
```
|
||||
|
||||
### Writing Style
|
||||
|
||||
Try to explain to the model why things are important in lieu of heavy-handed musty MUSTs. Use theory of mind and try to make the skill general and not super-narrow to specific examples. Start by writing a draft and then look at it with fresh eyes and improve it.
|
||||
|
||||
### Test Cases
|
||||
|
||||
After writing the skill draft, come up with 2-3 realistic test prompts — the kind of thing a real user would actually say. Share them with the user: [you don't have to use this exact language] "Here are a few test cases I'd like to try. Do these look right, or do you want to add more?" Then run them.
|
||||
|
||||
Save test cases to `evals/evals.json`. Don't write assertions yet — just the prompts. You'll draft assertions in the next step while the runs are in progress.
|
||||
|
||||
```json
|
||||
{
|
||||
"skill_name": "example-skill",
|
||||
"evals": [
|
||||
{
|
||||
"id": 1,
|
||||
"prompt": "User's task prompt",
|
||||
"expected_output": "Description of expected result",
|
||||
"files": []
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
See `references/schemas.md` for the full schema (including the `assertions` field, which you'll add later).
|
||||
|
||||
## Running and evaluating test cases
|
||||
|
||||
This section is one continuous sequence — don't stop partway through. Do NOT use `/skill-test` or any other testing skill.
|
||||
|
||||
Put results in `<skill-name>-workspace/` as a sibling to the skill directory. Within the workspace, organize results by iteration (`iteration-1/`, `iteration-2/`, etc.) and within that, each test case gets a directory (`eval-0/`, `eval-1/`, etc.). Don't create all of this upfront — just create directories as you go.
|
||||
|
||||
### Step 1: Spawn all runs (with-skill AND baseline) in the same turn
|
||||
|
||||
For each test case, spawn two subagents in the same turn — one with the skill, one without. This is important: don't spawn the with-skill runs first and then come back for baselines later. Launch everything at once so it all finishes around the same time.
|
||||
|
||||
**With-skill run:**
|
||||
|
||||
```
|
||||
Execute this task:
|
||||
- Skill path: <path-to-skill>
|
||||
- Task: <eval prompt>
|
||||
- Input files: <eval files if any, or "none">
|
||||
- Save outputs to: <workspace>/iteration-<N>/eval-<ID>/with_skill/outputs/
|
||||
- Outputs to save: <what the user cares about — e.g., "the .docx file", "the final CSV">
|
||||
```
|
||||
|
||||
**Baseline run** (same prompt, but the baseline depends on context):
|
||||
- **Creating a new skill**: no skill at all. Same prompt, no skill path, save to `without_skill/outputs/`.
|
||||
- **Improving an existing skill**: the old version. Before editing, snapshot the skill (`cp -r <skill-path> <workspace>/skill-snapshot/`), then point the baseline subagent at the snapshot. Save to `old_skill/outputs/`.
|
||||
|
||||
Write an `eval_metadata.json` for each test case (assertions can be empty for now). Give each eval a descriptive name based on what it's testing — not just "eval-0". Use this name for the directory too. If this iteration uses new or modified eval prompts, create these files for each new eval directory — don't assume they carry over from previous iterations.
|
||||
|
||||
```json
|
||||
{
|
||||
"eval_id": 0,
|
||||
"eval_name": "descriptive-name-here",
|
||||
"prompt": "The user's task prompt",
|
||||
"assertions": []
|
||||
}
|
||||
```
|
||||
|
||||
### Step 2: While runs are in progress, draft assertions
|
||||
|
||||
Don't just wait for the runs to finish — you can use this time productively. Draft quantitative assertions for each test case and explain them to the user. If assertions already exist in `evals/evals.json`, review them and explain what they check.
|
||||
|
||||
Good assertions are objectively verifiable and have descriptive names — they should read clearly in the benchmark viewer so someone glancing at the results immediately understands what each one checks. Subjective skills (writing style, design quality) are better evaluated qualitatively — don't force assertions onto things that need human judgment.
|
||||
|
||||
Update the `eval_metadata.json` files and `evals/evals.json` with the assertions once drafted. Also explain to the user what they'll see in the viewer — both the qualitative outputs and the quantitative benchmark.
|
||||
|
||||
### Step 3: As runs complete, capture timing data
|
||||
|
||||
When each subagent task completes, you receive a notification containing `total_tokens` and `duration_ms`. Save this data immediately to `timing.json` in the run directory:
|
||||
|
||||
```json
|
||||
{
|
||||
"total_tokens": 84852,
|
||||
"duration_ms": 23332,
|
||||
"total_duration_seconds": 23.3
|
||||
}
|
||||
```
|
||||
|
||||
This is the only opportunity to capture this data — it comes through the task notification and isn't persisted elsewhere. Process each notification as it arrives rather than trying to batch them.
|
||||
|
||||
### Step 4: Grade, aggregate, and launch the viewer
|
||||
|
||||
Once all runs are done:
|
||||
|
||||
1. **Grade each run** — spawn a grader subagent (or grade inline) that reads `agents/grader.md` and evaluates each assertion against the outputs. Save results to `grading.json` in each run directory. The grading.json expectations array must use the fields `text`, `passed`, and `evidence` (not `name`/`met`/`details` or other variants) — the viewer depends on these exact field names. For assertions that can be checked programmatically, write and run a script rather than eyeballing it — scripts are faster, more reliable, and can be reused across iterations.
|
||||
|
||||
2. **Aggregate into benchmark** — run the aggregation script from the skill-creator directory:
|
||||
```bash
|
||||
python -m scripts.aggregate_benchmark <workspace>/iteration-N --skill-name <name>
|
||||
```
|
||||
This produces `benchmark.json` and `benchmark.md` with pass_rate, time, and tokens for each configuration, with mean ± stddev and the delta. If generating benchmark.json manually, see `references/schemas.md` for the exact schema the viewer expects.
|
||||
Put each with_skill version before its baseline counterpart.
|
||||
|
||||
3. **Do an analyst pass** — read the benchmark data and surface patterns the aggregate stats might hide. See `agents/analyzer.md` (the "Analyzing Benchmark Results" section) for what to look for — things like assertions that always pass regardless of skill (non-discriminating), high-variance evals (possibly flaky), and time/token tradeoffs.
|
||||
|
||||
4. **Launch the viewer** with both qualitative outputs and quantitative data:
|
||||
```bash
|
||||
nohup python <skill-creator-path>/eval-viewer/generate_review.py \
|
||||
<workspace>/iteration-N \
|
||||
--skill-name "my-skill" \
|
||||
--benchmark <workspace>/iteration-N/benchmark.json \
|
||||
> /dev/null 2>&1 &
|
||||
VIEWER_PID=$!
|
||||
```
|
||||
For iteration 2+, also pass `--previous-workspace <workspace>/iteration-<N-1>`.
|
||||
|
||||
**Cowork / headless environments:** If `webbrowser.open()` is not available or the environment has no display, use `--static <output_path>` to write a standalone HTML file instead of starting a server. Feedback will be downloaded as a `feedback.json` file when the user clicks "Submit All Reviews". After download, copy `feedback.json` into the workspace directory for the next iteration to pick up.
|
||||
|
||||
Note: please use generate_review.py to create the viewer; there's no need to write custom HTML.
|
||||
|
||||
5. **Tell the user** something like: "I've opened the results in your browser. There are two tabs — 'Outputs' lets you click through each test case and leave feedback, 'Benchmark' shows the quantitative comparison. When you're done, come back here and let me know."
|
||||
|
||||
### What the user sees in the viewer
|
||||
|
||||
The "Outputs" tab shows one test case at a time:
|
||||
- **Prompt**: the task that was given
|
||||
- **Output**: the files the skill produced, rendered inline where possible
|
||||
- **Previous Output** (iteration 2+): collapsed section showing last iteration's output
|
||||
- **Formal Grades** (if grading was run): collapsed section showing assertion pass/fail
|
||||
- **Feedback**: a textbox that auto-saves as they type
|
||||
- **Previous Feedback** (iteration 2+): their comments from last time, shown below the textbox
|
||||
|
||||
The "Benchmark" tab shows the stats summary: pass rates, timing, and token usage for each configuration, with per-eval breakdowns and analyst observations.
|
||||
|
||||
Navigation is via prev/next buttons or arrow keys. When done, they click "Submit All Reviews" which saves all feedback to `feedback.json`.
|
||||
|
||||
### Step 5: Read the feedback
|
||||
|
||||
When the user tells you they're done, read `feedback.json`:
|
||||
|
||||
```json
|
||||
{
|
||||
"reviews": [
|
||||
{"run_id": "eval-0-with_skill", "feedback": "the chart is missing axis labels", "timestamp": "..."},
|
||||
{"run_id": "eval-1-with_skill", "feedback": "", "timestamp": "..."},
|
||||
{"run_id": "eval-2-with_skill", "feedback": "perfect, love this", "timestamp": "..."}
|
||||
],
|
||||
"status": "complete"
|
||||
}
|
||||
```
|
||||
|
||||
Empty feedback means the user thought it was fine. Focus your improvements on the test cases where the user had specific complaints.
|
||||
|
||||
Kill the viewer server when you're done with it:
|
||||
|
||||
```bash
|
||||
kill $VIEWER_PID 2>/dev/null
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Improving the skill
|
||||
|
||||
This is the heart of the loop. You've run the test cases, the user has reviewed the results, and now you need to make the skill better based on their feedback.
|
||||
|
||||
### How to think about improvements
|
||||
|
||||
1. **Generalize from the feedback.** The big picture thing that's happening here is that we're trying to create skills that can be used a million times (maybe literally, maybe even more who knows) across many different prompts. Here you and the user are iterating on only a few examples over and over again because it helps move faster. The user knows these examples in and out and it's quick for them to assess new outputs. But if the skill you and the user are codeveloping works only for those examples, it's useless. Rather than put in fiddly overfitty changes, or oppressively constrictive MUSTs, if there's some stubborn issue, you might try branching out and using different metaphors, or recommending different patterns of working. It's relatively cheap to try and maybe you'll land on something great.
|
||||
|
||||
2. **Keep the prompt lean.** Remove things that aren't pulling their weight. Make sure to read the transcripts, not just the final outputs — if it looks like the skill is making the model waste a bunch of time doing things that are unproductive, you can try getting rid of the parts of the skill that are making it do that and seeing what happens.
|
||||
|
||||
3. **Explain the why.** Try hard to explain the **why** behind everything you're asking the model to do. Today's LLMs are *smart*. They have good theory of mind and when given a good harness can go beyond rote instructions and really make things happen. Even if the feedback from the user is terse or frustrated, try to actually understand the task and why the user is writing what they wrote, and what they actually wrote, and then transmit this understanding into the instructions. If you find yourself writing ALWAYS or NEVER in all caps, or using super rigid structures, that's a yellow flag — if possible, reframe and explain the reasoning so that the model understands why the thing you're asking for is important. That's a more humane, powerful, and effective approach.
|
||||
|
||||
4. **Look for repeated work across test cases.** Read the transcripts from the test runs and notice if the subagents all independently wrote similar helper scripts or took the same multi-step approach to something. If all 3 test cases resulted in the subagent writing a `create_docx.py` or a `build_chart.py`, that's a strong signal the skill should bundle that script. Write it once, put it in `scripts/`, and tell the skill to use it. This saves every future invocation from reinventing the wheel.
|
||||
|
||||
This task is pretty important (we are trying to create billions a year in economic value here!) and your thinking time is not the blocker; take your time and really mull things over. I'd suggest writing a draft revision and then looking at it anew and making improvements. Really do your best to get into the head of the user and understand what they want and need.
|
||||
|
||||
### The iteration loop
|
||||
|
||||
After improving the skill:
|
||||
|
||||
1. Apply your improvements to the skill
|
||||
2. Rerun all test cases into a new `iteration-<N+1>/` directory, including baseline runs. If you're creating a new skill, the baseline is always `without_skill` (no skill) — that stays the same across iterations. If you're improving an existing skill, use your judgment on what makes sense as the baseline: the original version the user came in with, or the previous iteration.
|
||||
3. Launch the reviewer with `--previous-workspace` pointing at the previous iteration
|
||||
4. Wait for the user to review and tell you they're done
|
||||
5. Read the new feedback, improve again, repeat
|
||||
|
||||
Keep going until:
|
||||
- The user says they're happy
|
||||
- The feedback is all empty (everything looks good)
|
||||
- You're not making meaningful progress
|
||||
|
||||
---
|
||||
|
||||
## Advanced: Blind comparison
|
||||
|
||||
For situations where you want a more rigorous comparison between two versions of a skill (e.g., the user asks "is the new version actually better?"), there's a blind comparison system. Read `agents/comparator.md` and `agents/analyzer.md` for the details. The basic idea is: give two outputs to an independent agent without telling it which is which, and let it judge quality. Then analyze why the winner won.
|
||||
|
||||
This is optional, requires subagents, and most users won't need it. The human review loop is usually sufficient.
|
||||
|
||||
---
|
||||
|
||||
## Description Optimization
|
||||
|
||||
The description field in SKILL.md frontmatter is the primary mechanism that determines whether Claude invokes a skill. After creating or improving a skill, offer to optimize the description for better triggering accuracy.
|
||||
|
||||
### Step 1: Generate trigger eval queries
|
||||
|
||||
Create 20 eval queries — a mix of should-trigger and should-not-trigger. Save as JSON:
|
||||
|
||||
```json
|
||||
[
|
||||
{"query": "the user prompt", "should_trigger": true},
|
||||
{"query": "another prompt", "should_trigger": false}
|
||||
]
|
||||
```
|
||||
|
||||
The queries must be realistic and something a Claude Code or Claude.ai user would actually type. Not abstract requests, but requests that are concrete and specific and have a good amount of detail. For instance, file paths, personal context about the user's job or situation, column names and values, company names, URLs. A little bit of backstory. Some might be in lowercase or contain abbreviations or typos or casual speech. Use a mix of different lengths, and focus on edge cases rather than making them clear-cut (the user will get a chance to sign off on them).
|
||||
|
||||
Bad: `"Format this data"`, `"Extract text from PDF"`, `"Create a chart"`
|
||||
|
||||
Good: `"ok so my boss just sent me this xlsx file (its in my downloads, called something like 'Q4 sales final FINAL v2.xlsx') and she wants me to add a column that shows the profit margin as a percentage. The revenue is in column C and costs are in column D i think"`
|
||||
|
||||
For the **should-trigger** queries (8-10), think about coverage. You want different phrasings of the same intent — some formal, some casual. Include cases where the user doesn't explicitly name the skill or file type but clearly needs it. Throw in some uncommon use cases and cases where this skill competes with another but should win.
|
||||
|
||||
For the **should-not-trigger** queries (8-10), the most valuable ones are the near-misses — queries that share keywords or concepts with the skill but actually need something different. Think adjacent domains, ambiguous phrasing where a naive keyword match would trigger but shouldn't, and cases where the query touches on something the skill does but in a context where another tool is more appropriate.
|
||||
|
||||
The key thing to avoid: don't make should-not-trigger queries obviously irrelevant. "Write a fibonacci function" as a negative test for a PDF skill is too easy — it doesn't test anything. The negative cases should be genuinely tricky.
|
||||
|
||||
### Step 2: Review with user
|
||||
|
||||
Present the eval set to the user for review using the HTML template:
|
||||
|
||||
1. Read the template from `assets/eval_review.html`
|
||||
2. Replace the placeholders:
|
||||
- `__EVAL_DATA_PLACEHOLDER__` → the JSON array of eval items (no quotes around it — it's a JS variable assignment)
|
||||
- `__SKILL_NAME_PLACEHOLDER__` → the skill's name
|
||||
- `__SKILL_DESCRIPTION_PLACEHOLDER__` → the skill's current description
|
||||
3. Write to a temp file (e.g., `/tmp/eval_review_<skill-name>.html`) and open it: `open /tmp/eval_review_<skill-name>.html`
|
||||
4. The user can edit queries, toggle should-trigger, add/remove entries, then click "Export Eval Set"
|
||||
5. The file downloads to `~/Downloads/eval_set.json` — check the Downloads folder for the most recent version in case there are multiple (e.g., `eval_set (1).json`)
|
||||
|
||||
This step matters — bad eval queries lead to bad descriptions.
|
||||
|
||||
### Step 3: Run the optimization loop
|
||||
|
||||
Tell the user: "This will take some time — I'll run the optimization loop in the background and check on it periodically."
|
||||
|
||||
Save the eval set to the workspace, then run in the background:
|
||||
|
||||
```bash
|
||||
python -m scripts.run_loop \
|
||||
--eval-set <path-to-trigger-eval.json> \
|
||||
--skill-path <path-to-skill> \
|
||||
--model <model-id-powering-this-session> \
|
||||
--max-iterations 5 \
|
||||
--verbose
|
||||
```
|
||||
|
||||
Use the model ID from your system prompt (the one powering the current session) so the triggering test matches what the user actually experiences.
|
||||
|
||||
While it runs, periodically tail the output to give the user updates on which iteration it's on and what the scores look like.
|
||||
|
||||
This handles the full optimization loop automatically. It splits the eval set into 60% train and 40% held-out test, evaluates the current description (running each query 3 times to get a reliable trigger rate), then calls Claude with extended thinking to propose improvements based on what failed. It re-evaluates each new description on both train and test, iterating up to 5 times. When it's done, it opens an HTML report in the browser showing the results per iteration and returns JSON with `best_description` — selected by test score rather than train score to avoid overfitting.
|
||||
|
||||
### How skill triggering works
|
||||
|
||||
Understanding the triggering mechanism helps design better eval queries. Skills appear in Claude's `available_skills` list with their name + description, and Claude decides whether to consult a skill based on that description. The important thing to know is that Claude only consults skills for tasks it can't easily handle on its own — simple, one-step queries like "read this PDF" may not trigger a skill even if the description matches perfectly, because Claude can handle them directly with basic tools. Complex, multi-step, or specialized queries reliably trigger skills when the description matches.
|
||||
|
||||
This means your eval queries should be substantive enough that Claude would actually benefit from consulting a skill. Simple queries like "read file X" are poor test cases — they won't trigger skills regardless of description quality.
|
||||
|
||||
### Step 4: Apply the result
|
||||
|
||||
Take `best_description` from the JSON output and update the skill's SKILL.md frontmatter. Show the user before/after and report the scores.
|
||||
|
||||
---
|
||||
|
||||
### Package and Present (only if `present_files` tool is available)
|
||||
|
||||
Check whether you have access to the `present_files` tool. If you don't, skip this step. If you do, package the skill and present the .skill file to the user:
|
||||
|
||||
```bash
|
||||
python -m scripts.package_skill <path/to/skill-folder>
|
||||
```
|
||||
|
||||
After packaging, direct the user to the resulting `.skill` file path so they can install it.
|
||||
|
||||
---
|
||||
|
||||
## Claude.ai-specific instructions
|
||||
|
||||
In Claude.ai, the core workflow is the same (draft → test → review → improve → repeat), but because Claude.ai doesn't have subagents, some mechanics change. Here's what to adapt:
|
||||
|
||||
**Running test cases**: No subagents means no parallel execution. For each test case, read the skill's SKILL.md, then follow its instructions to accomplish the test prompt yourself. Do them one at a time. This is less rigorous than independent subagents (you wrote the skill and you're also running it, so you have full context), but it's a useful sanity check — and the human review step compensates. Skip the baseline runs — just use the skill to complete the task as requested.
|
||||
|
||||
**Reviewing results**: If you can't open a browser (e.g., Claude.ai's VM has no display, or you're on a remote server), skip the browser reviewer entirely. Instead, present results directly in the conversation. For each test case, show the prompt and the output. If the output is a file the user needs to see (like a .docx or .xlsx), save it to the filesystem and tell them where it is so they can download and inspect it. Ask for feedback inline: "How does this look? Anything you'd change?"
|
||||
|
||||
**Benchmarking**: Skip the quantitative benchmarking — it relies on baseline comparisons which aren't meaningful without subagents. Focus on qualitative feedback from the user.
|
||||
|
||||
**The iteration loop**: Same as before — improve the skill, rerun the test cases, ask for feedback — just without the browser reviewer in the middle. You can still organize results into iteration directories on the filesystem if you have one.
|
||||
|
||||
**Description optimization**: This section requires the `claude` CLI tool (specifically `claude -p`) which is only available in Claude Code. Skip it if you're on Claude.ai.
|
||||
|
||||
**Blind comparison**: Requires subagents. Skip it.
|
||||
|
||||
**Packaging**: The `package_skill.py` script works anywhere with Python and a filesystem. On Claude.ai, you can run it and the user can download the resulting `.skill` file.
|
||||
|
||||
---
|
||||
|
||||
## Cowork-Specific Instructions
|
||||
|
||||
If you're in Cowork, the main things to know are:
|
||||
|
||||
- You have subagents, so the main workflow (spawn test cases in parallel, run baselines, grade, etc.) all works. (However, if you run into severe problems with timeouts, it's OK to run the test prompts in series rather than parallel.)
|
||||
- You don't have a browser or display, so when generating the eval viewer, use `--static <output_path>` to write a standalone HTML file instead of starting a server. Then proffer a link that the user can click to open the HTML in their browser.
|
||||
- For whatever reason, the Cowork setup seems to disincline Claude from generating the eval viewer after running the tests, so just to reiterate: whether you're in Cowork or in Claude Code, after running tests, you should always generate the eval viewer for the human to look at examples before revising the skill yourself and trying to make corrections, using `generate_review.py` (not writing your own boutique html code). Sorry in advance but I'm gonna go all caps here: GENERATE THE EVAL VIEWER *BEFORE* evaluating inputs yourself. You want to get them in front of the human ASAP!
|
||||
- Feedback works differently: since there's no running server, the viewer's "Submit All Reviews" button will download `feedback.json` as a file. You can then read it from there (you may have to request access first).
|
||||
- Packaging works — `package_skill.py` just needs Python and a filesystem.
|
||||
- Description optimization (`run_loop.py` / `run_eval.py`) should work in Cowork just fine since it uses `claude -p` via subprocess, not a browser, but please save it until you've fully finished making the skill and the user agrees it's in good shape.
|
||||
|
||||
---
|
||||
|
||||
## Reference files
|
||||
|
||||
The agents/ directory contains instructions for specialized subagents. Read them when you need to spawn the relevant subagent.
|
||||
|
||||
- `agents/grader.md` — How to evaluate assertions against outputs
|
||||
- `agents/comparator.md` — How to do blind A/B comparison between two outputs
|
||||
- `agents/analyzer.md` — How to analyze why one version beat another
|
||||
|
||||
The references/ directory has additional documentation:
|
||||
- `references/schemas.md` — JSON structures for evals.json, grading.json, etc.
|
||||
|
||||
---
|
||||
|
||||
Repeating one more time the core loop here for emphasis:
|
||||
|
||||
- Figure out what the skill is about
|
||||
- Draft or edit the skill
|
||||
- Run claude-with-access-to-the-skill on test prompts
|
||||
- With the user, evaluate the outputs:
|
||||
- Create benchmark.json and run `eval-viewer/generate_review.py` to help the user review them
|
||||
- Run quantitative evals
|
||||
- Repeat until you and the user are satisfied
|
||||
- Package the final skill and return it to the user.
|
||||
|
||||
Please add steps to your TodoList, if you have such a thing, to make sure you don't forget. If you're in Cowork, please specifically put "Create evals JSON and run `eval-viewer/generate_review.py` so human can review test cases" in your TodoList to make sure it happens.
|
||||
|
||||
Good luck!
|
||||
274
.skills/skill-creator/agents/analyzer.md
Normal file
274
.skills/skill-creator/agents/analyzer.md
Normal file
|
|
@ -0,0 +1,274 @@
|
|||
# Post-hoc Analyzer Agent
|
||||
|
||||
Analyze blind comparison results to understand WHY the winner won and generate improvement suggestions.
|
||||
|
||||
## Role
|
||||
|
||||
After the blind comparator determines a winner, the Post-hoc Analyzer "unblids" the results by examining the skills and transcripts. The goal is to extract actionable insights: what made the winner better, and how can the loser be improved?
|
||||
|
||||
## Inputs
|
||||
|
||||
You receive these parameters in your prompt:
|
||||
|
||||
- **winner**: "A" or "B" (from blind comparison)
|
||||
- **winner_skill_path**: Path to the skill that produced the winning output
|
||||
- **winner_transcript_path**: Path to the execution transcript for the winner
|
||||
- **loser_skill_path**: Path to the skill that produced the losing output
|
||||
- **loser_transcript_path**: Path to the execution transcript for the loser
|
||||
- **comparison_result_path**: Path to the blind comparator's output JSON
|
||||
- **output_path**: Where to save the analysis results
|
||||
|
||||
## Process
|
||||
|
||||
### Step 1: Read Comparison Result
|
||||
|
||||
1. Read the blind comparator's output at comparison_result_path
|
||||
2. Note the winning side (A or B), the reasoning, and any scores
|
||||
3. Understand what the comparator valued in the winning output
|
||||
|
||||
### Step 2: Read Both Skills
|
||||
|
||||
1. Read the winner skill's SKILL.md and key referenced files
|
||||
2. Read the loser skill's SKILL.md and key referenced files
|
||||
3. Identify structural differences:
|
||||
- Instructions clarity and specificity
|
||||
- Script/tool usage patterns
|
||||
- Example coverage
|
||||
- Edge case handling
|
||||
|
||||
### Step 3: Read Both Transcripts
|
||||
|
||||
1. Read the winner's transcript
|
||||
2. Read the loser's transcript
|
||||
3. Compare execution patterns:
|
||||
- How closely did each follow their skill's instructions?
|
||||
- What tools were used differently?
|
||||
- Where did the loser diverge from optimal behavior?
|
||||
- Did either encounter errors or make recovery attempts?
|
||||
|
||||
### Step 4: Analyze Instruction Following
|
||||
|
||||
For each transcript, evaluate:
|
||||
- Did the agent follow the skill's explicit instructions?
|
||||
- Did the agent use the skill's provided tools/scripts?
|
||||
- Were there missed opportunities to leverage skill content?
|
||||
- Did the agent add unnecessary steps not in the skill?
|
||||
|
||||
Score instruction following 1-10 and note specific issues.
|
||||
|
||||
### Step 5: Identify Winner Strengths
|
||||
|
||||
Determine what made the winner better:
|
||||
- Clearer instructions that led to better behavior?
|
||||
- Better scripts/tools that produced better output?
|
||||
- More comprehensive examples that guided edge cases?
|
||||
- Better error handling guidance?
|
||||
|
||||
Be specific. Quote from skills/transcripts where relevant.
|
||||
|
||||
### Step 6: Identify Loser Weaknesses
|
||||
|
||||
Determine what held the loser back:
|
||||
- Ambiguous instructions that led to suboptimal choices?
|
||||
- Missing tools/scripts that forced workarounds?
|
||||
- Gaps in edge case coverage?
|
||||
- Poor error handling that caused failures?
|
||||
|
||||
### Step 7: Generate Improvement Suggestions
|
||||
|
||||
Based on the analysis, produce actionable suggestions for improving the loser skill:
|
||||
- Specific instruction changes to make
|
||||
- Tools/scripts to add or modify
|
||||
- Examples to include
|
||||
- Edge cases to address
|
||||
|
||||
Prioritize by impact. Focus on changes that would have changed the outcome.
|
||||
|
||||
### Step 8: Write Analysis Results
|
||||
|
||||
Save structured analysis to `{output_path}`.
|
||||
|
||||
## Output Format
|
||||
|
||||
Write a JSON file with this structure:
|
||||
|
||||
```json
|
||||
{
|
||||
"comparison_summary": {
|
||||
"winner": "A",
|
||||
"winner_skill": "path/to/winner/skill",
|
||||
"loser_skill": "path/to/loser/skill",
|
||||
"comparator_reasoning": "Brief summary of why comparator chose winner"
|
||||
},
|
||||
"winner_strengths": [
|
||||
"Clear step-by-step instructions for handling multi-page documents",
|
||||
"Included validation script that caught formatting errors",
|
||||
"Explicit guidance on fallback behavior when OCR fails"
|
||||
],
|
||||
"loser_weaknesses": [
|
||||
"Vague instruction 'process the document appropriately' led to inconsistent behavior",
|
||||
"No script for validation, agent had to improvise and made errors",
|
||||
"No guidance on OCR failure, agent gave up instead of trying alternatives"
|
||||
],
|
||||
"instruction_following": {
|
||||
"winner": {
|
||||
"score": 9,
|
||||
"issues": [
|
||||
"Minor: skipped optional logging step"
|
||||
]
|
||||
},
|
||||
"loser": {
|
||||
"score": 6,
|
||||
"issues": [
|
||||
"Did not use the skill's formatting template",
|
||||
"Invented own approach instead of following step 3",
|
||||
"Missed the 'always validate output' instruction"
|
||||
]
|
||||
}
|
||||
},
|
||||
"improvement_suggestions": [
|
||||
{
|
||||
"priority": "high",
|
||||
"category": "instructions",
|
||||
"suggestion": "Replace 'process the document appropriately' with explicit steps: 1) Extract text, 2) Identify sections, 3) Format per template",
|
||||
"expected_impact": "Would eliminate ambiguity that caused inconsistent behavior"
|
||||
},
|
||||
{
|
||||
"priority": "high",
|
||||
"category": "tools",
|
||||
"suggestion": "Add validate_output.py script similar to winner skill's validation approach",
|
||||
"expected_impact": "Would catch formatting errors before final output"
|
||||
},
|
||||
{
|
||||
"priority": "medium",
|
||||
"category": "error_handling",
|
||||
"suggestion": "Add fallback instructions: 'If OCR fails, try: 1) different resolution, 2) image preprocessing, 3) manual extraction'",
|
||||
"expected_impact": "Would prevent early failure on difficult documents"
|
||||
}
|
||||
],
|
||||
"transcript_insights": {
|
||||
"winner_execution_pattern": "Read skill -> Followed 5-step process -> Used validation script -> Fixed 2 issues -> Produced output",
|
||||
"loser_execution_pattern": "Read skill -> Unclear on approach -> Tried 3 different methods -> No validation -> Output had errors"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Guidelines
|
||||
|
||||
- **Be specific**: Quote from skills and transcripts, don't just say "instructions were unclear"
|
||||
- **Be actionable**: Suggestions should be concrete changes, not vague advice
|
||||
- **Focus on skill improvements**: The goal is to improve the losing skill, not critique the agent
|
||||
- **Prioritize by impact**: Which changes would most likely have changed the outcome?
|
||||
- **Consider causation**: Did the skill weakness actually cause the worse output, or is it incidental?
|
||||
- **Stay objective**: Analyze what happened, don't editorialize
|
||||
- **Think about generalization**: Would this improvement help on other evals too?
|
||||
|
||||
## Categories for Suggestions
|
||||
|
||||
Use these categories to organize improvement suggestions:
|
||||
|
||||
| Category | Description |
|
||||
|----------|-------------|
|
||||
| `instructions` | Changes to the skill's prose instructions |
|
||||
| `tools` | Scripts, templates, or utilities to add/modify |
|
||||
| `examples` | Example inputs/outputs to include |
|
||||
| `error_handling` | Guidance for handling failures |
|
||||
| `structure` | Reorganization of skill content |
|
||||
| `references` | External docs or resources to add |
|
||||
|
||||
## Priority Levels
|
||||
|
||||
- **high**: Would likely change the outcome of this comparison
|
||||
- **medium**: Would improve quality but may not change win/loss
|
||||
- **low**: Nice to have, marginal improvement
|
||||
|
||||
---
|
||||
|
||||
# Analyzing Benchmark Results
|
||||
|
||||
When analyzing benchmark results, the analyzer's purpose is to **surface patterns and anomalies** across multiple runs, not suggest skill improvements.
|
||||
|
||||
## Role
|
||||
|
||||
Review all benchmark run results and generate freeform notes that help the user understand skill performance. Focus on patterns that wouldn't be visible from aggregate metrics alone.
|
||||
|
||||
## Inputs
|
||||
|
||||
You receive these parameters in your prompt:
|
||||
|
||||
- **benchmark_data_path**: Path to the in-progress benchmark.json with all run results
|
||||
- **skill_path**: Path to the skill being benchmarked
|
||||
- **output_path**: Where to save the notes (as JSON array of strings)
|
||||
|
||||
## Process
|
||||
|
||||
### Step 1: Read Benchmark Data
|
||||
|
||||
1. Read the benchmark.json containing all run results
|
||||
2. Note the configurations tested (with_skill, without_skill)
|
||||
3. Understand the run_summary aggregates already calculated
|
||||
|
||||
### Step 2: Analyze Per-Assertion Patterns
|
||||
|
||||
For each expectation across all runs:
|
||||
- Does it **always pass** in both configurations? (may not differentiate skill value)
|
||||
- Does it **always fail** in both configurations? (may be broken or beyond capability)
|
||||
- Does it **always pass with skill but fail without**? (skill clearly adds value here)
|
||||
- Does it **always fail with skill but pass without**? (skill may be hurting)
|
||||
- Is it **highly variable**? (flaky expectation or non-deterministic behavior)
|
||||
|
||||
### Step 3: Analyze Cross-Eval Patterns
|
||||
|
||||
Look for patterns across evals:
|
||||
- Are certain eval types consistently harder/easier?
|
||||
- Do some evals show high variance while others are stable?
|
||||
- Are there surprising results that contradict expectations?
|
||||
|
||||
### Step 4: Analyze Metrics Patterns
|
||||
|
||||
Look at time_seconds, tokens, tool_calls:
|
||||
- Does the skill significantly increase execution time?
|
||||
- Is there high variance in resource usage?
|
||||
- Are there outlier runs that skew the aggregates?
|
||||
|
||||
### Step 5: Generate Notes
|
||||
|
||||
Write freeform observations as a list of strings. Each note should:
|
||||
- State a specific observation
|
||||
- Be grounded in the data (not speculation)
|
||||
- Help the user understand something the aggregate metrics don't show
|
||||
|
||||
Examples:
|
||||
- "Assertion 'Output is a PDF file' passes 100% in both configurations - may not differentiate skill value"
|
||||
- "Eval 3 shows high variance (50% ± 40%) - run 2 had an unusual failure that may be flaky"
|
||||
- "Without-skill runs consistently fail on table extraction expectations (0% pass rate)"
|
||||
- "Skill adds 13s average execution time but improves pass rate by 50%"
|
||||
- "Token usage is 80% higher with skill, primarily due to script output parsing"
|
||||
- "All 3 without-skill runs for eval 1 produced empty output"
|
||||
|
||||
### Step 6: Write Notes
|
||||
|
||||
Save notes to `{output_path}` as a JSON array of strings:
|
||||
|
||||
```json
|
||||
[
|
||||
"Assertion 'Output is a PDF file' passes 100% in both configurations - may not differentiate skill value",
|
||||
"Eval 3 shows high variance (50% ± 40%) - run 2 had an unusual failure",
|
||||
"Without-skill runs consistently fail on table extraction expectations",
|
||||
"Skill adds 13s average execution time but improves pass rate by 50%"
|
||||
]
|
||||
```
|
||||
|
||||
## Guidelines
|
||||
|
||||
**DO:**
|
||||
- Report what you observe in the data
|
||||
- Be specific about which evals, expectations, or runs you're referring to
|
||||
- Note patterns that aggregate metrics would hide
|
||||
- Provide context that helps interpret the numbers
|
||||
|
||||
**DO NOT:**
|
||||
- Suggest improvements to the skill (that's for the improvement step, not benchmarking)
|
||||
- Make subjective quality judgments ("the output was good/bad")
|
||||
- Speculate about causes without evidence
|
||||
- Repeat information already in the run_summary aggregates
|
||||
202
.skills/skill-creator/agents/comparator.md
Normal file
202
.skills/skill-creator/agents/comparator.md
Normal file
|
|
@ -0,0 +1,202 @@
|
|||
# Blind Comparator Agent
|
||||
|
||||
Compare two outputs WITHOUT knowing which skill produced them.
|
||||
|
||||
## Role
|
||||
|
||||
The Blind Comparator judges which output better accomplishes the eval task. You receive two outputs labeled A and B, but you do NOT know which skill produced which. This prevents bias toward a particular skill or approach.
|
||||
|
||||
Your judgment is based purely on output quality and task completion.
|
||||
|
||||
## Inputs
|
||||
|
||||
You receive these parameters in your prompt:
|
||||
|
||||
- **output_a_path**: Path to the first output file or directory
|
||||
- **output_b_path**: Path to the second output file or directory
|
||||
- **eval_prompt**: The original task/prompt that was executed
|
||||
- **expectations**: List of expectations to check (optional - may be empty)
|
||||
|
||||
## Process
|
||||
|
||||
### Step 1: Read Both Outputs
|
||||
|
||||
1. Examine output A (file or directory)
|
||||
2. Examine output B (file or directory)
|
||||
3. Note the type, structure, and content of each
|
||||
4. If outputs are directories, examine all relevant files inside
|
||||
|
||||
### Step 2: Understand the Task
|
||||
|
||||
1. Read the eval_prompt carefully
|
||||
2. Identify what the task requires:
|
||||
- What should be produced?
|
||||
- What qualities matter (accuracy, completeness, format)?
|
||||
- What would distinguish a good output from a poor one?
|
||||
|
||||
### Step 3: Generate Evaluation Rubric
|
||||
|
||||
Based on the task, generate a rubric with two dimensions:
|
||||
|
||||
**Content Rubric** (what the output contains):
|
||||
| Criterion | 1 (Poor) | 3 (Acceptable) | 5 (Excellent) |
|
||||
|-----------|----------|----------------|---------------|
|
||||
| Correctness | Major errors | Minor errors | Fully correct |
|
||||
| Completeness | Missing key elements | Mostly complete | All elements present |
|
||||
| Accuracy | Significant inaccuracies | Minor inaccuracies | Accurate throughout |
|
||||
|
||||
**Structure Rubric** (how the output is organized):
|
||||
| Criterion | 1 (Poor) | 3 (Acceptable) | 5 (Excellent) |
|
||||
|-----------|----------|----------------|---------------|
|
||||
| Organization | Disorganized | Reasonably organized | Clear, logical structure |
|
||||
| Formatting | Inconsistent/broken | Mostly consistent | Professional, polished |
|
||||
| Usability | Difficult to use | Usable with effort | Easy to use |
|
||||
|
||||
Adapt criteria to the specific task. For example:
|
||||
- PDF form → "Field alignment", "Text readability", "Data placement"
|
||||
- Document → "Section structure", "Heading hierarchy", "Paragraph flow"
|
||||
- Data output → "Schema correctness", "Data types", "Completeness"
|
||||
|
||||
### Step 4: Evaluate Each Output Against the Rubric
|
||||
|
||||
For each output (A and B):
|
||||
|
||||
1. **Score each criterion** on the rubric (1-5 scale)
|
||||
2. **Calculate dimension totals**: Content score, Structure score
|
||||
3. **Calculate overall score**: Average of dimension scores, scaled to 1-10
|
||||
|
||||
### Step 5: Check Assertions (if provided)
|
||||
|
||||
If expectations are provided:
|
||||
|
||||
1. Check each expectation against output A
|
||||
2. Check each expectation against output B
|
||||
3. Count pass rates for each output
|
||||
4. Use expectation scores as secondary evidence (not the primary decision factor)
|
||||
|
||||
### Step 6: Determine the Winner
|
||||
|
||||
Compare A and B based on (in priority order):
|
||||
|
||||
1. **Primary**: Overall rubric score (content + structure)
|
||||
2. **Secondary**: Assertion pass rates (if applicable)
|
||||
3. **Tiebreaker**: If truly equal, declare a TIE
|
||||
|
||||
Be decisive - ties should be rare. One output is usually better, even if marginally.
|
||||
|
||||
### Step 7: Write Comparison Results
|
||||
|
||||
Save results to a JSON file at the path specified (or `comparison.json` if not specified).
|
||||
|
||||
## Output Format
|
||||
|
||||
Write a JSON file with this structure:
|
||||
|
||||
```json
|
||||
{
|
||||
"winner": "A",
|
||||
"reasoning": "Output A provides a complete solution with proper formatting and all required fields. Output B is missing the date field and has formatting inconsistencies.",
|
||||
"rubric": {
|
||||
"A": {
|
||||
"content": {
|
||||
"correctness": 5,
|
||||
"completeness": 5,
|
||||
"accuracy": 4
|
||||
},
|
||||
"structure": {
|
||||
"organization": 4,
|
||||
"formatting": 5,
|
||||
"usability": 4
|
||||
},
|
||||
"content_score": 4.7,
|
||||
"structure_score": 4.3,
|
||||
"overall_score": 9.0
|
||||
},
|
||||
"B": {
|
||||
"content": {
|
||||
"correctness": 3,
|
||||
"completeness": 2,
|
||||
"accuracy": 3
|
||||
},
|
||||
"structure": {
|
||||
"organization": 3,
|
||||
"formatting": 2,
|
||||
"usability": 3
|
||||
},
|
||||
"content_score": 2.7,
|
||||
"structure_score": 2.7,
|
||||
"overall_score": 5.4
|
||||
}
|
||||
},
|
||||
"output_quality": {
|
||||
"A": {
|
||||
"score": 9,
|
||||
"strengths": ["Complete solution", "Well-formatted", "All fields present"],
|
||||
"weaknesses": ["Minor style inconsistency in header"]
|
||||
},
|
||||
"B": {
|
||||
"score": 5,
|
||||
"strengths": ["Readable output", "Correct basic structure"],
|
||||
"weaknesses": ["Missing date field", "Formatting inconsistencies", "Partial data extraction"]
|
||||
}
|
||||
},
|
||||
"expectation_results": {
|
||||
"A": {
|
||||
"passed": 4,
|
||||
"total": 5,
|
||||
"pass_rate": 0.80,
|
||||
"details": [
|
||||
{"text": "Output includes name", "passed": true},
|
||||
{"text": "Output includes date", "passed": true},
|
||||
{"text": "Format is PDF", "passed": true},
|
||||
{"text": "Contains signature", "passed": false},
|
||||
{"text": "Readable text", "passed": true}
|
||||
]
|
||||
},
|
||||
"B": {
|
||||
"passed": 3,
|
||||
"total": 5,
|
||||
"pass_rate": 0.60,
|
||||
"details": [
|
||||
{"text": "Output includes name", "passed": true},
|
||||
{"text": "Output includes date", "passed": false},
|
||||
{"text": "Format is PDF", "passed": true},
|
||||
{"text": "Contains signature", "passed": false},
|
||||
{"text": "Readable text", "passed": true}
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
If no expectations were provided, omit the `expectation_results` field entirely.
|
||||
|
||||
## Field Descriptions
|
||||
|
||||
- **winner**: "A", "B", or "TIE"
|
||||
- **reasoning**: Clear explanation of why the winner was chosen (or why it's a tie)
|
||||
- **rubric**: Structured rubric evaluation for each output
|
||||
- **content**: Scores for content criteria (correctness, completeness, accuracy)
|
||||
- **structure**: Scores for structure criteria (organization, formatting, usability)
|
||||
- **content_score**: Average of content criteria (1-5)
|
||||
- **structure_score**: Average of structure criteria (1-5)
|
||||
- **overall_score**: Combined score scaled to 1-10
|
||||
- **output_quality**: Summary quality assessment
|
||||
- **score**: 1-10 rating (should match rubric overall_score)
|
||||
- **strengths**: List of positive aspects
|
||||
- **weaknesses**: List of issues or shortcomings
|
||||
- **expectation_results**: (Only if expectations provided)
|
||||
- **passed**: Number of expectations that passed
|
||||
- **total**: Total number of expectations
|
||||
- **pass_rate**: Fraction passed (0.0 to 1.0)
|
||||
- **details**: Individual expectation results
|
||||
|
||||
## Guidelines
|
||||
|
||||
- **Stay blind**: DO NOT try to infer which skill produced which output. Judge purely on output quality.
|
||||
- **Be specific**: Cite specific examples when explaining strengths and weaknesses.
|
||||
- **Be decisive**: Choose a winner unless outputs are genuinely equivalent.
|
||||
- **Output quality first**: Assertion scores are secondary to overall task completion.
|
||||
- **Be objective**: Don't favor outputs based on style preferences; focus on correctness and completeness.
|
||||
- **Explain your reasoning**: The reasoning field should make it clear why you chose the winner.
|
||||
- **Handle edge cases**: If both outputs fail, pick the one that fails less badly. If both are excellent, pick the one that's marginally better.
|
||||
223
.skills/skill-creator/agents/grader.md
Normal file
223
.skills/skill-creator/agents/grader.md
Normal file
|
|
@ -0,0 +1,223 @@
|
|||
# Grader Agent
|
||||
|
||||
Evaluate expectations against an execution transcript and outputs.
|
||||
|
||||
## Role
|
||||
|
||||
The Grader reviews a transcript and output files, then determines whether each expectation passes or fails. Provide clear evidence for each judgment.
|
||||
|
||||
You have two jobs: grade the outputs, and critique the evals themselves. A passing grade on a weak assertion is worse than useless — it creates false confidence. When you notice an assertion that's trivially satisfied, or an important outcome that no assertion checks, say so.
|
||||
|
||||
## Inputs
|
||||
|
||||
You receive these parameters in your prompt:
|
||||
|
||||
- **expectations**: List of expectations to evaluate (strings)
|
||||
- **transcript_path**: Path to the execution transcript (markdown file)
|
||||
- **outputs_dir**: Directory containing output files from execution
|
||||
|
||||
## Process
|
||||
|
||||
### Step 1: Read the Transcript
|
||||
|
||||
1. Read the transcript file completely
|
||||
2. Note the eval prompt, execution steps, and final result
|
||||
3. Identify any issues or errors documented
|
||||
|
||||
### Step 2: Examine Output Files
|
||||
|
||||
1. List files in outputs_dir
|
||||
2. Read/examine each file relevant to the expectations. If outputs aren't plain text, use the inspection tools provided in your prompt — don't rely solely on what the transcript says the executor produced.
|
||||
3. Note contents, structure, and quality
|
||||
|
||||
### Step 3: Evaluate Each Assertion
|
||||
|
||||
For each expectation:
|
||||
|
||||
1. **Search for evidence** in the transcript and outputs
|
||||
2. **Determine verdict**:
|
||||
- **PASS**: Clear evidence the expectation is true AND the evidence reflects genuine task completion, not just surface-level compliance
|
||||
- **FAIL**: No evidence, or evidence contradicts the expectation, or the evidence is superficial (e.g., correct filename but empty/wrong content)
|
||||
3. **Cite the evidence**: Quote the specific text or describe what you found
|
||||
|
||||
### Step 4: Extract and Verify Claims
|
||||
|
||||
Beyond the predefined expectations, extract implicit claims from the outputs and verify them:
|
||||
|
||||
1. **Extract claims** from the transcript and outputs:
|
||||
- Factual statements ("The form has 12 fields")
|
||||
- Process claims ("Used pypdf to fill the form")
|
||||
- Quality claims ("All fields were filled correctly")
|
||||
|
||||
2. **Verify each claim**:
|
||||
- **Factual claims**: Can be checked against the outputs or external sources
|
||||
- **Process claims**: Can be verified from the transcript
|
||||
- **Quality claims**: Evaluate whether the claim is justified
|
||||
|
||||
3. **Flag unverifiable claims**: Note claims that cannot be verified with available information
|
||||
|
||||
This catches issues that predefined expectations might miss.
|
||||
|
||||
### Step 5: Read User Notes
|
||||
|
||||
If `{outputs_dir}/user_notes.md` exists:
|
||||
1. Read it and note any uncertainties or issues flagged by the executor
|
||||
2. Include relevant concerns in the grading output
|
||||
3. These may reveal problems even when expectations pass
|
||||
|
||||
### Step 6: Critique the Evals
|
||||
|
||||
After grading, consider whether the evals themselves could be improved. Only surface suggestions when there's a clear gap.
|
||||
|
||||
Good suggestions test meaningful outcomes — assertions that are hard to satisfy without actually doing the work correctly. Think about what makes an assertion *discriminating*: it passes when the skill genuinely succeeds and fails when it doesn't.
|
||||
|
||||
Suggestions worth raising:
|
||||
- An assertion that passed but would also pass for a clearly wrong output (e.g., checking filename existence but not file content)
|
||||
- An important outcome you observed — good or bad — that no assertion covers at all
|
||||
- An assertion that can't actually be verified from the available outputs
|
||||
|
||||
Keep the bar high. The goal is to flag things the eval author would say "good catch" about, not to nitpick every assertion.
|
||||
|
||||
### Step 7: Write Grading Results
|
||||
|
||||
Save results to `{outputs_dir}/../grading.json` (sibling to outputs_dir).
|
||||
|
||||
## Grading Criteria
|
||||
|
||||
**PASS when**:
|
||||
- The transcript or outputs clearly demonstrate the expectation is true
|
||||
- Specific evidence can be cited
|
||||
- The evidence reflects genuine substance, not just surface compliance (e.g., a file exists AND contains correct content, not just the right filename)
|
||||
|
||||
**FAIL when**:
|
||||
- No evidence found for the expectation
|
||||
- Evidence contradicts the expectation
|
||||
- The expectation cannot be verified from available information
|
||||
- The evidence is superficial — the assertion is technically satisfied but the underlying task outcome is wrong or incomplete
|
||||
- The output appears to meet the assertion by coincidence rather than by actually doing the work
|
||||
|
||||
**When uncertain**: The burden of proof to pass is on the expectation.
|
||||
|
||||
### Step 8: Read Executor Metrics and Timing
|
||||
|
||||
1. If `{outputs_dir}/metrics.json` exists, read it and include in grading output
|
||||
2. If `{outputs_dir}/../timing.json` exists, read it and include timing data
|
||||
|
||||
## Output Format
|
||||
|
||||
Write a JSON file with this structure:
|
||||
|
||||
```json
|
||||
{
|
||||
"expectations": [
|
||||
{
|
||||
"text": "The output includes the name 'John Smith'",
|
||||
"passed": true,
|
||||
"evidence": "Found in transcript Step 3: 'Extracted names: John Smith, Sarah Johnson'"
|
||||
},
|
||||
{
|
||||
"text": "The spreadsheet has a SUM formula in cell B10",
|
||||
"passed": false,
|
||||
"evidence": "No spreadsheet was created. The output was a text file."
|
||||
},
|
||||
{
|
||||
"text": "The assistant used the skill's OCR script",
|
||||
"passed": true,
|
||||
"evidence": "Transcript Step 2 shows: 'Tool: Bash - python ocr_script.py image.png'"
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"passed": 2,
|
||||
"failed": 1,
|
||||
"total": 3,
|
||||
"pass_rate": 0.67
|
||||
},
|
||||
"execution_metrics": {
|
||||
"tool_calls": {
|
||||
"Read": 5,
|
||||
"Write": 2,
|
||||
"Bash": 8
|
||||
},
|
||||
"total_tool_calls": 15,
|
||||
"total_steps": 6,
|
||||
"errors_encountered": 0,
|
||||
"output_chars": 12450,
|
||||
"transcript_chars": 3200
|
||||
},
|
||||
"timing": {
|
||||
"executor_duration_seconds": 165.0,
|
||||
"grader_duration_seconds": 26.0,
|
||||
"total_duration_seconds": 191.0
|
||||
},
|
||||
"claims": [
|
||||
{
|
||||
"claim": "The form has 12 fillable fields",
|
||||
"type": "factual",
|
||||
"verified": true,
|
||||
"evidence": "Counted 12 fields in field_info.json"
|
||||
},
|
||||
{
|
||||
"claim": "All required fields were populated",
|
||||
"type": "quality",
|
||||
"verified": false,
|
||||
"evidence": "Reference section was left blank despite data being available"
|
||||
}
|
||||
],
|
||||
"user_notes_summary": {
|
||||
"uncertainties": ["Used 2023 data, may be stale"],
|
||||
"needs_review": [],
|
||||
"workarounds": ["Fell back to text overlay for non-fillable fields"]
|
||||
},
|
||||
"eval_feedback": {
|
||||
"suggestions": [
|
||||
{
|
||||
"assertion": "The output includes the name 'John Smith'",
|
||||
"reason": "A hallucinated document that mentions the name would also pass — consider checking it appears as the primary contact with matching phone and email from the input"
|
||||
},
|
||||
{
|
||||
"reason": "No assertion checks whether the extracted phone numbers match the input — I observed incorrect numbers in the output that went uncaught"
|
||||
}
|
||||
],
|
||||
"overall": "Assertions check presence but not correctness. Consider adding content verification."
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Field Descriptions
|
||||
|
||||
- **expectations**: Array of graded expectations
|
||||
- **text**: The original expectation text
|
||||
- **passed**: Boolean - true if expectation passes
|
||||
- **evidence**: Specific quote or description supporting the verdict
|
||||
- **summary**: Aggregate statistics
|
||||
- **passed**: Count of passed expectations
|
||||
- **failed**: Count of failed expectations
|
||||
- **total**: Total expectations evaluated
|
||||
- **pass_rate**: Fraction passed (0.0 to 1.0)
|
||||
- **execution_metrics**: Copied from executor's metrics.json (if available)
|
||||
- **output_chars**: Total character count of output files (proxy for tokens)
|
||||
- **transcript_chars**: Character count of transcript
|
||||
- **timing**: Wall clock timing from timing.json (if available)
|
||||
- **executor_duration_seconds**: Time spent in executor subagent
|
||||
- **total_duration_seconds**: Total elapsed time for the run
|
||||
- **claims**: Extracted and verified claims from the output
|
||||
- **claim**: The statement being verified
|
||||
- **type**: "factual", "process", or "quality"
|
||||
- **verified**: Boolean - whether the claim holds
|
||||
- **evidence**: Supporting or contradicting evidence
|
||||
- **user_notes_summary**: Issues flagged by the executor
|
||||
- **uncertainties**: Things the executor wasn't sure about
|
||||
- **needs_review**: Items requiring human attention
|
||||
- **workarounds**: Places where the skill didn't work as expected
|
||||
- **eval_feedback**: Improvement suggestions for the evals (only when warranted)
|
||||
- **suggestions**: List of concrete suggestions, each with a `reason` and optionally an `assertion` it relates to
|
||||
- **overall**: Brief assessment — can be "No suggestions, evals look solid" if nothing to flag
|
||||
|
||||
## Guidelines
|
||||
|
||||
- **Be objective**: Base verdicts on evidence, not assumptions
|
||||
- **Be specific**: Quote the exact text that supports your verdict
|
||||
- **Be thorough**: Check both transcript and output files
|
||||
- **Be consistent**: Apply the same standard to each expectation
|
||||
- **Explain failures**: Make it clear why evidence was insufficient
|
||||
- **No partial credit**: Each expectation is pass or fail, not partial
|
||||
146
.skills/skill-creator/assets/eval_review.html
Normal file
146
.skills/skill-creator/assets/eval_review.html
Normal file
|
|
@ -0,0 +1,146 @@
|
|||
<!DOCTYPE html>
|
||||
<html lang="en">
|
||||
<head>
|
||||
<meta charset="UTF-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
||||
<title>Eval Set Review - __SKILL_NAME_PLACEHOLDER__</title>
|
||||
<link rel="preconnect" href="https://fonts.googleapis.com">
|
||||
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
|
||||
<link href="https://fonts.googleapis.com/css2?family=Poppins:wght@500;600&family=Lora:wght@400;500&display=swap" rel="stylesheet">
|
||||
<style>
|
||||
* { box-sizing: border-box; margin: 0; padding: 0; }
|
||||
body { font-family: 'Lora', Georgia, serif; background: #faf9f5; padding: 2rem; color: #141413; }
|
||||
h1 { font-family: 'Poppins', sans-serif; margin-bottom: 0.5rem; font-size: 1.5rem; }
|
||||
.description { color: #b0aea5; margin-bottom: 1.5rem; font-style: italic; max-width: 900px; }
|
||||
.controls { margin-bottom: 1rem; display: flex; gap: 0.5rem; }
|
||||
.btn { font-family: 'Poppins', sans-serif; padding: 0.5rem 1rem; border: none; border-radius: 6px; cursor: pointer; font-size: 0.875rem; font-weight: 500; }
|
||||
.btn-add { background: #6a9bcc; color: white; }
|
||||
.btn-add:hover { background: #5889b8; }
|
||||
.btn-export { background: #d97757; color: white; }
|
||||
.btn-export:hover { background: #c4613f; }
|
||||
table { width: 100%; max-width: 1100px; border-collapse: collapse; background: white; border-radius: 6px; overflow: hidden; box-shadow: 0 1px 3px rgba(0,0,0,0.08); }
|
||||
th { font-family: 'Poppins', sans-serif; background: #141413; color: #faf9f5; padding: 0.75rem 1rem; text-align: left; font-size: 0.875rem; }
|
||||
td { padding: 0.75rem 1rem; border-bottom: 1px solid #e8e6dc; vertical-align: top; }
|
||||
tr:nth-child(even) td { background: #faf9f5; }
|
||||
tr:hover td { background: #f3f1ea; }
|
||||
.section-header td { background: #e8e6dc; font-family: 'Poppins', sans-serif; font-weight: 500; font-size: 0.8rem; color: #141413; text-transform: uppercase; letter-spacing: 0.05em; }
|
||||
.query-input { width: 100%; padding: 0.4rem; border: 1px solid #e8e6dc; border-radius: 4px; font-size: 0.875rem; font-family: 'Lora', Georgia, serif; resize: vertical; min-height: 60px; }
|
||||
.query-input:focus { outline: none; border-color: #d97757; box-shadow: 0 0 0 2px rgba(217,119,87,0.15); }
|
||||
.toggle { position: relative; display: inline-block; width: 44px; height: 24px; }
|
||||
.toggle input { opacity: 0; width: 0; height: 0; }
|
||||
.toggle .slider { position: absolute; inset: 0; background: #b0aea5; border-radius: 24px; cursor: pointer; transition: 0.2s; }
|
||||
.toggle .slider::before { content: ""; position: absolute; width: 18px; height: 18px; left: 3px; bottom: 3px; background: white; border-radius: 50%; transition: 0.2s; }
|
||||
.toggle input:checked + .slider { background: #d97757; }
|
||||
.toggle input:checked + .slider::before { transform: translateX(20px); }
|
||||
.btn-delete { background: #c44; color: white; padding: 0.3rem 0.6rem; border: none; border-radius: 4px; cursor: pointer; font-size: 0.75rem; font-family: 'Poppins', sans-serif; }
|
||||
.btn-delete:hover { background: #a33; }
|
||||
.summary { margin-top: 1rem; color: #b0aea5; font-size: 0.875rem; }
|
||||
</style>
|
||||
</head>
|
||||
<body>
|
||||
<h1>Eval Set Review: <span id="skill-name">__SKILL_NAME_PLACEHOLDER__</span></h1>
|
||||
<p class="description">Current description: <span id="skill-desc">__SKILL_DESCRIPTION_PLACEHOLDER__</span></p>
|
||||
|
||||
<div class="controls">
|
||||
<button class="btn btn-add" onclick="addRow()">+ Add Query</button>
|
||||
<button class="btn btn-export" onclick="exportEvalSet()">Export Eval Set</button>
|
||||
</div>
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr>
|
||||
<th style="width:65%">Query</th>
|
||||
<th style="width:18%">Should Trigger</th>
|
||||
<th style="width:10%">Actions</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody id="eval-body"></tbody>
|
||||
</table>
|
||||
|
||||
<p class="summary" id="summary"></p>
|
||||
|
||||
<script>
|
||||
const EVAL_DATA = __EVAL_DATA_PLACEHOLDER__;
|
||||
|
||||
let evalItems = [...EVAL_DATA];
|
||||
|
||||
function render() {
|
||||
const tbody = document.getElementById('eval-body');
|
||||
tbody.innerHTML = '';
|
||||
|
||||
// Sort: should-trigger first, then should-not-trigger
|
||||
const sorted = evalItems
|
||||
.map((item, origIdx) => ({ ...item, origIdx }))
|
||||
.sort((a, b) => (b.should_trigger ? 1 : 0) - (a.should_trigger ? 1 : 0));
|
||||
|
||||
let lastGroup = null;
|
||||
sorted.forEach(item => {
|
||||
const group = item.should_trigger ? 'trigger' : 'no-trigger';
|
||||
if (group !== lastGroup) {
|
||||
const headerRow = document.createElement('tr');
|
||||
headerRow.className = 'section-header';
|
||||
headerRow.innerHTML = `<td colspan="3">${item.should_trigger ? 'Should Trigger' : 'Should NOT Trigger'}</td>`;
|
||||
tbody.appendChild(headerRow);
|
||||
lastGroup = group;
|
||||
}
|
||||
|
||||
const idx = item.origIdx;
|
||||
const tr = document.createElement('tr');
|
||||
tr.innerHTML = `
|
||||
<td><textarea class="query-input" onchange="updateQuery(${idx}, this.value)">${escapeHtml(item.query)}</textarea></td>
|
||||
<td>
|
||||
<label class="toggle">
|
||||
<input type="checkbox" ${item.should_trigger ? 'checked' : ''} onchange="updateTrigger(${idx}, this.checked)">
|
||||
<span class="slider"></span>
|
||||
</label>
|
||||
<span style="margin-left:8px;font-size:0.8rem;color:#b0aea5">${item.should_trigger ? 'Yes' : 'No'}</span>
|
||||
</td>
|
||||
<td><button class="btn-delete" onclick="deleteRow(${idx})">Delete</button></td>
|
||||
`;
|
||||
tbody.appendChild(tr);
|
||||
});
|
||||
updateSummary();
|
||||
}
|
||||
|
||||
function escapeHtml(text) {
|
||||
const div = document.createElement('div');
|
||||
div.textContent = text;
|
||||
return div.innerHTML;
|
||||
}
|
||||
|
||||
function updateQuery(idx, value) { evalItems[idx].query = value; updateSummary(); }
|
||||
function updateTrigger(idx, value) { evalItems[idx].should_trigger = value; render(); }
|
||||
function deleteRow(idx) { evalItems.splice(idx, 1); render(); }
|
||||
|
||||
function addRow() {
|
||||
evalItems.push({ query: '', should_trigger: true });
|
||||
render();
|
||||
const inputs = document.querySelectorAll('.query-input');
|
||||
inputs[inputs.length - 1].focus();
|
||||
}
|
||||
|
||||
function updateSummary() {
|
||||
const trigger = evalItems.filter(i => i.should_trigger).length;
|
||||
const noTrigger = evalItems.filter(i => !i.should_trigger).length;
|
||||
document.getElementById('summary').textContent =
|
||||
`${evalItems.length} queries total: ${trigger} should trigger, ${noTrigger} should not trigger`;
|
||||
}
|
||||
|
||||
function exportEvalSet() {
|
||||
const valid = evalItems.filter(i => i.query.trim() !== '');
|
||||
const data = valid.map(i => ({ query: i.query.trim(), should_trigger: i.should_trigger }));
|
||||
const blob = new Blob([JSON.stringify(data, null, 2)], { type: 'application/json' });
|
||||
const url = URL.createObjectURL(blob);
|
||||
const a = document.createElement('a');
|
||||
a.href = url;
|
||||
a.download = 'eval_set.json';
|
||||
document.body.appendChild(a);
|
||||
a.click();
|
||||
document.body.removeChild(a);
|
||||
URL.revokeObjectURL(url);
|
||||
}
|
||||
|
||||
render();
|
||||
</script>
|
||||
</body>
|
||||
</html>
|
||||
471
.skills/skill-creator/eval-viewer/generate_review.py
Normal file
471
.skills/skill-creator/eval-viewer/generate_review.py
Normal file
|
|
@ -0,0 +1,471 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Generate and serve a review page for eval results.
|
||||
|
||||
Reads the workspace directory, discovers runs (directories with outputs/),
|
||||
embeds all output data into a self-contained HTML page, and serves it via
|
||||
a tiny HTTP server. Feedback auto-saves to feedback.json in the workspace.
|
||||
|
||||
Usage:
|
||||
python generate_review.py <workspace-path> [--port PORT] [--skill-name NAME]
|
||||
python generate_review.py <workspace-path> --previous-feedback /path/to/old/feedback.json
|
||||
|
||||
No dependencies beyond the Python stdlib are required.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import base64
|
||||
import json
|
||||
import mimetypes
|
||||
import os
|
||||
import re
|
||||
import signal
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
import webbrowser
|
||||
from functools import partial
|
||||
from http.server import HTTPServer, BaseHTTPRequestHandler
|
||||
from pathlib import Path
|
||||
|
||||
# Files to exclude from output listings
|
||||
METADATA_FILES = {"transcript.md", "user_notes.md", "metrics.json"}
|
||||
|
||||
# Extensions we render as inline text
|
||||
TEXT_EXTENSIONS = {
|
||||
".txt", ".md", ".json", ".csv", ".py", ".js", ".ts", ".tsx", ".jsx",
|
||||
".yaml", ".yml", ".xml", ".html", ".css", ".sh", ".rb", ".go", ".rs",
|
||||
".java", ".c", ".cpp", ".h", ".hpp", ".sql", ".r", ".toml",
|
||||
}
|
||||
|
||||
# Extensions we render as inline images
|
||||
IMAGE_EXTENSIONS = {".png", ".jpg", ".jpeg", ".gif", ".svg", ".webp"}
|
||||
|
||||
# MIME type overrides for common types
|
||||
MIME_OVERRIDES = {
|
||||
".svg": "image/svg+xml",
|
||||
".xlsx": "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet",
|
||||
".docx": "application/vnd.openxmlformats-officedocument.wordprocessingml.document",
|
||||
".pptx": "application/vnd.openxmlformats-officedocument.presentationml.presentation",
|
||||
}
|
||||
|
||||
|
||||
def get_mime_type(path: Path) -> str:
|
||||
ext = path.suffix.lower()
|
||||
if ext in MIME_OVERRIDES:
|
||||
return MIME_OVERRIDES[ext]
|
||||
mime, _ = mimetypes.guess_type(str(path))
|
||||
return mime or "application/octet-stream"
|
||||
|
||||
|
||||
def find_runs(workspace: Path) -> list[dict]:
|
||||
"""Recursively find directories that contain an outputs/ subdirectory."""
|
||||
runs: list[dict] = []
|
||||
_find_runs_recursive(workspace, workspace, runs)
|
||||
runs.sort(key=lambda r: (r.get("eval_id", float("inf")), r["id"]))
|
||||
return runs
|
||||
|
||||
|
||||
def _find_runs_recursive(root: Path, current: Path, runs: list[dict]) -> None:
|
||||
if not current.is_dir():
|
||||
return
|
||||
|
||||
outputs_dir = current / "outputs"
|
||||
if outputs_dir.is_dir():
|
||||
run = build_run(root, current)
|
||||
if run:
|
||||
runs.append(run)
|
||||
return
|
||||
|
||||
skip = {"node_modules", ".git", "__pycache__", "skill", "inputs"}
|
||||
for child in sorted(current.iterdir()):
|
||||
if child.is_dir() and child.name not in skip:
|
||||
_find_runs_recursive(root, child, runs)
|
||||
|
||||
|
||||
def build_run(root: Path, run_dir: Path) -> dict | None:
|
||||
"""Build a run dict with prompt, outputs, and grading data."""
|
||||
prompt = ""
|
||||
eval_id = None
|
||||
|
||||
# Try eval_metadata.json
|
||||
for candidate in [run_dir / "eval_metadata.json", run_dir.parent / "eval_metadata.json"]:
|
||||
if candidate.exists():
|
||||
try:
|
||||
metadata = json.loads(candidate.read_text())
|
||||
prompt = metadata.get("prompt", "")
|
||||
eval_id = metadata.get("eval_id")
|
||||
except (json.JSONDecodeError, OSError):
|
||||
pass
|
||||
if prompt:
|
||||
break
|
||||
|
||||
# Fall back to transcript.md
|
||||
if not prompt:
|
||||
for candidate in [run_dir / "transcript.md", run_dir / "outputs" / "transcript.md"]:
|
||||
if candidate.exists():
|
||||
try:
|
||||
text = candidate.read_text()
|
||||
match = re.search(r"## Eval Prompt\n\n([\s\S]*?)(?=\n##|$)", text)
|
||||
if match:
|
||||
prompt = match.group(1).strip()
|
||||
except OSError:
|
||||
pass
|
||||
if prompt:
|
||||
break
|
||||
|
||||
if not prompt:
|
||||
prompt = "(No prompt found)"
|
||||
|
||||
run_id = str(run_dir.relative_to(root)).replace("/", "-").replace("\\", "-")
|
||||
|
||||
# Collect output files
|
||||
outputs_dir = run_dir / "outputs"
|
||||
output_files: list[dict] = []
|
||||
if outputs_dir.is_dir():
|
||||
for f in sorted(outputs_dir.iterdir()):
|
||||
if f.is_file() and f.name not in METADATA_FILES:
|
||||
output_files.append(embed_file(f))
|
||||
|
||||
# Load grading if present
|
||||
grading = None
|
||||
for candidate in [run_dir / "grading.json", run_dir.parent / "grading.json"]:
|
||||
if candidate.exists():
|
||||
try:
|
||||
grading = json.loads(candidate.read_text())
|
||||
except (json.JSONDecodeError, OSError):
|
||||
pass
|
||||
if grading:
|
||||
break
|
||||
|
||||
return {
|
||||
"id": run_id,
|
||||
"prompt": prompt,
|
||||
"eval_id": eval_id,
|
||||
"outputs": output_files,
|
||||
"grading": grading,
|
||||
}
|
||||
|
||||
|
||||
def embed_file(path: Path) -> dict:
|
||||
"""Read a file and return an embedded representation."""
|
||||
ext = path.suffix.lower()
|
||||
mime = get_mime_type(path)
|
||||
|
||||
if ext in TEXT_EXTENSIONS:
|
||||
try:
|
||||
content = path.read_text(errors="replace")
|
||||
except OSError:
|
||||
content = "(Error reading file)"
|
||||
return {
|
||||
"name": path.name,
|
||||
"type": "text",
|
||||
"content": content,
|
||||
}
|
||||
elif ext in IMAGE_EXTENSIONS:
|
||||
try:
|
||||
raw = path.read_bytes()
|
||||
b64 = base64.b64encode(raw).decode("ascii")
|
||||
except OSError:
|
||||
return {"name": path.name, "type": "error", "content": "(Error reading file)"}
|
||||
return {
|
||||
"name": path.name,
|
||||
"type": "image",
|
||||
"mime": mime,
|
||||
"data_uri": f"data:{mime};base64,{b64}",
|
||||
}
|
||||
elif ext == ".pdf":
|
||||
try:
|
||||
raw = path.read_bytes()
|
||||
b64 = base64.b64encode(raw).decode("ascii")
|
||||
except OSError:
|
||||
return {"name": path.name, "type": "error", "content": "(Error reading file)"}
|
||||
return {
|
||||
"name": path.name,
|
||||
"type": "pdf",
|
||||
"data_uri": f"data:{mime};base64,{b64}",
|
||||
}
|
||||
elif ext == ".xlsx":
|
||||
try:
|
||||
raw = path.read_bytes()
|
||||
b64 = base64.b64encode(raw).decode("ascii")
|
||||
except OSError:
|
||||
return {"name": path.name, "type": "error", "content": "(Error reading file)"}
|
||||
return {
|
||||
"name": path.name,
|
||||
"type": "xlsx",
|
||||
"data_b64": b64,
|
||||
}
|
||||
else:
|
||||
# Binary / unknown — base64 download link
|
||||
try:
|
||||
raw = path.read_bytes()
|
||||
b64 = base64.b64encode(raw).decode("ascii")
|
||||
except OSError:
|
||||
return {"name": path.name, "type": "error", "content": "(Error reading file)"}
|
||||
return {
|
||||
"name": path.name,
|
||||
"type": "binary",
|
||||
"mime": mime,
|
||||
"data_uri": f"data:{mime};base64,{b64}",
|
||||
}
|
||||
|
||||
|
||||
def load_previous_iteration(workspace: Path) -> dict[str, dict]:
|
||||
"""Load previous iteration's feedback and outputs.
|
||||
|
||||
Returns a map of run_id -> {"feedback": str, "outputs": list[dict]}.
|
||||
"""
|
||||
result: dict[str, dict] = {}
|
||||
|
||||
# Load feedback
|
||||
feedback_map: dict[str, str] = {}
|
||||
feedback_path = workspace / "feedback.json"
|
||||
if feedback_path.exists():
|
||||
try:
|
||||
data = json.loads(feedback_path.read_text())
|
||||
feedback_map = {
|
||||
r["run_id"]: r["feedback"]
|
||||
for r in data.get("reviews", [])
|
||||
if r.get("feedback", "").strip()
|
||||
}
|
||||
except (json.JSONDecodeError, OSError, KeyError):
|
||||
pass
|
||||
|
||||
# Load runs (to get outputs)
|
||||
prev_runs = find_runs(workspace)
|
||||
for run in prev_runs:
|
||||
result[run["id"]] = {
|
||||
"feedback": feedback_map.get(run["id"], ""),
|
||||
"outputs": run.get("outputs", []),
|
||||
}
|
||||
|
||||
# Also add feedback for run_ids that had feedback but no matching run
|
||||
for run_id, fb in feedback_map.items():
|
||||
if run_id not in result:
|
||||
result[run_id] = {"feedback": fb, "outputs": []}
|
||||
|
||||
return result
|
||||
|
||||
|
||||
def generate_html(
|
||||
runs: list[dict],
|
||||
skill_name: str,
|
||||
previous: dict[str, dict] | None = None,
|
||||
benchmark: dict | None = None,
|
||||
) -> str:
|
||||
"""Generate the complete standalone HTML page with embedded data."""
|
||||
template_path = Path(__file__).parent / "viewer.html"
|
||||
template = template_path.read_text()
|
||||
|
||||
# Build previous_feedback and previous_outputs maps for the template
|
||||
previous_feedback: dict[str, str] = {}
|
||||
previous_outputs: dict[str, list[dict]] = {}
|
||||
if previous:
|
||||
for run_id, data in previous.items():
|
||||
if data.get("feedback"):
|
||||
previous_feedback[run_id] = data["feedback"]
|
||||
if data.get("outputs"):
|
||||
previous_outputs[run_id] = data["outputs"]
|
||||
|
||||
embedded = {
|
||||
"skill_name": skill_name,
|
||||
"runs": runs,
|
||||
"previous_feedback": previous_feedback,
|
||||
"previous_outputs": previous_outputs,
|
||||
}
|
||||
if benchmark:
|
||||
embedded["benchmark"] = benchmark
|
||||
|
||||
data_json = json.dumps(embedded)
|
||||
|
||||
return template.replace("/*__EMBEDDED_DATA__*/", f"const EMBEDDED_DATA = {data_json};")
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# HTTP server (stdlib only, zero dependencies)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def _kill_port(port: int) -> None:
|
||||
"""Kill any process listening on the given port."""
|
||||
try:
|
||||
result = subprocess.run(
|
||||
["lsof", "-ti", f":{port}"],
|
||||
capture_output=True, text=True, timeout=5,
|
||||
)
|
||||
for pid_str in result.stdout.strip().split("\n"):
|
||||
if pid_str.strip():
|
||||
try:
|
||||
os.kill(int(pid_str.strip()), signal.SIGTERM)
|
||||
except (ProcessLookupError, ValueError):
|
||||
pass
|
||||
if result.stdout.strip():
|
||||
time.sleep(0.5)
|
||||
except subprocess.TimeoutExpired:
|
||||
pass
|
||||
except FileNotFoundError:
|
||||
print("Note: lsof not found, cannot check if port is in use", file=sys.stderr)
|
||||
|
||||
class ReviewHandler(BaseHTTPRequestHandler):
|
||||
"""Serves the review HTML and handles feedback saves.
|
||||
|
||||
Regenerates the HTML on each page load so that refreshing the browser
|
||||
picks up new eval outputs without restarting the server.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
workspace: Path,
|
||||
skill_name: str,
|
||||
feedback_path: Path,
|
||||
previous: dict[str, dict],
|
||||
benchmark_path: Path | None,
|
||||
*args,
|
||||
**kwargs,
|
||||
):
|
||||
self.workspace = workspace
|
||||
self.skill_name = skill_name
|
||||
self.feedback_path = feedback_path
|
||||
self.previous = previous
|
||||
self.benchmark_path = benchmark_path
|
||||
super().__init__(*args, **kwargs)
|
||||
|
||||
def do_GET(self) -> None:
|
||||
if self.path == "/" or self.path == "/index.html":
|
||||
# Regenerate HTML on each request (re-scans workspace for new outputs)
|
||||
runs = find_runs(self.workspace)
|
||||
benchmark = None
|
||||
if self.benchmark_path and self.benchmark_path.exists():
|
||||
try:
|
||||
benchmark = json.loads(self.benchmark_path.read_text())
|
||||
except (json.JSONDecodeError, OSError):
|
||||
pass
|
||||
html = generate_html(runs, self.skill_name, self.previous, benchmark)
|
||||
content = html.encode("utf-8")
|
||||
self.send_response(200)
|
||||
self.send_header("Content-Type", "text/html; charset=utf-8")
|
||||
self.send_header("Content-Length", str(len(content)))
|
||||
self.end_headers()
|
||||
self.wfile.write(content)
|
||||
elif self.path == "/api/feedback":
|
||||
data = b"{}"
|
||||
if self.feedback_path.exists():
|
||||
data = self.feedback_path.read_bytes()
|
||||
self.send_response(200)
|
||||
self.send_header("Content-Type", "application/json")
|
||||
self.send_header("Content-Length", str(len(data)))
|
||||
self.end_headers()
|
||||
self.wfile.write(data)
|
||||
else:
|
||||
self.send_error(404)
|
||||
|
||||
def do_POST(self) -> None:
|
||||
if self.path == "/api/feedback":
|
||||
length = int(self.headers.get("Content-Length", 0))
|
||||
body = self.rfile.read(length)
|
||||
try:
|
||||
data = json.loads(body)
|
||||
if not isinstance(data, dict) or "reviews" not in data:
|
||||
raise ValueError("Expected JSON object with 'reviews' key")
|
||||
self.feedback_path.write_text(json.dumps(data, indent=2) + "\n")
|
||||
resp = b'{"ok":true}'
|
||||
self.send_response(200)
|
||||
except (json.JSONDecodeError, OSError, ValueError) as e:
|
||||
resp = json.dumps({"error": str(e)}).encode()
|
||||
self.send_response(500)
|
||||
self.send_header("Content-Type", "application/json")
|
||||
self.send_header("Content-Length", str(len(resp)))
|
||||
self.end_headers()
|
||||
self.wfile.write(resp)
|
||||
else:
|
||||
self.send_error(404)
|
||||
|
||||
def log_message(self, format: str, *args: object) -> None:
|
||||
# Suppress request logging to keep terminal clean
|
||||
pass
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(description="Generate and serve eval review")
|
||||
parser.add_argument("workspace", type=Path, help="Path to workspace directory")
|
||||
parser.add_argument("--port", "-p", type=int, default=3117, help="Server port (default: 3117)")
|
||||
parser.add_argument("--skill-name", "-n", type=str, default=None, help="Skill name for header")
|
||||
parser.add_argument(
|
||||
"--previous-workspace", type=Path, default=None,
|
||||
help="Path to previous iteration's workspace (shows old outputs and feedback as context)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--benchmark", type=Path, default=None,
|
||||
help="Path to benchmark.json to show in the Benchmark tab",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--static", "-s", type=Path, default=None,
|
||||
help="Write standalone HTML to this path instead of starting a server",
|
||||
)
|
||||
args = parser.parse_args()
|
||||
|
||||
workspace = args.workspace.resolve()
|
||||
if not workspace.is_dir():
|
||||
print(f"Error: {workspace} is not a directory", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
runs = find_runs(workspace)
|
||||
if not runs:
|
||||
print(f"No runs found in {workspace}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
skill_name = args.skill_name or workspace.name.replace("-workspace", "")
|
||||
feedback_path = workspace / "feedback.json"
|
||||
|
||||
previous: dict[str, dict] = {}
|
||||
if args.previous_workspace:
|
||||
previous = load_previous_iteration(args.previous_workspace.resolve())
|
||||
|
||||
benchmark_path = args.benchmark.resolve() if args.benchmark else None
|
||||
benchmark = None
|
||||
if benchmark_path and benchmark_path.exists():
|
||||
try:
|
||||
benchmark = json.loads(benchmark_path.read_text())
|
||||
except (json.JSONDecodeError, OSError):
|
||||
pass
|
||||
|
||||
if args.static:
|
||||
html = generate_html(runs, skill_name, previous, benchmark)
|
||||
args.static.parent.mkdir(parents=True, exist_ok=True)
|
||||
args.static.write_text(html)
|
||||
print(f"\n Static viewer written to: {args.static}\n")
|
||||
sys.exit(0)
|
||||
|
||||
# Kill any existing process on the target port
|
||||
port = args.port
|
||||
_kill_port(port)
|
||||
handler = partial(ReviewHandler, workspace, skill_name, feedback_path, previous, benchmark_path)
|
||||
try:
|
||||
server = HTTPServer(("127.0.0.1", port), handler)
|
||||
except OSError:
|
||||
# Port still in use after kill attempt — find a free one
|
||||
server = HTTPServer(("127.0.0.1", 0), handler)
|
||||
port = server.server_address[1]
|
||||
|
||||
url = f"http://localhost:{port}"
|
||||
print(f"\n Eval Viewer")
|
||||
print(f" ─────────────────────────────────")
|
||||
print(f" URL: {url}")
|
||||
print(f" Workspace: {workspace}")
|
||||
print(f" Feedback: {feedback_path}")
|
||||
if previous:
|
||||
print(f" Previous: {args.previous_workspace} ({len(previous)} runs)")
|
||||
if benchmark_path:
|
||||
print(f" Benchmark: {benchmark_path}")
|
||||
print(f"\n Press Ctrl+C to stop.\n")
|
||||
|
||||
webbrowser.open(url)
|
||||
|
||||
try:
|
||||
server.serve_forever()
|
||||
except KeyboardInterrupt:
|
||||
print("\nStopped.")
|
||||
server.server_close()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
1325
.skills/skill-creator/eval-viewer/viewer.html
Normal file
1325
.skills/skill-creator/eval-viewer/viewer.html
Normal file
File diff suppressed because it is too large
Load diff
430
.skills/skill-creator/references/schemas.md
Normal file
430
.skills/skill-creator/references/schemas.md
Normal file
|
|
@ -0,0 +1,430 @@
|
|||
# JSON Schemas
|
||||
|
||||
This document defines the JSON schemas used by skill-creator.
|
||||
|
||||
---
|
||||
|
||||
## evals.json
|
||||
|
||||
Defines the evals for a skill. Located at `evals/evals.json` within the skill directory.
|
||||
|
||||
```json
|
||||
{
|
||||
"skill_name": "example-skill",
|
||||
"evals": [
|
||||
{
|
||||
"id": 1,
|
||||
"prompt": "User's example prompt",
|
||||
"expected_output": "Description of expected result",
|
||||
"files": ["evals/files/sample1.pdf"],
|
||||
"expectations": [
|
||||
"The output includes X",
|
||||
"The skill used script Y"
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Fields:**
|
||||
- `skill_name`: Name matching the skill's frontmatter
|
||||
- `evals[].id`: Unique integer identifier
|
||||
- `evals[].prompt`: The task to execute
|
||||
- `evals[].expected_output`: Human-readable description of success
|
||||
- `evals[].files`: Optional list of input file paths (relative to skill root)
|
||||
- `evals[].expectations`: List of verifiable statements
|
||||
|
||||
---
|
||||
|
||||
## history.json
|
||||
|
||||
Tracks version progression in Improve mode. Located at workspace root.
|
||||
|
||||
```json
|
||||
{
|
||||
"started_at": "2026-01-15T10:30:00Z",
|
||||
"skill_name": "pdf",
|
||||
"current_best": "v2",
|
||||
"iterations": [
|
||||
{
|
||||
"version": "v0",
|
||||
"parent": null,
|
||||
"expectation_pass_rate": 0.65,
|
||||
"grading_result": "baseline",
|
||||
"is_current_best": false
|
||||
},
|
||||
{
|
||||
"version": "v1",
|
||||
"parent": "v0",
|
||||
"expectation_pass_rate": 0.75,
|
||||
"grading_result": "won",
|
||||
"is_current_best": false
|
||||
},
|
||||
{
|
||||
"version": "v2",
|
||||
"parent": "v1",
|
||||
"expectation_pass_rate": 0.85,
|
||||
"grading_result": "won",
|
||||
"is_current_best": true
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Fields:**
|
||||
- `started_at`: ISO timestamp of when improvement started
|
||||
- `skill_name`: Name of the skill being improved
|
||||
- `current_best`: Version identifier of the best performer
|
||||
- `iterations[].version`: Version identifier (v0, v1, ...)
|
||||
- `iterations[].parent`: Parent version this was derived from
|
||||
- `iterations[].expectation_pass_rate`: Pass rate from grading
|
||||
- `iterations[].grading_result`: "baseline", "won", "lost", or "tie"
|
||||
- `iterations[].is_current_best`: Whether this is the current best version
|
||||
|
||||
---
|
||||
|
||||
## grading.json
|
||||
|
||||
Output from the grader agent. Located at `<run-dir>/grading.json`.
|
||||
|
||||
```json
|
||||
{
|
||||
"expectations": [
|
||||
{
|
||||
"text": "The output includes the name 'John Smith'",
|
||||
"passed": true,
|
||||
"evidence": "Found in transcript Step 3: 'Extracted names: John Smith, Sarah Johnson'"
|
||||
},
|
||||
{
|
||||
"text": "The spreadsheet has a SUM formula in cell B10",
|
||||
"passed": false,
|
||||
"evidence": "No spreadsheet was created. The output was a text file."
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"passed": 2,
|
||||
"failed": 1,
|
||||
"total": 3,
|
||||
"pass_rate": 0.67
|
||||
},
|
||||
"execution_metrics": {
|
||||
"tool_calls": {
|
||||
"Read": 5,
|
||||
"Write": 2,
|
||||
"Bash": 8
|
||||
},
|
||||
"total_tool_calls": 15,
|
||||
"total_steps": 6,
|
||||
"errors_encountered": 0,
|
||||
"output_chars": 12450,
|
||||
"transcript_chars": 3200
|
||||
},
|
||||
"timing": {
|
||||
"executor_duration_seconds": 165.0,
|
||||
"grader_duration_seconds": 26.0,
|
||||
"total_duration_seconds": 191.0
|
||||
},
|
||||
"claims": [
|
||||
{
|
||||
"claim": "The form has 12 fillable fields",
|
||||
"type": "factual",
|
||||
"verified": true,
|
||||
"evidence": "Counted 12 fields in field_info.json"
|
||||
}
|
||||
],
|
||||
"user_notes_summary": {
|
||||
"uncertainties": ["Used 2023 data, may be stale"],
|
||||
"needs_review": [],
|
||||
"workarounds": ["Fell back to text overlay for non-fillable fields"]
|
||||
},
|
||||
"eval_feedback": {
|
||||
"suggestions": [
|
||||
{
|
||||
"assertion": "The output includes the name 'John Smith'",
|
||||
"reason": "A hallucinated document that mentions the name would also pass"
|
||||
}
|
||||
],
|
||||
"overall": "Assertions check presence but not correctness."
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Fields:**
|
||||
- `expectations[]`: Graded expectations with evidence
|
||||
- `summary`: Aggregate pass/fail counts
|
||||
- `execution_metrics`: Tool usage and output size (from executor's metrics.json)
|
||||
- `timing`: Wall clock timing (from timing.json)
|
||||
- `claims`: Extracted and verified claims from the output
|
||||
- `user_notes_summary`: Issues flagged by the executor
|
||||
- `eval_feedback`: (optional) Improvement suggestions for the evals, only present when the grader identifies issues worth raising
|
||||
|
||||
---
|
||||
|
||||
## metrics.json
|
||||
|
||||
Output from the executor agent. Located at `<run-dir>/outputs/metrics.json`.
|
||||
|
||||
```json
|
||||
{
|
||||
"tool_calls": {
|
||||
"Read": 5,
|
||||
"Write": 2,
|
||||
"Bash": 8,
|
||||
"Edit": 1,
|
||||
"Glob": 2,
|
||||
"Grep": 0
|
||||
},
|
||||
"total_tool_calls": 18,
|
||||
"total_steps": 6,
|
||||
"files_created": ["filled_form.pdf", "field_values.json"],
|
||||
"errors_encountered": 0,
|
||||
"output_chars": 12450,
|
||||
"transcript_chars": 3200
|
||||
}
|
||||
```
|
||||
|
||||
**Fields:**
|
||||
- `tool_calls`: Count per tool type
|
||||
- `total_tool_calls`: Sum of all tool calls
|
||||
- `total_steps`: Number of major execution steps
|
||||
- `files_created`: List of output files created
|
||||
- `errors_encountered`: Number of errors during execution
|
||||
- `output_chars`: Total character count of output files
|
||||
- `transcript_chars`: Character count of transcript
|
||||
|
||||
---
|
||||
|
||||
## timing.json
|
||||
|
||||
Wall clock timing for a run. Located at `<run-dir>/timing.json`.
|
||||
|
||||
**How to capture:** When a subagent task completes, the task notification includes `total_tokens` and `duration_ms`. Save these immediately — they are not persisted anywhere else and cannot be recovered after the fact.
|
||||
|
||||
```json
|
||||
{
|
||||
"total_tokens": 84852,
|
||||
"duration_ms": 23332,
|
||||
"total_duration_seconds": 23.3,
|
||||
"executor_start": "2026-01-15T10:30:00Z",
|
||||
"executor_end": "2026-01-15T10:32:45Z",
|
||||
"executor_duration_seconds": 165.0,
|
||||
"grader_start": "2026-01-15T10:32:46Z",
|
||||
"grader_end": "2026-01-15T10:33:12Z",
|
||||
"grader_duration_seconds": 26.0
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## benchmark.json
|
||||
|
||||
Output from Benchmark mode. Located at `benchmarks/<timestamp>/benchmark.json`.
|
||||
|
||||
```json
|
||||
{
|
||||
"metadata": {
|
||||
"skill_name": "pdf",
|
||||
"skill_path": "/path/to/pdf",
|
||||
"executor_model": "claude-sonnet-4-20250514",
|
||||
"analyzer_model": "most-capable-model",
|
||||
"timestamp": "2026-01-15T10:30:00Z",
|
||||
"evals_run": [1, 2, 3],
|
||||
"runs_per_configuration": 3
|
||||
},
|
||||
|
||||
"runs": [
|
||||
{
|
||||
"eval_id": 1,
|
||||
"eval_name": "Ocean",
|
||||
"configuration": "with_skill",
|
||||
"run_number": 1,
|
||||
"result": {
|
||||
"pass_rate": 0.85,
|
||||
"passed": 6,
|
||||
"failed": 1,
|
||||
"total": 7,
|
||||
"time_seconds": 42.5,
|
||||
"tokens": 3800,
|
||||
"tool_calls": 18,
|
||||
"errors": 0
|
||||
},
|
||||
"expectations": [
|
||||
{"text": "...", "passed": true, "evidence": "..."}
|
||||
],
|
||||
"notes": [
|
||||
"Used 2023 data, may be stale",
|
||||
"Fell back to text overlay for non-fillable fields"
|
||||
]
|
||||
}
|
||||
],
|
||||
|
||||
"run_summary": {
|
||||
"with_skill": {
|
||||
"pass_rate": {"mean": 0.85, "stddev": 0.05, "min": 0.80, "max": 0.90},
|
||||
"time_seconds": {"mean": 45.0, "stddev": 12.0, "min": 32.0, "max": 58.0},
|
||||
"tokens": {"mean": 3800, "stddev": 400, "min": 3200, "max": 4100}
|
||||
},
|
||||
"without_skill": {
|
||||
"pass_rate": {"mean": 0.35, "stddev": 0.08, "min": 0.28, "max": 0.45},
|
||||
"time_seconds": {"mean": 32.0, "stddev": 8.0, "min": 24.0, "max": 42.0},
|
||||
"tokens": {"mean": 2100, "stddev": 300, "min": 1800, "max": 2500}
|
||||
},
|
||||
"delta": {
|
||||
"pass_rate": "+0.50",
|
||||
"time_seconds": "+13.0",
|
||||
"tokens": "+1700"
|
||||
}
|
||||
},
|
||||
|
||||
"notes": [
|
||||
"Assertion 'Output is a PDF file' passes 100% in both configurations - may not differentiate skill value",
|
||||
"Eval 3 shows high variance (50% ± 40%) - may be flaky or model-dependent",
|
||||
"Without-skill runs consistently fail on table extraction expectations",
|
||||
"Skill adds 13s average execution time but improves pass rate by 50%"
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Fields:**
|
||||
- `metadata`: Information about the benchmark run
|
||||
- `skill_name`: Name of the skill
|
||||
- `timestamp`: When the benchmark was run
|
||||
- `evals_run`: List of eval names or IDs
|
||||
- `runs_per_configuration`: Number of runs per config (e.g. 3)
|
||||
- `runs[]`: Individual run results
|
||||
- `eval_id`: Numeric eval identifier
|
||||
- `eval_name`: Human-readable eval name (used as section header in the viewer)
|
||||
- `configuration`: Must be `"with_skill"` or `"without_skill"` (the viewer uses this exact string for grouping and color coding)
|
||||
- `run_number`: Integer run number (1, 2, 3...)
|
||||
- `result`: Nested object with `pass_rate`, `passed`, `total`, `time_seconds`, `tokens`, `errors`
|
||||
- `run_summary`: Statistical aggregates per configuration
|
||||
- `with_skill` / `without_skill`: Each contains `pass_rate`, `time_seconds`, `tokens` objects with `mean` and `stddev` fields
|
||||
- `delta`: Difference strings like `"+0.50"`, `"+13.0"`, `"+1700"`
|
||||
- `notes`: Freeform observations from the analyzer
|
||||
|
||||
**Important:** The viewer reads these field names exactly. Using `config` instead of `configuration`, or putting `pass_rate` at the top level of a run instead of nested under `result`, will cause the viewer to show empty/zero values. Always reference this schema when generating benchmark.json manually.
|
||||
|
||||
---
|
||||
|
||||
## comparison.json
|
||||
|
||||
Output from blind comparator. Located at `<grading-dir>/comparison-N.json`.
|
||||
|
||||
```json
|
||||
{
|
||||
"winner": "A",
|
||||
"reasoning": "Output A provides a complete solution with proper formatting and all required fields. Output B is missing the date field and has formatting inconsistencies.",
|
||||
"rubric": {
|
||||
"A": {
|
||||
"content": {
|
||||
"correctness": 5,
|
||||
"completeness": 5,
|
||||
"accuracy": 4
|
||||
},
|
||||
"structure": {
|
||||
"organization": 4,
|
||||
"formatting": 5,
|
||||
"usability": 4
|
||||
},
|
||||
"content_score": 4.7,
|
||||
"structure_score": 4.3,
|
||||
"overall_score": 9.0
|
||||
},
|
||||
"B": {
|
||||
"content": {
|
||||
"correctness": 3,
|
||||
"completeness": 2,
|
||||
"accuracy": 3
|
||||
},
|
||||
"structure": {
|
||||
"organization": 3,
|
||||
"formatting": 2,
|
||||
"usability": 3
|
||||
},
|
||||
"content_score": 2.7,
|
||||
"structure_score": 2.7,
|
||||
"overall_score": 5.4
|
||||
}
|
||||
},
|
||||
"output_quality": {
|
||||
"A": {
|
||||
"score": 9,
|
||||
"strengths": ["Complete solution", "Well-formatted", "All fields present"],
|
||||
"weaknesses": ["Minor style inconsistency in header"]
|
||||
},
|
||||
"B": {
|
||||
"score": 5,
|
||||
"strengths": ["Readable output", "Correct basic structure"],
|
||||
"weaknesses": ["Missing date field", "Formatting inconsistencies", "Partial data extraction"]
|
||||
}
|
||||
},
|
||||
"expectation_results": {
|
||||
"A": {
|
||||
"passed": 4,
|
||||
"total": 5,
|
||||
"pass_rate": 0.80,
|
||||
"details": [
|
||||
{"text": "Output includes name", "passed": true}
|
||||
]
|
||||
},
|
||||
"B": {
|
||||
"passed": 3,
|
||||
"total": 5,
|
||||
"pass_rate": 0.60,
|
||||
"details": [
|
||||
{"text": "Output includes name", "passed": true}
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## analysis.json
|
||||
|
||||
Output from post-hoc analyzer. Located at `<grading-dir>/analysis.json`.
|
||||
|
||||
```json
|
||||
{
|
||||
"comparison_summary": {
|
||||
"winner": "A",
|
||||
"winner_skill": "path/to/winner/skill",
|
||||
"loser_skill": "path/to/loser/skill",
|
||||
"comparator_reasoning": "Brief summary of why comparator chose winner"
|
||||
},
|
||||
"winner_strengths": [
|
||||
"Clear step-by-step instructions for handling multi-page documents",
|
||||
"Included validation script that caught formatting errors"
|
||||
],
|
||||
"loser_weaknesses": [
|
||||
"Vague instruction 'process the document appropriately' led to inconsistent behavior",
|
||||
"No script for validation, agent had to improvise"
|
||||
],
|
||||
"instruction_following": {
|
||||
"winner": {
|
||||
"score": 9,
|
||||
"issues": ["Minor: skipped optional logging step"]
|
||||
},
|
||||
"loser": {
|
||||
"score": 6,
|
||||
"issues": [
|
||||
"Did not use the skill's formatting template",
|
||||
"Invented own approach instead of following step 3"
|
||||
]
|
||||
}
|
||||
},
|
||||
"improvement_suggestions": [
|
||||
{
|
||||
"priority": "high",
|
||||
"category": "instructions",
|
||||
"suggestion": "Replace 'process the document appropriately' with explicit steps",
|
||||
"expected_impact": "Would eliminate ambiguity that caused inconsistent behavior"
|
||||
}
|
||||
],
|
||||
"transcript_insights": {
|
||||
"winner_execution_pattern": "Read skill -> Followed 5-step process -> Used validation script",
|
||||
"loser_execution_pattern": "Read skill -> Unclear on approach -> Tried 3 different methods"
|
||||
}
|
||||
}
|
||||
```
|
||||
0
.skills/skill-creator/scripts/__init__.py
Normal file
0
.skills/skill-creator/scripts/__init__.py
Normal file
401
.skills/skill-creator/scripts/aggregate_benchmark.py
Normal file
401
.skills/skill-creator/scripts/aggregate_benchmark.py
Normal file
|
|
@ -0,0 +1,401 @@
|
|||
#!/usr/bin/env python3
|
||||
"""
|
||||
Aggregate individual run results into benchmark summary statistics.
|
||||
|
||||
Reads grading.json files from run directories and produces:
|
||||
- run_summary with mean, stddev, min, max for each metric
|
||||
- delta between with_skill and without_skill configurations
|
||||
|
||||
Usage:
|
||||
python aggregate_benchmark.py <benchmark_dir>
|
||||
|
||||
Example:
|
||||
python aggregate_benchmark.py benchmarks/2026-01-15T10-30-00/
|
||||
|
||||
The script supports two directory layouts:
|
||||
|
||||
Workspace layout (from skill-creator iterations):
|
||||
<benchmark_dir>/
|
||||
└── eval-N/
|
||||
├── with_skill/
|
||||
│ ├── run-1/grading.json
|
||||
│ └── run-2/grading.json
|
||||
└── without_skill/
|
||||
├── run-1/grading.json
|
||||
└── run-2/grading.json
|
||||
|
||||
Legacy layout (with runs/ subdirectory):
|
||||
<benchmark_dir>/
|
||||
└── runs/
|
||||
└── eval-N/
|
||||
├── with_skill/
|
||||
│ └── run-1/grading.json
|
||||
└── without_skill/
|
||||
└── run-1/grading.json
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import math
|
||||
import sys
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def calculate_stats(values: list[float]) -> dict:
|
||||
"""Calculate mean, stddev, min, max for a list of values."""
|
||||
if not values:
|
||||
return {"mean": 0.0, "stddev": 0.0, "min": 0.0, "max": 0.0}
|
||||
|
||||
n = len(values)
|
||||
mean = sum(values) / n
|
||||
|
||||
if n > 1:
|
||||
variance = sum((x - mean) ** 2 for x in values) / (n - 1)
|
||||
stddev = math.sqrt(variance)
|
||||
else:
|
||||
stddev = 0.0
|
||||
|
||||
return {
|
||||
"mean": round(mean, 4),
|
||||
"stddev": round(stddev, 4),
|
||||
"min": round(min(values), 4),
|
||||
"max": round(max(values), 4)
|
||||
}
|
||||
|
||||
|
||||
def load_run_results(benchmark_dir: Path) -> dict:
|
||||
"""
|
||||
Load all run results from a benchmark directory.
|
||||
|
||||
Returns dict keyed by config name (e.g. "with_skill"/"without_skill",
|
||||
or "new_skill"/"old_skill"), each containing a list of run results.
|
||||
"""
|
||||
# Support both layouts: eval dirs directly under benchmark_dir, or under runs/
|
||||
runs_dir = benchmark_dir / "runs"
|
||||
if runs_dir.exists():
|
||||
search_dir = runs_dir
|
||||
elif list(benchmark_dir.glob("eval-*")):
|
||||
search_dir = benchmark_dir
|
||||
else:
|
||||
print(f"No eval directories found in {benchmark_dir} or {benchmark_dir / 'runs'}")
|
||||
return {}
|
||||
|
||||
results: dict[str, list] = {}
|
||||
|
||||
for eval_idx, eval_dir in enumerate(sorted(search_dir.glob("eval-*"))):
|
||||
metadata_path = eval_dir / "eval_metadata.json"
|
||||
if metadata_path.exists():
|
||||
try:
|
||||
with open(metadata_path) as mf:
|
||||
eval_id = json.load(mf).get("eval_id", eval_idx)
|
||||
except (json.JSONDecodeError, OSError):
|
||||
eval_id = eval_idx
|
||||
else:
|
||||
try:
|
||||
eval_id = int(eval_dir.name.split("-")[1])
|
||||
except ValueError:
|
||||
eval_id = eval_idx
|
||||
|
||||
# Discover config directories dynamically rather than hardcoding names
|
||||
for config_dir in sorted(eval_dir.iterdir()):
|
||||
if not config_dir.is_dir():
|
||||
continue
|
||||
# Skip non-config directories (inputs, outputs, etc.)
|
||||
if not list(config_dir.glob("run-*")):
|
||||
continue
|
||||
config = config_dir.name
|
||||
if config not in results:
|
||||
results[config] = []
|
||||
|
||||
for run_dir in sorted(config_dir.glob("run-*")):
|
||||
run_number = int(run_dir.name.split("-")[1])
|
||||
grading_file = run_dir / "grading.json"
|
||||
|
||||
if not grading_file.exists():
|
||||
print(f"Warning: grading.json not found in {run_dir}")
|
||||
continue
|
||||
|
||||
try:
|
||||
with open(grading_file) as f:
|
||||
grading = json.load(f)
|
||||
except json.JSONDecodeError as e:
|
||||
print(f"Warning: Invalid JSON in {grading_file}: {e}")
|
||||
continue
|
||||
|
||||
# Extract metrics
|
||||
result = {
|
||||
"eval_id": eval_id,
|
||||
"run_number": run_number,
|
||||
"pass_rate": grading.get("summary", {}).get("pass_rate", 0.0),
|
||||
"passed": grading.get("summary", {}).get("passed", 0),
|
||||
"failed": grading.get("summary", {}).get("failed", 0),
|
||||
"total": grading.get("summary", {}).get("total", 0),
|
||||
}
|
||||
|
||||
# Extract timing — check grading.json first, then sibling timing.json
|
||||
timing = grading.get("timing", {})
|
||||
result["time_seconds"] = timing.get("total_duration_seconds", 0.0)
|
||||
timing_file = run_dir / "timing.json"
|
||||
if result["time_seconds"] == 0.0 and timing_file.exists():
|
||||
try:
|
||||
with open(timing_file) as tf:
|
||||
timing_data = json.load(tf)
|
||||
result["time_seconds"] = timing_data.get("total_duration_seconds", 0.0)
|
||||
result["tokens"] = timing_data.get("total_tokens", 0)
|
||||
except json.JSONDecodeError:
|
||||
pass
|
||||
|
||||
# Extract metrics if available
|
||||
metrics = grading.get("execution_metrics", {})
|
||||
result["tool_calls"] = metrics.get("total_tool_calls", 0)
|
||||
if not result.get("tokens"):
|
||||
result["tokens"] = metrics.get("output_chars", 0)
|
||||
result["errors"] = metrics.get("errors_encountered", 0)
|
||||
|
||||
# Extract expectations — viewer requires fields: text, passed, evidence
|
||||
raw_expectations = grading.get("expectations", [])
|
||||
for exp in raw_expectations:
|
||||
if "text" not in exp or "passed" not in exp:
|
||||
print(f"Warning: expectation in {grading_file} missing required fields (text, passed, evidence): {exp}")
|
||||
result["expectations"] = raw_expectations
|
||||
|
||||
# Extract notes from user_notes_summary
|
||||
notes_summary = grading.get("user_notes_summary", {})
|
||||
notes = []
|
||||
notes.extend(notes_summary.get("uncertainties", []))
|
||||
notes.extend(notes_summary.get("needs_review", []))
|
||||
notes.extend(notes_summary.get("workarounds", []))
|
||||
result["notes"] = notes
|
||||
|
||||
results[config].append(result)
|
||||
|
||||
return results
|
||||
|
||||
|
||||
def aggregate_results(results: dict) -> dict:
|
||||
"""
|
||||
Aggregate run results into summary statistics.
|
||||
|
||||
Returns run_summary with stats for each configuration and delta.
|
||||
"""
|
||||
run_summary = {}
|
||||
configs = list(results.keys())
|
||||
|
||||
for config in configs:
|
||||
runs = results.get(config, [])
|
||||
|
||||
if not runs:
|
||||
run_summary[config] = {
|
||||
"pass_rate": {"mean": 0.0, "stddev": 0.0, "min": 0.0, "max": 0.0},
|
||||
"time_seconds": {"mean": 0.0, "stddev": 0.0, "min": 0.0, "max": 0.0},
|
||||
"tokens": {"mean": 0, "stddev": 0, "min": 0, "max": 0}
|
||||
}
|
||||
continue
|
||||
|
||||
pass_rates = [r["pass_rate"] for r in runs]
|
||||
times = [r["time_seconds"] for r in runs]
|
||||
tokens = [r.get("tokens", 0) for r in runs]
|
||||
|
||||
run_summary[config] = {
|
||||
"pass_rate": calculate_stats(pass_rates),
|
||||
"time_seconds": calculate_stats(times),
|
||||
"tokens": calculate_stats(tokens)
|
||||
}
|
||||
|
||||
# Calculate delta between the first two configs (if two exist)
|
||||
if len(configs) >= 2:
|
||||
primary = run_summary.get(configs[0], {})
|
||||
baseline = run_summary.get(configs[1], {})
|
||||
else:
|
||||
primary = run_summary.get(configs[0], {}) if configs else {}
|
||||
baseline = {}
|
||||
|
||||
delta_pass_rate = primary.get("pass_rate", {}).get("mean", 0) - baseline.get("pass_rate", {}).get("mean", 0)
|
||||
delta_time = primary.get("time_seconds", {}).get("mean", 0) - baseline.get("time_seconds", {}).get("mean", 0)
|
||||
delta_tokens = primary.get("tokens", {}).get("mean", 0) - baseline.get("tokens", {}).get("mean", 0)
|
||||
|
||||
run_summary["delta"] = {
|
||||
"pass_rate": f"{delta_pass_rate:+.2f}",
|
||||
"time_seconds": f"{delta_time:+.1f}",
|
||||
"tokens": f"{delta_tokens:+.0f}"
|
||||
}
|
||||
|
||||
return run_summary
|
||||
|
||||
|
||||
def generate_benchmark(benchmark_dir: Path, skill_name: str = "", skill_path: str = "") -> dict:
|
||||
"""
|
||||
Generate complete benchmark.json from run results.
|
||||
"""
|
||||
results = load_run_results(benchmark_dir)
|
||||
run_summary = aggregate_results(results)
|
||||
|
||||
# Build runs array for benchmark.json
|
||||
runs = []
|
||||
for config in results:
|
||||
for result in results[config]:
|
||||
runs.append({
|
||||
"eval_id": result["eval_id"],
|
||||
"configuration": config,
|
||||
"run_number": result["run_number"],
|
||||
"result": {
|
||||
"pass_rate": result["pass_rate"],
|
||||
"passed": result["passed"],
|
||||
"failed": result["failed"],
|
||||
"total": result["total"],
|
||||
"time_seconds": result["time_seconds"],
|
||||
"tokens": result.get("tokens", 0),
|
||||
"tool_calls": result.get("tool_calls", 0),
|
||||
"errors": result.get("errors", 0)
|
||||
},
|
||||
"expectations": result["expectations"],
|
||||
"notes": result["notes"]
|
||||
})
|
||||
|
||||
# Determine eval IDs from results
|
||||
eval_ids = sorted(set(
|
||||
r["eval_id"]
|
||||
for config in results.values()
|
||||
for r in config
|
||||
))
|
||||
|
||||
benchmark = {
|
||||
"metadata": {
|
||||
"skill_name": skill_name or "<skill-name>",
|
||||
"skill_path": skill_path or "<path/to/skill>",
|
||||
"executor_model": "<model-name>",
|
||||
"analyzer_model": "<model-name>",
|
||||
"timestamp": datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ"),
|
||||
"evals_run": eval_ids,
|
||||
"runs_per_configuration": 3
|
||||
},
|
||||
"runs": runs,
|
||||
"run_summary": run_summary,
|
||||
"notes": [] # To be filled by analyzer
|
||||
}
|
||||
|
||||
return benchmark
|
||||
|
||||
|
||||
def generate_markdown(benchmark: dict) -> str:
|
||||
"""Generate human-readable benchmark.md from benchmark data."""
|
||||
metadata = benchmark["metadata"]
|
||||
run_summary = benchmark["run_summary"]
|
||||
|
||||
# Determine config names (excluding "delta")
|
||||
configs = [k for k in run_summary if k != "delta"]
|
||||
config_a = configs[0] if len(configs) >= 1 else "config_a"
|
||||
config_b = configs[1] if len(configs) >= 2 else "config_b"
|
||||
label_a = config_a.replace("_", " ").title()
|
||||
label_b = config_b.replace("_", " ").title()
|
||||
|
||||
lines = [
|
||||
f"# Skill Benchmark: {metadata['skill_name']}",
|
||||
"",
|
||||
f"**Model**: {metadata['executor_model']}",
|
||||
f"**Date**: {metadata['timestamp']}",
|
||||
f"**Evals**: {', '.join(map(str, metadata['evals_run']))} ({metadata['runs_per_configuration']} runs each per configuration)",
|
||||
"",
|
||||
"## Summary",
|
||||
"",
|
||||
f"| Metric | {label_a} | {label_b} | Delta |",
|
||||
"|--------|------------|---------------|-------|",
|
||||
]
|
||||
|
||||
a_summary = run_summary.get(config_a, {})
|
||||
b_summary = run_summary.get(config_b, {})
|
||||
delta = run_summary.get("delta", {})
|
||||
|
||||
# Format pass rate
|
||||
a_pr = a_summary.get("pass_rate", {})
|
||||
b_pr = b_summary.get("pass_rate", {})
|
||||
lines.append(f"| Pass Rate | {a_pr.get('mean', 0)*100:.0f}% ± {a_pr.get('stddev', 0)*100:.0f}% | {b_pr.get('mean', 0)*100:.0f}% ± {b_pr.get('stddev', 0)*100:.0f}% | {delta.get('pass_rate', '—')} |")
|
||||
|
||||
# Format time
|
||||
a_time = a_summary.get("time_seconds", {})
|
||||
b_time = b_summary.get("time_seconds", {})
|
||||
lines.append(f"| Time | {a_time.get('mean', 0):.1f}s ± {a_time.get('stddev', 0):.1f}s | {b_time.get('mean', 0):.1f}s ± {b_time.get('stddev', 0):.1f}s | {delta.get('time_seconds', '—')}s |")
|
||||
|
||||
# Format tokens
|
||||
a_tokens = a_summary.get("tokens", {})
|
||||
b_tokens = b_summary.get("tokens", {})
|
||||
lines.append(f"| Tokens | {a_tokens.get('mean', 0):.0f} ± {a_tokens.get('stddev', 0):.0f} | {b_tokens.get('mean', 0):.0f} ± {b_tokens.get('stddev', 0):.0f} | {delta.get('tokens', '—')} |")
|
||||
|
||||
# Notes section
|
||||
if benchmark.get("notes"):
|
||||
lines.extend([
|
||||
"",
|
||||
"## Notes",
|
||||
""
|
||||
])
|
||||
for note in benchmark["notes"]:
|
||||
lines.append(f"- {note}")
|
||||
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(
|
||||
description="Aggregate benchmark run results into summary statistics"
|
||||
)
|
||||
parser.add_argument(
|
||||
"benchmark_dir",
|
||||
type=Path,
|
||||
help="Path to the benchmark directory"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--skill-name",
|
||||
default="",
|
||||
help="Name of the skill being benchmarked"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--skill-path",
|
||||
default="",
|
||||
help="Path to the skill being benchmarked"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--output", "-o",
|
||||
type=Path,
|
||||
help="Output path for benchmark.json (default: <benchmark_dir>/benchmark.json)"
|
||||
)
|
||||
|
||||
args = parser.parse_args()
|
||||
|
||||
if not args.benchmark_dir.exists():
|
||||
print(f"Directory not found: {args.benchmark_dir}")
|
||||
sys.exit(1)
|
||||
|
||||
# Generate benchmark
|
||||
benchmark = generate_benchmark(args.benchmark_dir, args.skill_name, args.skill_path)
|
||||
|
||||
# Determine output paths
|
||||
output_json = args.output or (args.benchmark_dir / "benchmark.json")
|
||||
output_md = output_json.with_suffix(".md")
|
||||
|
||||
# Write benchmark.json
|
||||
with open(output_json, "w") as f:
|
||||
json.dump(benchmark, f, indent=2)
|
||||
print(f"Generated: {output_json}")
|
||||
|
||||
# Write benchmark.md
|
||||
markdown = generate_markdown(benchmark)
|
||||
with open(output_md, "w") as f:
|
||||
f.write(markdown)
|
||||
print(f"Generated: {output_md}")
|
||||
|
||||
# Print summary
|
||||
run_summary = benchmark["run_summary"]
|
||||
configs = [k for k in run_summary if k != "delta"]
|
||||
delta = run_summary.get("delta", {})
|
||||
|
||||
print(f"\nSummary:")
|
||||
for config in configs:
|
||||
pr = run_summary[config]["pass_rate"]["mean"]
|
||||
label = config.replace("_", " ").title()
|
||||
print(f" {label}: {pr*100:.1f}% pass rate")
|
||||
print(f" Delta: {delta.get('pass_rate', '—')}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
326
.skills/skill-creator/scripts/generate_report.py
Normal file
326
.skills/skill-creator/scripts/generate_report.py
Normal file
|
|
@ -0,0 +1,326 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Generate an HTML report from run_loop.py output.
|
||||
|
||||
Takes the JSON output from run_loop.py and generates a visual HTML report
|
||||
showing each description attempt with check/x for each test case.
|
||||
Distinguishes between train and test queries.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import html
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def generate_html(data: dict, auto_refresh: bool = False, skill_name: str = "") -> str:
|
||||
"""Generate HTML report from loop output data. If auto_refresh is True, adds a meta refresh tag."""
|
||||
history = data.get("history", [])
|
||||
holdout = data.get("holdout", 0)
|
||||
title_prefix = html.escape(skill_name + " \u2014 ") if skill_name else ""
|
||||
|
||||
# Get all unique queries from train and test sets, with should_trigger info
|
||||
train_queries: list[dict] = []
|
||||
test_queries: list[dict] = []
|
||||
if history:
|
||||
for r in history[0].get("train_results", history[0].get("results", [])):
|
||||
train_queries.append({"query": r["query"], "should_trigger": r.get("should_trigger", True)})
|
||||
if history[0].get("test_results"):
|
||||
for r in history[0].get("test_results", []):
|
||||
test_queries.append({"query": r["query"], "should_trigger": r.get("should_trigger", True)})
|
||||
|
||||
refresh_tag = ' <meta http-equiv="refresh" content="5">\n' if auto_refresh else ""
|
||||
|
||||
html_parts = ["""<!DOCTYPE html>
|
||||
<html>
|
||||
<head>
|
||||
<meta charset="utf-8">
|
||||
""" + refresh_tag + """ <title>""" + title_prefix + """Skill Description Optimization</title>
|
||||
<link rel="preconnect" href="https://fonts.googleapis.com">
|
||||
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
|
||||
<link href="https://fonts.googleapis.com/css2?family=Poppins:wght@500;600&family=Lora:wght@400;500&display=swap" rel="stylesheet">
|
||||
<style>
|
||||
body {
|
||||
font-family: 'Lora', Georgia, serif;
|
||||
max-width: 100%;
|
||||
margin: 0 auto;
|
||||
padding: 20px;
|
||||
background: #faf9f5;
|
||||
color: #141413;
|
||||
}
|
||||
h1 { font-family: 'Poppins', sans-serif; color: #141413; }
|
||||
.explainer {
|
||||
background: white;
|
||||
padding: 15px;
|
||||
border-radius: 6px;
|
||||
margin-bottom: 20px;
|
||||
border: 1px solid #e8e6dc;
|
||||
color: #b0aea5;
|
||||
font-size: 0.875rem;
|
||||
line-height: 1.6;
|
||||
}
|
||||
.summary {
|
||||
background: white;
|
||||
padding: 15px;
|
||||
border-radius: 6px;
|
||||
margin-bottom: 20px;
|
||||
border: 1px solid #e8e6dc;
|
||||
}
|
||||
.summary p { margin: 5px 0; }
|
||||
.best { color: #788c5d; font-weight: bold; }
|
||||
.table-container {
|
||||
overflow-x: auto;
|
||||
width: 100%;
|
||||
}
|
||||
table {
|
||||
border-collapse: collapse;
|
||||
background: white;
|
||||
border: 1px solid #e8e6dc;
|
||||
border-radius: 6px;
|
||||
font-size: 12px;
|
||||
min-width: 100%;
|
||||
}
|
||||
th, td {
|
||||
padding: 8px;
|
||||
text-align: left;
|
||||
border: 1px solid #e8e6dc;
|
||||
white-space: normal;
|
||||
word-wrap: break-word;
|
||||
}
|
||||
th {
|
||||
font-family: 'Poppins', sans-serif;
|
||||
background: #141413;
|
||||
color: #faf9f5;
|
||||
font-weight: 500;
|
||||
}
|
||||
th.test-col {
|
||||
background: #6a9bcc;
|
||||
}
|
||||
th.query-col { min-width: 200px; }
|
||||
td.description {
|
||||
font-family: monospace;
|
||||
font-size: 11px;
|
||||
word-wrap: break-word;
|
||||
max-width: 400px;
|
||||
}
|
||||
td.result {
|
||||
text-align: center;
|
||||
font-size: 16px;
|
||||
min-width: 40px;
|
||||
}
|
||||
td.test-result {
|
||||
background: #f0f6fc;
|
||||
}
|
||||
.pass { color: #788c5d; }
|
||||
.fail { color: #c44; }
|
||||
.rate {
|
||||
font-size: 9px;
|
||||
color: #b0aea5;
|
||||
display: block;
|
||||
}
|
||||
tr:hover { background: #faf9f5; }
|
||||
.score {
|
||||
display: inline-block;
|
||||
padding: 2px 6px;
|
||||
border-radius: 4px;
|
||||
font-weight: bold;
|
||||
font-size: 11px;
|
||||
}
|
||||
.score-good { background: #eef2e8; color: #788c5d; }
|
||||
.score-ok { background: #fef3c7; color: #d97706; }
|
||||
.score-bad { background: #fceaea; color: #c44; }
|
||||
.train-label { color: #b0aea5; font-size: 10px; }
|
||||
.test-label { color: #6a9bcc; font-size: 10px; font-weight: bold; }
|
||||
.best-row { background: #f5f8f2; }
|
||||
th.positive-col { border-bottom: 3px solid #788c5d; }
|
||||
th.negative-col { border-bottom: 3px solid #c44; }
|
||||
th.test-col.positive-col { border-bottom: 3px solid #788c5d; }
|
||||
th.test-col.negative-col { border-bottom: 3px solid #c44; }
|
||||
.legend { font-family: 'Poppins', sans-serif; display: flex; gap: 20px; margin-bottom: 10px; font-size: 13px; align-items: center; }
|
||||
.legend-item { display: flex; align-items: center; gap: 6px; }
|
||||
.legend-swatch { width: 16px; height: 16px; border-radius: 3px; display: inline-block; }
|
||||
.swatch-positive { background: #141413; border-bottom: 3px solid #788c5d; }
|
||||
.swatch-negative { background: #141413; border-bottom: 3px solid #c44; }
|
||||
.swatch-test { background: #6a9bcc; }
|
||||
.swatch-train { background: #141413; }
|
||||
</style>
|
||||
</head>
|
||||
<body>
|
||||
<h1>""" + title_prefix + """Skill Description Optimization</h1>
|
||||
<div class="explainer">
|
||||
<strong>Optimizing your skill's description.</strong> This page updates automatically as Claude tests different versions of your skill's description. Each row is an iteration — a new description attempt. The columns show test queries: green checkmarks mean the skill triggered correctly (or correctly didn't trigger), red crosses mean it got it wrong. The "Train" score shows performance on queries used to improve the description; the "Test" score shows performance on held-out queries the optimizer hasn't seen. When it's done, Claude will apply the best-performing description to your skill.
|
||||
</div>
|
||||
"""]
|
||||
|
||||
# Summary section
|
||||
best_test_score = data.get('best_test_score')
|
||||
best_train_score = data.get('best_train_score')
|
||||
html_parts.append(f"""
|
||||
<div class="summary">
|
||||
<p><strong>Original:</strong> {html.escape(data.get('original_description', 'N/A'))}</p>
|
||||
<p class="best"><strong>Best:</strong> {html.escape(data.get('best_description', 'N/A'))}</p>
|
||||
<p><strong>Best Score:</strong> {data.get('best_score', 'N/A')} {'(test)' if best_test_score else '(train)'}</p>
|
||||
<p><strong>Iterations:</strong> {data.get('iterations_run', 0)} | <strong>Train:</strong> {data.get('train_size', '?')} | <strong>Test:</strong> {data.get('test_size', '?')}</p>
|
||||
</div>
|
||||
""")
|
||||
|
||||
# Legend
|
||||
html_parts.append("""
|
||||
<div class="legend">
|
||||
<span style="font-weight:600">Query columns:</span>
|
||||
<span class="legend-item"><span class="legend-swatch swatch-positive"></span> Should trigger</span>
|
||||
<span class="legend-item"><span class="legend-swatch swatch-negative"></span> Should NOT trigger</span>
|
||||
<span class="legend-item"><span class="legend-swatch swatch-train"></span> Train</span>
|
||||
<span class="legend-item"><span class="legend-swatch swatch-test"></span> Test</span>
|
||||
</div>
|
||||
""")
|
||||
|
||||
# Table header
|
||||
html_parts.append("""
|
||||
<div class="table-container">
|
||||
<table>
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Iter</th>
|
||||
<th>Train</th>
|
||||
<th>Test</th>
|
||||
<th class="query-col">Description</th>
|
||||
""")
|
||||
|
||||
# Add column headers for train queries
|
||||
for qinfo in train_queries:
|
||||
polarity = "positive-col" if qinfo["should_trigger"] else "negative-col"
|
||||
html_parts.append(f' <th class="{polarity}">{html.escape(qinfo["query"])}</th>\n')
|
||||
|
||||
# Add column headers for test queries (different color)
|
||||
for qinfo in test_queries:
|
||||
polarity = "positive-col" if qinfo["should_trigger"] else "negative-col"
|
||||
html_parts.append(f' <th class="test-col {polarity}">{html.escape(qinfo["query"])}</th>\n')
|
||||
|
||||
html_parts.append(""" </tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
""")
|
||||
|
||||
# Find best iteration for highlighting
|
||||
if test_queries:
|
||||
best_iter = max(history, key=lambda h: h.get("test_passed") or 0).get("iteration")
|
||||
else:
|
||||
best_iter = max(history, key=lambda h: h.get("train_passed", h.get("passed", 0))).get("iteration")
|
||||
|
||||
# Add rows for each iteration
|
||||
for h in history:
|
||||
iteration = h.get("iteration", "?")
|
||||
train_passed = h.get("train_passed", h.get("passed", 0))
|
||||
train_total = h.get("train_total", h.get("total", 0))
|
||||
test_passed = h.get("test_passed")
|
||||
test_total = h.get("test_total")
|
||||
description = h.get("description", "")
|
||||
train_results = h.get("train_results", h.get("results", []))
|
||||
test_results = h.get("test_results", [])
|
||||
|
||||
# Create lookups for results by query
|
||||
train_by_query = {r["query"]: r for r in train_results}
|
||||
test_by_query = {r["query"]: r for r in test_results} if test_results else {}
|
||||
|
||||
# Compute aggregate correct/total runs across all retries
|
||||
def aggregate_runs(results: list[dict]) -> tuple[int, int]:
|
||||
correct = 0
|
||||
total = 0
|
||||
for r in results:
|
||||
runs = r.get("runs", 0)
|
||||
triggers = r.get("triggers", 0)
|
||||
total += runs
|
||||
if r.get("should_trigger", True):
|
||||
correct += triggers
|
||||
else:
|
||||
correct += runs - triggers
|
||||
return correct, total
|
||||
|
||||
train_correct, train_runs = aggregate_runs(train_results)
|
||||
test_correct, test_runs = aggregate_runs(test_results)
|
||||
|
||||
# Determine score classes
|
||||
def score_class(correct: int, total: int) -> str:
|
||||
if total > 0:
|
||||
ratio = correct / total
|
||||
if ratio >= 0.8:
|
||||
return "score-good"
|
||||
elif ratio >= 0.5:
|
||||
return "score-ok"
|
||||
return "score-bad"
|
||||
|
||||
train_class = score_class(train_correct, train_runs)
|
||||
test_class = score_class(test_correct, test_runs)
|
||||
|
||||
row_class = "best-row" if iteration == best_iter else ""
|
||||
|
||||
html_parts.append(f""" <tr class="{row_class}">
|
||||
<td>{iteration}</td>
|
||||
<td><span class="score {train_class}">{train_correct}/{train_runs}</span></td>
|
||||
<td><span class="score {test_class}">{test_correct}/{test_runs}</span></td>
|
||||
<td class="description">{html.escape(description)}</td>
|
||||
""")
|
||||
|
||||
# Add result for each train query
|
||||
for qinfo in train_queries:
|
||||
r = train_by_query.get(qinfo["query"], {})
|
||||
did_pass = r.get("pass", False)
|
||||
triggers = r.get("triggers", 0)
|
||||
runs = r.get("runs", 0)
|
||||
|
||||
icon = "✓" if did_pass else "✗"
|
||||
css_class = "pass" if did_pass else "fail"
|
||||
|
||||
html_parts.append(f' <td class="result {css_class}">{icon}<span class="rate">{triggers}/{runs}</span></td>\n')
|
||||
|
||||
# Add result for each test query (with different background)
|
||||
for qinfo in test_queries:
|
||||
r = test_by_query.get(qinfo["query"], {})
|
||||
did_pass = r.get("pass", False)
|
||||
triggers = r.get("triggers", 0)
|
||||
runs = r.get("runs", 0)
|
||||
|
||||
icon = "✓" if did_pass else "✗"
|
||||
css_class = "pass" if did_pass else "fail"
|
||||
|
||||
html_parts.append(f' <td class="result test-result {css_class}">{icon}<span class="rate">{triggers}/{runs}</span></td>\n')
|
||||
|
||||
html_parts.append(" </tr>\n")
|
||||
|
||||
html_parts.append(""" </tbody>
|
||||
</table>
|
||||
</div>
|
||||
""")
|
||||
|
||||
html_parts.append("""
|
||||
</body>
|
||||
</html>
|
||||
""")
|
||||
|
||||
return "".join(html_parts)
|
||||
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(description="Generate HTML report from run_loop output")
|
||||
parser.add_argument("input", help="Path to JSON output from run_loop.py (or - for stdin)")
|
||||
parser.add_argument("-o", "--output", default=None, help="Output HTML file (default: stdout)")
|
||||
parser.add_argument("--skill-name", default="", help="Skill name to include in the report title")
|
||||
args = parser.parse_args()
|
||||
|
||||
if args.input == "-":
|
||||
data = json.load(sys.stdin)
|
||||
else:
|
||||
data = json.loads(Path(args.input).read_text())
|
||||
|
||||
html_output = generate_html(data, skill_name=args.skill_name)
|
||||
|
||||
if args.output:
|
||||
Path(args.output).write_text(html_output)
|
||||
print(f"Report written to {args.output}", file=sys.stderr)
|
||||
else:
|
||||
print(html_output)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
248
.skills/skill-creator/scripts/improve_description.py
Normal file
248
.skills/skill-creator/scripts/improve_description.py
Normal file
|
|
@ -0,0 +1,248 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Improve a skill description based on eval results.
|
||||
|
||||
Takes eval results (from run_eval.py) and generates an improved description
|
||||
using Claude with extended thinking.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import re
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import anthropic
|
||||
|
||||
from scripts.utils import parse_skill_md
|
||||
|
||||
|
||||
def improve_description(
|
||||
client: anthropic.Anthropic,
|
||||
skill_name: str,
|
||||
skill_content: str,
|
||||
current_description: str,
|
||||
eval_results: dict,
|
||||
history: list[dict],
|
||||
model: str,
|
||||
test_results: dict | None = None,
|
||||
log_dir: Path | None = None,
|
||||
iteration: int | None = None,
|
||||
) -> str:
|
||||
"""Call Claude to improve the description based on eval results."""
|
||||
failed_triggers = [
|
||||
r for r in eval_results["results"]
|
||||
if r["should_trigger"] and not r["pass"]
|
||||
]
|
||||
false_triggers = [
|
||||
r for r in eval_results["results"]
|
||||
if not r["should_trigger"] and not r["pass"]
|
||||
]
|
||||
|
||||
# Build scores summary
|
||||
train_score = f"{eval_results['summary']['passed']}/{eval_results['summary']['total']}"
|
||||
if test_results:
|
||||
test_score = f"{test_results['summary']['passed']}/{test_results['summary']['total']}"
|
||||
scores_summary = f"Train: {train_score}, Test: {test_score}"
|
||||
else:
|
||||
scores_summary = f"Train: {train_score}"
|
||||
|
||||
prompt = f"""You are optimizing a skill description for a Claude Code skill called "{skill_name}". A "skill" is sort of like a prompt, but with progressive disclosure -- there's a title and description that Claude sees when deciding whether to use the skill, and then if it does use the skill, it reads the .md file which has lots more details and potentially links to other resources in the skill folder like helper files and scripts and additional documentation or examples.
|
||||
|
||||
The description appears in Claude's "available_skills" list. When a user sends a query, Claude decides whether to invoke the skill based solely on the title and on this description. Your goal is to write a description that triggers for relevant queries, and doesn't trigger for irrelevant ones.
|
||||
|
||||
Here's the current description:
|
||||
<current_description>
|
||||
"{current_description}"
|
||||
</current_description>
|
||||
|
||||
Current scores ({scores_summary}):
|
||||
<scores_summary>
|
||||
"""
|
||||
if failed_triggers:
|
||||
prompt += "FAILED TO TRIGGER (should have triggered but didn't):\n"
|
||||
for r in failed_triggers:
|
||||
prompt += f' - "{r["query"]}" (triggered {r["triggers"]}/{r["runs"]} times)\n'
|
||||
prompt += "\n"
|
||||
|
||||
if false_triggers:
|
||||
prompt += "FALSE TRIGGERS (triggered but shouldn't have):\n"
|
||||
for r in false_triggers:
|
||||
prompt += f' - "{r["query"]}" (triggered {r["triggers"]}/{r["runs"]} times)\n'
|
||||
prompt += "\n"
|
||||
|
||||
if history:
|
||||
prompt += "PREVIOUS ATTEMPTS (do NOT repeat these — try something structurally different):\n\n"
|
||||
for h in history:
|
||||
train_s = f"{h.get('train_passed', h.get('passed', 0))}/{h.get('train_total', h.get('total', 0))}"
|
||||
test_s = f"{h.get('test_passed', '?')}/{h.get('test_total', '?')}" if h.get('test_passed') is not None else None
|
||||
score_str = f"train={train_s}" + (f", test={test_s}" if test_s else "")
|
||||
prompt += f'<attempt {score_str}>\n'
|
||||
prompt += f'Description: "{h["description"]}"\n'
|
||||
if "results" in h:
|
||||
prompt += "Train results:\n"
|
||||
for r in h["results"]:
|
||||
status = "PASS" if r["pass"] else "FAIL"
|
||||
prompt += f' [{status}] "{r["query"][:80]}" (triggered {r["triggers"]}/{r["runs"]})\n'
|
||||
if h.get("note"):
|
||||
prompt += f'Note: {h["note"]}\n'
|
||||
prompt += "</attempt>\n\n"
|
||||
|
||||
prompt += f"""</scores_summary>
|
||||
|
||||
Skill content (for context on what the skill does):
|
||||
<skill_content>
|
||||
{skill_content}
|
||||
</skill_content>
|
||||
|
||||
Based on the failures, write a new and improved description that is more likely to trigger correctly. When I say "based on the failures", it's a bit of a tricky line to walk because we don't want to overfit to the specific cases you're seeing. So what I DON'T want you to do is produce an ever-expanding list of specific queries that this skill should or shouldn't trigger for. Instead, try to generalize from the failures to broader categories of user intent and situations where this skill would be useful or not useful. The reason for this is twofold:
|
||||
|
||||
1. Avoid overfitting
|
||||
2. The list might get loooong and it's injected into ALL queries and there might be a lot of skills, so we don't want to blow too much space on any given description.
|
||||
|
||||
Concretely, your description should not be more than about 100-200 words, even if that comes at the cost of accuracy.
|
||||
|
||||
Here are some tips that we've found to work well in writing these descriptions:
|
||||
- The skill should be phrased in the imperative -- "Use this skill for" rather than "this skill does"
|
||||
- The skill description should focus on the user's intent, what they are trying to achieve, vs. the implementation details of how the skill works.
|
||||
- The description competes with other skills for Claude's attention — make it distinctive and immediately recognizable.
|
||||
- If you're getting lots of failures after repeated attempts, change things up. Try different sentence structures or wordings.
|
||||
|
||||
I'd encourage you to be creative and mix up the style in different iterations since you'll have multiple opportunities to try different approaches and we'll just grab the highest-scoring one at the end.
|
||||
|
||||
Please respond with only the new description text in <new_description> tags, nothing else."""
|
||||
|
||||
response = client.messages.create(
|
||||
model=model,
|
||||
max_tokens=16000,
|
||||
thinking={
|
||||
"type": "enabled",
|
||||
"budget_tokens": 10000,
|
||||
},
|
||||
messages=[{"role": "user", "content": prompt}],
|
||||
)
|
||||
|
||||
# Extract thinking and text from response
|
||||
thinking_text = ""
|
||||
text = ""
|
||||
for block in response.content:
|
||||
if block.type == "thinking":
|
||||
thinking_text = block.thinking
|
||||
elif block.type == "text":
|
||||
text = block.text
|
||||
|
||||
# Parse out the <new_description> tags
|
||||
match = re.search(r"<new_description>(.*?)</new_description>", text, re.DOTALL)
|
||||
description = match.group(1).strip().strip('"') if match else text.strip().strip('"')
|
||||
|
||||
# Log the transcript
|
||||
transcript: dict = {
|
||||
"iteration": iteration,
|
||||
"prompt": prompt,
|
||||
"thinking": thinking_text,
|
||||
"response": text,
|
||||
"parsed_description": description,
|
||||
"char_count": len(description),
|
||||
"over_limit": len(description) > 1024,
|
||||
}
|
||||
|
||||
# If over 1024 chars, ask the model to shorten it
|
||||
if len(description) > 1024:
|
||||
shorten_prompt = f"Your description is {len(description)} characters, which exceeds the hard 1024 character limit. Please rewrite it to be under 1024 characters while preserving the most important trigger words and intent coverage. Respond with only the new description in <new_description> tags."
|
||||
shorten_response = client.messages.create(
|
||||
model=model,
|
||||
max_tokens=16000,
|
||||
thinking={
|
||||
"type": "enabled",
|
||||
"budget_tokens": 10000,
|
||||
},
|
||||
messages=[
|
||||
{"role": "user", "content": prompt},
|
||||
{"role": "assistant", "content": text},
|
||||
{"role": "user", "content": shorten_prompt},
|
||||
],
|
||||
)
|
||||
|
||||
shorten_thinking = ""
|
||||
shorten_text = ""
|
||||
for block in shorten_response.content:
|
||||
if block.type == "thinking":
|
||||
shorten_thinking = block.thinking
|
||||
elif block.type == "text":
|
||||
shorten_text = block.text
|
||||
|
||||
match = re.search(r"<new_description>(.*?)</new_description>", shorten_text, re.DOTALL)
|
||||
shortened = match.group(1).strip().strip('"') if match else shorten_text.strip().strip('"')
|
||||
|
||||
transcript["rewrite_prompt"] = shorten_prompt
|
||||
transcript["rewrite_thinking"] = shorten_thinking
|
||||
transcript["rewrite_response"] = shorten_text
|
||||
transcript["rewrite_description"] = shortened
|
||||
transcript["rewrite_char_count"] = len(shortened)
|
||||
description = shortened
|
||||
|
||||
transcript["final_description"] = description
|
||||
|
||||
if log_dir:
|
||||
log_dir.mkdir(parents=True, exist_ok=True)
|
||||
log_file = log_dir / f"improve_iter_{iteration or 'unknown'}.json"
|
||||
log_file.write_text(json.dumps(transcript, indent=2))
|
||||
|
||||
return description
|
||||
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(description="Improve a skill description based on eval results")
|
||||
parser.add_argument("--eval-results", required=True, help="Path to eval results JSON (from run_eval.py)")
|
||||
parser.add_argument("--skill-path", required=True, help="Path to skill directory")
|
||||
parser.add_argument("--history", default=None, help="Path to history JSON (previous attempts)")
|
||||
parser.add_argument("--model", required=True, help="Model for improvement")
|
||||
parser.add_argument("--verbose", action="store_true", help="Print thinking to stderr")
|
||||
args = parser.parse_args()
|
||||
|
||||
skill_path = Path(args.skill_path)
|
||||
if not (skill_path / "SKILL.md").exists():
|
||||
print(f"Error: No SKILL.md found at {skill_path}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
eval_results = json.loads(Path(args.eval_results).read_text())
|
||||
history = []
|
||||
if args.history:
|
||||
history = json.loads(Path(args.history).read_text())
|
||||
|
||||
name, _, content = parse_skill_md(skill_path)
|
||||
current_description = eval_results["description"]
|
||||
|
||||
if args.verbose:
|
||||
print(f"Current: {current_description}", file=sys.stderr)
|
||||
print(f"Score: {eval_results['summary']['passed']}/{eval_results['summary']['total']}", file=sys.stderr)
|
||||
|
||||
client = anthropic.Anthropic()
|
||||
new_description = improve_description(
|
||||
client=client,
|
||||
skill_name=name,
|
||||
skill_content=content,
|
||||
current_description=current_description,
|
||||
eval_results=eval_results,
|
||||
history=history,
|
||||
model=args.model,
|
||||
)
|
||||
|
||||
if args.verbose:
|
||||
print(f"Improved: {new_description}", file=sys.stderr)
|
||||
|
||||
# Output as JSON with both the new description and updated history
|
||||
output = {
|
||||
"description": new_description,
|
||||
"history": history + [{
|
||||
"description": current_description,
|
||||
"passed": eval_results["summary"]["passed"],
|
||||
"failed": eval_results["summary"]["failed"],
|
||||
"total": eval_results["summary"]["total"],
|
||||
"results": eval_results["results"],
|
||||
}],
|
||||
}
|
||||
print(json.dumps(output, indent=2))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
136
.skills/skill-creator/scripts/package_skill.py
Normal file
136
.skills/skill-creator/scripts/package_skill.py
Normal file
|
|
@ -0,0 +1,136 @@
|
|||
#!/usr/bin/env python3
|
||||
"""
|
||||
Skill Packager - Creates a distributable .skill file of a skill folder
|
||||
|
||||
Usage:
|
||||
python utils/package_skill.py <path/to/skill-folder> [output-directory]
|
||||
|
||||
Example:
|
||||
python utils/package_skill.py skills/public/my-skill
|
||||
python utils/package_skill.py skills/public/my-skill ./dist
|
||||
"""
|
||||
|
||||
import fnmatch
|
||||
import sys
|
||||
import zipfile
|
||||
from pathlib import Path
|
||||
from scripts.quick_validate import validate_skill
|
||||
|
||||
# Patterns to exclude when packaging skills.
|
||||
EXCLUDE_DIRS = {"__pycache__", "node_modules"}
|
||||
EXCLUDE_GLOBS = {"*.pyc"}
|
||||
EXCLUDE_FILES = {".DS_Store"}
|
||||
# Directories excluded only at the skill root (not when nested deeper).
|
||||
ROOT_EXCLUDE_DIRS = {"evals"}
|
||||
|
||||
|
||||
def should_exclude(rel_path: Path) -> bool:
|
||||
"""Check if a path should be excluded from packaging."""
|
||||
parts = rel_path.parts
|
||||
if any(part in EXCLUDE_DIRS for part in parts):
|
||||
return True
|
||||
# rel_path is relative to skill_path.parent, so parts[0] is the skill
|
||||
# folder name and parts[1] (if present) is the first subdir.
|
||||
if len(parts) > 1 and parts[1] in ROOT_EXCLUDE_DIRS:
|
||||
return True
|
||||
name = rel_path.name
|
||||
if name in EXCLUDE_FILES:
|
||||
return True
|
||||
return any(fnmatch.fnmatch(name, pat) for pat in EXCLUDE_GLOBS)
|
||||
|
||||
|
||||
def package_skill(skill_path, output_dir=None):
|
||||
"""
|
||||
Package a skill folder into a .skill file.
|
||||
|
||||
Args:
|
||||
skill_path: Path to the skill folder
|
||||
output_dir: Optional output directory for the .skill file (defaults to current directory)
|
||||
|
||||
Returns:
|
||||
Path to the created .skill file, or None if error
|
||||
"""
|
||||
skill_path = Path(skill_path).resolve()
|
||||
|
||||
# Validate skill folder exists
|
||||
if not skill_path.exists():
|
||||
print(f"❌ Error: Skill folder not found: {skill_path}")
|
||||
return None
|
||||
|
||||
if not skill_path.is_dir():
|
||||
print(f"❌ Error: Path is not a directory: {skill_path}")
|
||||
return None
|
||||
|
||||
# Validate SKILL.md exists
|
||||
skill_md = skill_path / "SKILL.md"
|
||||
if not skill_md.exists():
|
||||
print(f"❌ Error: SKILL.md not found in {skill_path}")
|
||||
return None
|
||||
|
||||
# Run validation before packaging
|
||||
print("🔍 Validating skill...")
|
||||
valid, message = validate_skill(skill_path)
|
||||
if not valid:
|
||||
print(f"❌ Validation failed: {message}")
|
||||
print(" Please fix the validation errors before packaging.")
|
||||
return None
|
||||
print(f"✅ {message}\n")
|
||||
|
||||
# Determine output location
|
||||
skill_name = skill_path.name
|
||||
if output_dir:
|
||||
output_path = Path(output_dir).resolve()
|
||||
output_path.mkdir(parents=True, exist_ok=True)
|
||||
else:
|
||||
output_path = Path.cwd()
|
||||
|
||||
skill_filename = output_path / f"{skill_name}.skill"
|
||||
|
||||
# Create the .skill file (zip format)
|
||||
try:
|
||||
with zipfile.ZipFile(skill_filename, 'w', zipfile.ZIP_DEFLATED) as zipf:
|
||||
# Walk through the skill directory, excluding build artifacts
|
||||
for file_path in skill_path.rglob('*'):
|
||||
if not file_path.is_file():
|
||||
continue
|
||||
arcname = file_path.relative_to(skill_path.parent)
|
||||
if should_exclude(arcname):
|
||||
print(f" Skipped: {arcname}")
|
||||
continue
|
||||
zipf.write(file_path, arcname)
|
||||
print(f" Added: {arcname}")
|
||||
|
||||
print(f"\n✅ Successfully packaged skill to: {skill_filename}")
|
||||
return skill_filename
|
||||
|
||||
except Exception as e:
|
||||
print(f"❌ Error creating .skill file: {e}")
|
||||
return None
|
||||
|
||||
|
||||
def main():
|
||||
if len(sys.argv) < 2:
|
||||
print("Usage: python utils/package_skill.py <path/to/skill-folder> [output-directory]")
|
||||
print("\nExample:")
|
||||
print(" python utils/package_skill.py skills/public/my-skill")
|
||||
print(" python utils/package_skill.py skills/public/my-skill ./dist")
|
||||
sys.exit(1)
|
||||
|
||||
skill_path = sys.argv[1]
|
||||
output_dir = sys.argv[2] if len(sys.argv) > 2 else None
|
||||
|
||||
print(f"📦 Packaging skill: {skill_path}")
|
||||
if output_dir:
|
||||
print(f" Output directory: {output_dir}")
|
||||
print()
|
||||
|
||||
result = package_skill(skill_path, output_dir)
|
||||
|
||||
if result:
|
||||
sys.exit(0)
|
||||
else:
|
||||
sys.exit(1)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
103
.skills/skill-creator/scripts/quick_validate.py
Normal file
103
.skills/skill-creator/scripts/quick_validate.py
Normal file
|
|
@ -0,0 +1,103 @@
|
|||
#!/usr/bin/env python3
|
||||
"""
|
||||
Quick validation script for skills - minimal version
|
||||
"""
|
||||
|
||||
import sys
|
||||
import os
|
||||
import re
|
||||
import yaml
|
||||
from pathlib import Path
|
||||
|
||||
def validate_skill(skill_path):
|
||||
"""Basic validation of a skill"""
|
||||
skill_path = Path(skill_path)
|
||||
|
||||
# Check SKILL.md exists
|
||||
skill_md = skill_path / 'SKILL.md'
|
||||
if not skill_md.exists():
|
||||
return False, "SKILL.md not found"
|
||||
|
||||
# Read and validate frontmatter
|
||||
content = skill_md.read_text()
|
||||
if not content.startswith('---'):
|
||||
return False, "No YAML frontmatter found"
|
||||
|
||||
# Extract frontmatter
|
||||
match = re.match(r'^---\n(.*?)\n---', content, re.DOTALL)
|
||||
if not match:
|
||||
return False, "Invalid frontmatter format"
|
||||
|
||||
frontmatter_text = match.group(1)
|
||||
|
||||
# Parse YAML frontmatter
|
||||
try:
|
||||
frontmatter = yaml.safe_load(frontmatter_text)
|
||||
if not isinstance(frontmatter, dict):
|
||||
return False, "Frontmatter must be a YAML dictionary"
|
||||
except yaml.YAMLError as e:
|
||||
return False, f"Invalid YAML in frontmatter: {e}"
|
||||
|
||||
# Define allowed properties
|
||||
ALLOWED_PROPERTIES = {'name', 'description', 'license', 'allowed-tools', 'metadata', 'compatibility'}
|
||||
|
||||
# Check for unexpected properties (excluding nested keys under metadata)
|
||||
unexpected_keys = set(frontmatter.keys()) - ALLOWED_PROPERTIES
|
||||
if unexpected_keys:
|
||||
return False, (
|
||||
f"Unexpected key(s) in SKILL.md frontmatter: {', '.join(sorted(unexpected_keys))}. "
|
||||
f"Allowed properties are: {', '.join(sorted(ALLOWED_PROPERTIES))}"
|
||||
)
|
||||
|
||||
# Check required fields
|
||||
if 'name' not in frontmatter:
|
||||
return False, "Missing 'name' in frontmatter"
|
||||
if 'description' not in frontmatter:
|
||||
return False, "Missing 'description' in frontmatter"
|
||||
|
||||
# Extract name for validation
|
||||
name = frontmatter.get('name', '')
|
||||
if not isinstance(name, str):
|
||||
return False, f"Name must be a string, got {type(name).__name__}"
|
||||
name = name.strip()
|
||||
if name:
|
||||
# Check naming convention (kebab-case: lowercase with hyphens)
|
||||
if not re.match(r'^[a-z0-9-]+$', name):
|
||||
return False, f"Name '{name}' should be kebab-case (lowercase letters, digits, and hyphens only)"
|
||||
if name.startswith('-') or name.endswith('-') or '--' in name:
|
||||
return False, f"Name '{name}' cannot start/end with hyphen or contain consecutive hyphens"
|
||||
# Check name length (max 64 characters per spec)
|
||||
if len(name) > 64:
|
||||
return False, f"Name is too long ({len(name)} characters). Maximum is 64 characters."
|
||||
|
||||
# Extract and validate description
|
||||
description = frontmatter.get('description', '')
|
||||
if not isinstance(description, str):
|
||||
return False, f"Description must be a string, got {type(description).__name__}"
|
||||
description = description.strip()
|
||||
if description:
|
||||
# Check for angle brackets
|
||||
if '<' in description or '>' in description:
|
||||
return False, "Description cannot contain angle brackets (< or >)"
|
||||
# Check description length (max 1024 characters per spec)
|
||||
if len(description) > 1024:
|
||||
return False, f"Description is too long ({len(description)} characters). Maximum is 1024 characters."
|
||||
|
||||
# Validate compatibility field if present (optional)
|
||||
compatibility = frontmatter.get('compatibility', '')
|
||||
if compatibility:
|
||||
if not isinstance(compatibility, str):
|
||||
return False, f"Compatibility must be a string, got {type(compatibility).__name__}"
|
||||
if len(compatibility) > 500:
|
||||
return False, f"Compatibility is too long ({len(compatibility)} characters). Maximum is 500 characters."
|
||||
|
||||
return True, "Skill is valid!"
|
||||
|
||||
if __name__ == "__main__":
|
||||
if len(sys.argv) != 2:
|
||||
print("Usage: python quick_validate.py <skill_directory>")
|
||||
sys.exit(1)
|
||||
|
||||
valid, message = validate_skill(sys.argv[1])
|
||||
print(message)
|
||||
sys.exit(0 if valid else 1)
|
||||
310
.skills/skill-creator/scripts/run_eval.py
Normal file
310
.skills/skill-creator/scripts/run_eval.py
Normal file
|
|
@ -0,0 +1,310 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Run trigger evaluation for a skill description.
|
||||
|
||||
Tests whether a skill's description causes Claude to trigger (read the skill)
|
||||
for a set of queries. Outputs results as JSON.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import select
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
import uuid
|
||||
from concurrent.futures import ProcessPoolExecutor, as_completed
|
||||
from pathlib import Path
|
||||
|
||||
from scripts.utils import parse_skill_md
|
||||
|
||||
|
||||
def find_project_root() -> Path:
|
||||
"""Find the project root by walking up from cwd looking for .claude/.
|
||||
|
||||
Mimics how Claude Code discovers its project root, so the command file
|
||||
we create ends up where claude -p will look for it.
|
||||
"""
|
||||
current = Path.cwd()
|
||||
for parent in [current, *current.parents]:
|
||||
if (parent / ".claude").is_dir():
|
||||
return parent
|
||||
return current
|
||||
|
||||
|
||||
def run_single_query(
|
||||
query: str,
|
||||
skill_name: str,
|
||||
skill_description: str,
|
||||
timeout: int,
|
||||
project_root: str,
|
||||
model: str | None = None,
|
||||
) -> bool:
|
||||
"""Run a single query and return whether the skill was triggered.
|
||||
|
||||
Creates a command file in .claude/commands/ so it appears in Claude's
|
||||
available_skills list, then runs `claude -p` with the raw query.
|
||||
Uses --include-partial-messages to detect triggering early from
|
||||
stream events (content_block_start) rather than waiting for the
|
||||
full assistant message, which only arrives after tool execution.
|
||||
"""
|
||||
unique_id = uuid.uuid4().hex[:8]
|
||||
clean_name = f"{skill_name}-skill-{unique_id}"
|
||||
project_commands_dir = Path(project_root) / ".claude" / "commands"
|
||||
command_file = project_commands_dir / f"{clean_name}.md"
|
||||
|
||||
try:
|
||||
project_commands_dir.mkdir(parents=True, exist_ok=True)
|
||||
# Use YAML block scalar to avoid breaking on quotes in description
|
||||
indented_desc = "\n ".join(skill_description.split("\n"))
|
||||
command_content = (
|
||||
f"---\n"
|
||||
f"description: |\n"
|
||||
f" {indented_desc}\n"
|
||||
f"---\n\n"
|
||||
f"# {skill_name}\n\n"
|
||||
f"This skill handles: {skill_description}\n"
|
||||
)
|
||||
command_file.write_text(command_content)
|
||||
|
||||
cmd = [
|
||||
"claude",
|
||||
"-p", query,
|
||||
"--output-format", "stream-json",
|
||||
"--verbose",
|
||||
"--include-partial-messages",
|
||||
]
|
||||
if model:
|
||||
cmd.extend(["--model", model])
|
||||
|
||||
# Remove CLAUDECODE env var to allow nesting claude -p inside a
|
||||
# Claude Code session. The guard is for interactive terminal conflicts;
|
||||
# programmatic subprocess usage is safe.
|
||||
env = {k: v for k, v in os.environ.items() if k != "CLAUDECODE"}
|
||||
|
||||
process = subprocess.Popen(
|
||||
cmd,
|
||||
stdout=subprocess.PIPE,
|
||||
stderr=subprocess.DEVNULL,
|
||||
cwd=project_root,
|
||||
env=env,
|
||||
)
|
||||
|
||||
triggered = False
|
||||
start_time = time.time()
|
||||
buffer = ""
|
||||
# Track state for stream event detection
|
||||
pending_tool_name = None
|
||||
accumulated_json = ""
|
||||
|
||||
try:
|
||||
while time.time() - start_time < timeout:
|
||||
if process.poll() is not None:
|
||||
remaining = process.stdout.read()
|
||||
if remaining:
|
||||
buffer += remaining.decode("utf-8", errors="replace")
|
||||
break
|
||||
|
||||
ready, _, _ = select.select([process.stdout], [], [], 1.0)
|
||||
if not ready:
|
||||
continue
|
||||
|
||||
chunk = os.read(process.stdout.fileno(), 8192)
|
||||
if not chunk:
|
||||
break
|
||||
buffer += chunk.decode("utf-8", errors="replace")
|
||||
|
||||
while "\n" in buffer:
|
||||
line, buffer = buffer.split("\n", 1)
|
||||
line = line.strip()
|
||||
if not line:
|
||||
continue
|
||||
|
||||
try:
|
||||
event = json.loads(line)
|
||||
except json.JSONDecodeError:
|
||||
continue
|
||||
|
||||
# Early detection via stream events
|
||||
if event.get("type") == "stream_event":
|
||||
se = event.get("event", {})
|
||||
se_type = se.get("type", "")
|
||||
|
||||
if se_type == "content_block_start":
|
||||
cb = se.get("content_block", {})
|
||||
if cb.get("type") == "tool_use":
|
||||
tool_name = cb.get("name", "")
|
||||
if tool_name in ("Skill", "Read"):
|
||||
pending_tool_name = tool_name
|
||||
accumulated_json = ""
|
||||
else:
|
||||
return False
|
||||
|
||||
elif se_type == "content_block_delta" and pending_tool_name:
|
||||
delta = se.get("delta", {})
|
||||
if delta.get("type") == "input_json_delta":
|
||||
accumulated_json += delta.get("partial_json", "")
|
||||
if clean_name in accumulated_json:
|
||||
return True
|
||||
|
||||
elif se_type in ("content_block_stop", "message_stop"):
|
||||
if pending_tool_name:
|
||||
return clean_name in accumulated_json
|
||||
if se_type == "message_stop":
|
||||
return False
|
||||
|
||||
# Fallback: full assistant message
|
||||
elif event.get("type") == "assistant":
|
||||
message = event.get("message", {})
|
||||
for content_item in message.get("content", []):
|
||||
if content_item.get("type") != "tool_use":
|
||||
continue
|
||||
tool_name = content_item.get("name", "")
|
||||
tool_input = content_item.get("input", {})
|
||||
if tool_name == "Skill" and clean_name in tool_input.get("skill", ""):
|
||||
triggered = True
|
||||
elif tool_name == "Read" and clean_name in tool_input.get("file_path", ""):
|
||||
triggered = True
|
||||
return triggered
|
||||
|
||||
elif event.get("type") == "result":
|
||||
return triggered
|
||||
finally:
|
||||
# Clean up process on any exit path (return, exception, timeout)
|
||||
if process.poll() is None:
|
||||
process.kill()
|
||||
process.wait()
|
||||
|
||||
return triggered
|
||||
finally:
|
||||
if command_file.exists():
|
||||
command_file.unlink()
|
||||
|
||||
|
||||
def run_eval(
|
||||
eval_set: list[dict],
|
||||
skill_name: str,
|
||||
description: str,
|
||||
num_workers: int,
|
||||
timeout: int,
|
||||
project_root: Path,
|
||||
runs_per_query: int = 1,
|
||||
trigger_threshold: float = 0.5,
|
||||
model: str | None = None,
|
||||
) -> dict:
|
||||
"""Run the full eval set and return results."""
|
||||
results = []
|
||||
|
||||
with ProcessPoolExecutor(max_workers=num_workers) as executor:
|
||||
future_to_info = {}
|
||||
for item in eval_set:
|
||||
for run_idx in range(runs_per_query):
|
||||
future = executor.submit(
|
||||
run_single_query,
|
||||
item["query"],
|
||||
skill_name,
|
||||
description,
|
||||
timeout,
|
||||
str(project_root),
|
||||
model,
|
||||
)
|
||||
future_to_info[future] = (item, run_idx)
|
||||
|
||||
query_triggers: dict[str, list[bool]] = {}
|
||||
query_items: dict[str, dict] = {}
|
||||
for future in as_completed(future_to_info):
|
||||
item, _ = future_to_info[future]
|
||||
query = item["query"]
|
||||
query_items[query] = item
|
||||
if query not in query_triggers:
|
||||
query_triggers[query] = []
|
||||
try:
|
||||
query_triggers[query].append(future.result())
|
||||
except Exception as e:
|
||||
print(f"Warning: query failed: {e}", file=sys.stderr)
|
||||
query_triggers[query].append(False)
|
||||
|
||||
for query, triggers in query_triggers.items():
|
||||
item = query_items[query]
|
||||
trigger_rate = sum(triggers) / len(triggers)
|
||||
should_trigger = item["should_trigger"]
|
||||
if should_trigger:
|
||||
did_pass = trigger_rate >= trigger_threshold
|
||||
else:
|
||||
did_pass = trigger_rate < trigger_threshold
|
||||
results.append({
|
||||
"query": query,
|
||||
"should_trigger": should_trigger,
|
||||
"trigger_rate": trigger_rate,
|
||||
"triggers": sum(triggers),
|
||||
"runs": len(triggers),
|
||||
"pass": did_pass,
|
||||
})
|
||||
|
||||
passed = sum(1 for r in results if r["pass"])
|
||||
total = len(results)
|
||||
|
||||
return {
|
||||
"skill_name": skill_name,
|
||||
"description": description,
|
||||
"results": results,
|
||||
"summary": {
|
||||
"total": total,
|
||||
"passed": passed,
|
||||
"failed": total - passed,
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(description="Run trigger evaluation for a skill description")
|
||||
parser.add_argument("--eval-set", required=True, help="Path to eval set JSON file")
|
||||
parser.add_argument("--skill-path", required=True, help="Path to skill directory")
|
||||
parser.add_argument("--description", default=None, help="Override description to test")
|
||||
parser.add_argument("--num-workers", type=int, default=10, help="Number of parallel workers")
|
||||
parser.add_argument("--timeout", type=int, default=30, help="Timeout per query in seconds")
|
||||
parser.add_argument("--runs-per-query", type=int, default=3, help="Number of runs per query")
|
||||
parser.add_argument("--trigger-threshold", type=float, default=0.5, help="Trigger rate threshold")
|
||||
parser.add_argument("--model", default=None, help="Model to use for claude -p (default: user's configured model)")
|
||||
parser.add_argument("--verbose", action="store_true", help="Print progress to stderr")
|
||||
args = parser.parse_args()
|
||||
|
||||
eval_set = json.loads(Path(args.eval_set).read_text())
|
||||
skill_path = Path(args.skill_path)
|
||||
|
||||
if not (skill_path / "SKILL.md").exists():
|
||||
print(f"Error: No SKILL.md found at {skill_path}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
name, original_description, content = parse_skill_md(skill_path)
|
||||
description = args.description or original_description
|
||||
project_root = find_project_root()
|
||||
|
||||
if args.verbose:
|
||||
print(f"Evaluating: {description}", file=sys.stderr)
|
||||
|
||||
output = run_eval(
|
||||
eval_set=eval_set,
|
||||
skill_name=name,
|
||||
description=description,
|
||||
num_workers=args.num_workers,
|
||||
timeout=args.timeout,
|
||||
project_root=project_root,
|
||||
runs_per_query=args.runs_per_query,
|
||||
trigger_threshold=args.trigger_threshold,
|
||||
model=args.model,
|
||||
)
|
||||
|
||||
if args.verbose:
|
||||
summary = output["summary"]
|
||||
print(f"Results: {summary['passed']}/{summary['total']} passed", file=sys.stderr)
|
||||
for r in output["results"]:
|
||||
status = "PASS" if r["pass"] else "FAIL"
|
||||
rate_str = f"{r['triggers']}/{r['runs']}"
|
||||
print(f" [{status}] rate={rate_str} expected={r['should_trigger']}: {r['query'][:70]}", file=sys.stderr)
|
||||
|
||||
print(json.dumps(output, indent=2))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
332
.skills/skill-creator/scripts/run_loop.py
Normal file
332
.skills/skill-creator/scripts/run_loop.py
Normal file
|
|
@ -0,0 +1,332 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Run the eval + improve loop until all pass or max iterations reached.
|
||||
|
||||
Combines run_eval.py and improve_description.py in a loop, tracking history
|
||||
and returning the best description found. Supports train/test split to prevent
|
||||
overfitting.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import random
|
||||
import sys
|
||||
import tempfile
|
||||
import time
|
||||
import webbrowser
|
||||
from pathlib import Path
|
||||
|
||||
import anthropic
|
||||
|
||||
from scripts.generate_report import generate_html
|
||||
from scripts.improve_description import improve_description
|
||||
from scripts.run_eval import find_project_root, run_eval
|
||||
from scripts.utils import parse_skill_md
|
||||
|
||||
|
||||
def split_eval_set(eval_set: list[dict], holdout: float, seed: int = 42) -> tuple[list[dict], list[dict]]:
|
||||
"""Split eval set into train and test sets, stratified by should_trigger."""
|
||||
random.seed(seed)
|
||||
|
||||
# Separate by should_trigger
|
||||
trigger = [e for e in eval_set if e["should_trigger"]]
|
||||
no_trigger = [e for e in eval_set if not e["should_trigger"]]
|
||||
|
||||
# Shuffle each group
|
||||
random.shuffle(trigger)
|
||||
random.shuffle(no_trigger)
|
||||
|
||||
# Calculate split points
|
||||
n_trigger_test = max(1, int(len(trigger) * holdout))
|
||||
n_no_trigger_test = max(1, int(len(no_trigger) * holdout))
|
||||
|
||||
# Split
|
||||
test_set = trigger[:n_trigger_test] + no_trigger[:n_no_trigger_test]
|
||||
train_set = trigger[n_trigger_test:] + no_trigger[n_no_trigger_test:]
|
||||
|
||||
return train_set, test_set
|
||||
|
||||
|
||||
def run_loop(
|
||||
eval_set: list[dict],
|
||||
skill_path: Path,
|
||||
description_override: str | None,
|
||||
num_workers: int,
|
||||
timeout: int,
|
||||
max_iterations: int,
|
||||
runs_per_query: int,
|
||||
trigger_threshold: float,
|
||||
holdout: float,
|
||||
model: str,
|
||||
verbose: bool,
|
||||
live_report_path: Path | None = None,
|
||||
log_dir: Path | None = None,
|
||||
) -> dict:
|
||||
"""Run the eval + improvement loop."""
|
||||
project_root = find_project_root()
|
||||
name, original_description, content = parse_skill_md(skill_path)
|
||||
current_description = description_override or original_description
|
||||
|
||||
# Split into train/test if holdout > 0
|
||||
if holdout > 0:
|
||||
train_set, test_set = split_eval_set(eval_set, holdout)
|
||||
if verbose:
|
||||
print(f"Split: {len(train_set)} train, {len(test_set)} test (holdout={holdout})", file=sys.stderr)
|
||||
else:
|
||||
train_set = eval_set
|
||||
test_set = []
|
||||
|
||||
client = anthropic.Anthropic()
|
||||
history = []
|
||||
exit_reason = "unknown"
|
||||
|
||||
for iteration in range(1, max_iterations + 1):
|
||||
if verbose:
|
||||
print(f"\n{'='*60}", file=sys.stderr)
|
||||
print(f"Iteration {iteration}/{max_iterations}", file=sys.stderr)
|
||||
print(f"Description: {current_description}", file=sys.stderr)
|
||||
print(f"{'='*60}", file=sys.stderr)
|
||||
|
||||
# Evaluate train + test together in one batch for parallelism
|
||||
all_queries = train_set + test_set
|
||||
t0 = time.time()
|
||||
all_results = run_eval(
|
||||
eval_set=all_queries,
|
||||
skill_name=name,
|
||||
description=current_description,
|
||||
num_workers=num_workers,
|
||||
timeout=timeout,
|
||||
project_root=project_root,
|
||||
runs_per_query=runs_per_query,
|
||||
trigger_threshold=trigger_threshold,
|
||||
model=model,
|
||||
)
|
||||
eval_elapsed = time.time() - t0
|
||||
|
||||
# Split results back into train/test by matching queries
|
||||
train_queries_set = {q["query"] for q in train_set}
|
||||
train_result_list = [r for r in all_results["results"] if r["query"] in train_queries_set]
|
||||
test_result_list = [r for r in all_results["results"] if r["query"] not in train_queries_set]
|
||||
|
||||
train_passed = sum(1 for r in train_result_list if r["pass"])
|
||||
train_total = len(train_result_list)
|
||||
train_summary = {"passed": train_passed, "failed": train_total - train_passed, "total": train_total}
|
||||
train_results = {"results": train_result_list, "summary": train_summary}
|
||||
|
||||
if test_set:
|
||||
test_passed = sum(1 for r in test_result_list if r["pass"])
|
||||
test_total = len(test_result_list)
|
||||
test_summary = {"passed": test_passed, "failed": test_total - test_passed, "total": test_total}
|
||||
test_results = {"results": test_result_list, "summary": test_summary}
|
||||
else:
|
||||
test_results = None
|
||||
test_summary = None
|
||||
|
||||
history.append({
|
||||
"iteration": iteration,
|
||||
"description": current_description,
|
||||
"train_passed": train_summary["passed"],
|
||||
"train_failed": train_summary["failed"],
|
||||
"train_total": train_summary["total"],
|
||||
"train_results": train_results["results"],
|
||||
"test_passed": test_summary["passed"] if test_summary else None,
|
||||
"test_failed": test_summary["failed"] if test_summary else None,
|
||||
"test_total": test_summary["total"] if test_summary else None,
|
||||
"test_results": test_results["results"] if test_results else None,
|
||||
# For backward compat with report generator
|
||||
"passed": train_summary["passed"],
|
||||
"failed": train_summary["failed"],
|
||||
"total": train_summary["total"],
|
||||
"results": train_results["results"],
|
||||
})
|
||||
|
||||
# Write live report if path provided
|
||||
if live_report_path:
|
||||
partial_output = {
|
||||
"original_description": original_description,
|
||||
"best_description": current_description,
|
||||
"best_score": "in progress",
|
||||
"iterations_run": len(history),
|
||||
"holdout": holdout,
|
||||
"train_size": len(train_set),
|
||||
"test_size": len(test_set),
|
||||
"history": history,
|
||||
}
|
||||
live_report_path.write_text(generate_html(partial_output, auto_refresh=True, skill_name=name))
|
||||
|
||||
if verbose:
|
||||
def print_eval_stats(label, results, elapsed):
|
||||
pos = [r for r in results if r["should_trigger"]]
|
||||
neg = [r for r in results if not r["should_trigger"]]
|
||||
tp = sum(r["triggers"] for r in pos)
|
||||
pos_runs = sum(r["runs"] for r in pos)
|
||||
fn = pos_runs - tp
|
||||
fp = sum(r["triggers"] for r in neg)
|
||||
neg_runs = sum(r["runs"] for r in neg)
|
||||
tn = neg_runs - fp
|
||||
total = tp + tn + fp + fn
|
||||
precision = tp / (tp + fp) if (tp + fp) > 0 else 1.0
|
||||
recall = tp / (tp + fn) if (tp + fn) > 0 else 1.0
|
||||
accuracy = (tp + tn) / total if total > 0 else 0.0
|
||||
print(f"{label}: {tp+tn}/{total} correct, precision={precision:.0%} recall={recall:.0%} accuracy={accuracy:.0%} ({elapsed:.1f}s)", file=sys.stderr)
|
||||
for r in results:
|
||||
status = "PASS" if r["pass"] else "FAIL"
|
||||
rate_str = f"{r['triggers']}/{r['runs']}"
|
||||
print(f" [{status}] rate={rate_str} expected={r['should_trigger']}: {r['query'][:60]}", file=sys.stderr)
|
||||
|
||||
print_eval_stats("Train", train_results["results"], eval_elapsed)
|
||||
if test_summary:
|
||||
print_eval_stats("Test ", test_results["results"], 0)
|
||||
|
||||
if train_summary["failed"] == 0:
|
||||
exit_reason = f"all_passed (iteration {iteration})"
|
||||
if verbose:
|
||||
print(f"\nAll train queries passed on iteration {iteration}!", file=sys.stderr)
|
||||
break
|
||||
|
||||
if iteration == max_iterations:
|
||||
exit_reason = f"max_iterations ({max_iterations})"
|
||||
if verbose:
|
||||
print(f"\nMax iterations reached ({max_iterations}).", file=sys.stderr)
|
||||
break
|
||||
|
||||
# Improve the description based on train results
|
||||
if verbose:
|
||||
print(f"\nImproving description...", file=sys.stderr)
|
||||
|
||||
t0 = time.time()
|
||||
# Strip test scores from history so improvement model can't see them
|
||||
blinded_history = [
|
||||
{k: v for k, v in h.items() if not k.startswith("test_")}
|
||||
for h in history
|
||||
]
|
||||
new_description = improve_description(
|
||||
client=client,
|
||||
skill_name=name,
|
||||
skill_content=content,
|
||||
current_description=current_description,
|
||||
eval_results=train_results,
|
||||
history=blinded_history,
|
||||
model=model,
|
||||
log_dir=log_dir,
|
||||
iteration=iteration,
|
||||
)
|
||||
improve_elapsed = time.time() - t0
|
||||
|
||||
if verbose:
|
||||
print(f"Proposed ({improve_elapsed:.1f}s): {new_description}", file=sys.stderr)
|
||||
|
||||
current_description = new_description
|
||||
|
||||
# Find the best iteration by TEST score (or train if no test set)
|
||||
if test_set:
|
||||
best = max(history, key=lambda h: h["test_passed"] or 0)
|
||||
best_score = f"{best['test_passed']}/{best['test_total']}"
|
||||
else:
|
||||
best = max(history, key=lambda h: h["train_passed"])
|
||||
best_score = f"{best['train_passed']}/{best['train_total']}"
|
||||
|
||||
if verbose:
|
||||
print(f"\nExit reason: {exit_reason}", file=sys.stderr)
|
||||
print(f"Best score: {best_score} (iteration {best['iteration']})", file=sys.stderr)
|
||||
|
||||
return {
|
||||
"exit_reason": exit_reason,
|
||||
"original_description": original_description,
|
||||
"best_description": best["description"],
|
||||
"best_score": best_score,
|
||||
"best_train_score": f"{best['train_passed']}/{best['train_total']}",
|
||||
"best_test_score": f"{best['test_passed']}/{best['test_total']}" if test_set else None,
|
||||
"final_description": current_description,
|
||||
"iterations_run": len(history),
|
||||
"holdout": holdout,
|
||||
"train_size": len(train_set),
|
||||
"test_size": len(test_set),
|
||||
"history": history,
|
||||
}
|
||||
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(description="Run eval + improve loop")
|
||||
parser.add_argument("--eval-set", required=True, help="Path to eval set JSON file")
|
||||
parser.add_argument("--skill-path", required=True, help="Path to skill directory")
|
||||
parser.add_argument("--description", default=None, help="Override starting description")
|
||||
parser.add_argument("--num-workers", type=int, default=10, help="Number of parallel workers")
|
||||
parser.add_argument("--timeout", type=int, default=30, help="Timeout per query in seconds")
|
||||
parser.add_argument("--max-iterations", type=int, default=5, help="Max improvement iterations")
|
||||
parser.add_argument("--runs-per-query", type=int, default=3, help="Number of runs per query")
|
||||
parser.add_argument("--trigger-threshold", type=float, default=0.5, help="Trigger rate threshold")
|
||||
parser.add_argument("--holdout", type=float, default=0.4, help="Fraction of eval set to hold out for testing (0 to disable)")
|
||||
parser.add_argument("--model", required=True, help="Model for improvement")
|
||||
parser.add_argument("--verbose", action="store_true", help="Print progress to stderr")
|
||||
parser.add_argument("--report", default="auto", help="Generate HTML report at this path (default: 'auto' for temp file, 'none' to disable)")
|
||||
parser.add_argument("--results-dir", default=None, help="Save all outputs (results.json, report.html, log.txt) to a timestamped subdirectory here")
|
||||
args = parser.parse_args()
|
||||
|
||||
eval_set = json.loads(Path(args.eval_set).read_text())
|
||||
skill_path = Path(args.skill_path)
|
||||
|
||||
if not (skill_path / "SKILL.md").exists():
|
||||
print(f"Error: No SKILL.md found at {skill_path}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
name, _, _ = parse_skill_md(skill_path)
|
||||
|
||||
# Set up live report path
|
||||
if args.report != "none":
|
||||
if args.report == "auto":
|
||||
timestamp = time.strftime("%Y%m%d_%H%M%S")
|
||||
live_report_path = Path(tempfile.gettempdir()) / f"skill_description_report_{skill_path.name}_{timestamp}.html"
|
||||
else:
|
||||
live_report_path = Path(args.report)
|
||||
# Open the report immediately so the user can watch
|
||||
live_report_path.write_text("<html><body><h1>Starting optimization loop...</h1><meta http-equiv='refresh' content='5'></body></html>")
|
||||
webbrowser.open(str(live_report_path))
|
||||
else:
|
||||
live_report_path = None
|
||||
|
||||
# Determine output directory (create before run_loop so logs can be written)
|
||||
if args.results_dir:
|
||||
timestamp = time.strftime("%Y-%m-%d_%H%M%S")
|
||||
results_dir = Path(args.results_dir) / timestamp
|
||||
results_dir.mkdir(parents=True, exist_ok=True)
|
||||
else:
|
||||
results_dir = None
|
||||
|
||||
log_dir = results_dir / "logs" if results_dir else None
|
||||
|
||||
output = run_loop(
|
||||
eval_set=eval_set,
|
||||
skill_path=skill_path,
|
||||
description_override=args.description,
|
||||
num_workers=args.num_workers,
|
||||
timeout=args.timeout,
|
||||
max_iterations=args.max_iterations,
|
||||
runs_per_query=args.runs_per_query,
|
||||
trigger_threshold=args.trigger_threshold,
|
||||
holdout=args.holdout,
|
||||
model=args.model,
|
||||
verbose=args.verbose,
|
||||
live_report_path=live_report_path,
|
||||
log_dir=log_dir,
|
||||
)
|
||||
|
||||
# Save JSON output
|
||||
json_output = json.dumps(output, indent=2)
|
||||
print(json_output)
|
||||
if results_dir:
|
||||
(results_dir / "results.json").write_text(json_output)
|
||||
|
||||
# Write final HTML report (without auto-refresh)
|
||||
if live_report_path:
|
||||
live_report_path.write_text(generate_html(output, auto_refresh=False, skill_name=name))
|
||||
print(f"\nReport: {live_report_path}", file=sys.stderr)
|
||||
|
||||
if results_dir and live_report_path:
|
||||
(results_dir / "report.html").write_text(generate_html(output, auto_refresh=False, skill_name=name))
|
||||
|
||||
if results_dir:
|
||||
print(f"Results saved to: {results_dir}", file=sys.stderr)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
47
.skills/skill-creator/scripts/utils.py
Normal file
47
.skills/skill-creator/scripts/utils.py
Normal file
|
|
@ -0,0 +1,47 @@
|
|||
"""Shared utilities for skill-creator scripts."""
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
|
||||
def parse_skill_md(skill_path: Path) -> tuple[str, str, str]:
|
||||
"""Parse a SKILL.md file, returning (name, description, full_content)."""
|
||||
content = (skill_path / "SKILL.md").read_text()
|
||||
lines = content.split("\n")
|
||||
|
||||
if lines[0].strip() != "---":
|
||||
raise ValueError("SKILL.md missing frontmatter (no opening ---)")
|
||||
|
||||
end_idx = None
|
||||
for i, line in enumerate(lines[1:], start=1):
|
||||
if line.strip() == "---":
|
||||
end_idx = i
|
||||
break
|
||||
|
||||
if end_idx is None:
|
||||
raise ValueError("SKILL.md missing frontmatter (no closing ---)")
|
||||
|
||||
name = ""
|
||||
description = ""
|
||||
frontmatter_lines = lines[1:end_idx]
|
||||
i = 0
|
||||
while i < len(frontmatter_lines):
|
||||
line = frontmatter_lines[i]
|
||||
if line.startswith("name:"):
|
||||
name = line[len("name:"):].strip().strip('"').strip("'")
|
||||
elif line.startswith("description:"):
|
||||
value = line[len("description:"):].strip()
|
||||
# Handle YAML multiline indicators (>, |, >-, |-)
|
||||
if value in (">", "|", ">-", "|-"):
|
||||
continuation_lines: list[str] = []
|
||||
i += 1
|
||||
while i < len(frontmatter_lines) and (frontmatter_lines[i].startswith(" ") or frontmatter_lines[i].startswith("\t")):
|
||||
continuation_lines.append(frontmatter_lines[i].strip())
|
||||
i += 1
|
||||
description = " ".join(continuation_lines)
|
||||
continue
|
||||
else:
|
||||
description = value.strip('"').strip("'")
|
||||
i += 1
|
||||
|
||||
return name, description, content
|
||||
97
.skills/writing-skills/SKILL.md
Normal file
97
.skills/writing-skills/SKILL.md
Normal file
|
|
@ -0,0 +1,97 @@
|
|||
---
|
||||
name: writing-skills
|
||||
description: "Use when creating, updating, or improving agent skills."
|
||||
---
|
||||
|
||||
# Writing Skills (Excellence)
|
||||
|
||||
Dispatcher for skill creation excellence. Use the decision tree below to find the right template and standards.
|
||||
|
||||
## Quick Decision Tree
|
||||
|
||||
### What do you need to do?
|
||||
|
||||
1. **Create a NEW skill:**
|
||||
- Is it simple (single file, <200 lines)? -> [Tier 1 Architecture](references/tier-1-simple/README.md)
|
||||
- Is it complex (multi-concept, 200-1000 lines)? -> [Tier 2 Architecture](references/tier-2-expanded/README.md)
|
||||
- Is it a massive platform (10+ products, AWS, Convex)? -> [Tier 3 Architecture](references/tier-3-platform/README.md)
|
||||
|
||||
2. **Improve an EXISTING skill:**
|
||||
- Fix "it's too long" -> [Modularize (Tier 3)](references/templates/tier-3-platform.md)
|
||||
- Fix "AI ignores rules" -> [Anti-Rationalization](references/anti-rationalization/README.md)
|
||||
- Fix "users can't find it" -> [CSO (Search Optimization)](references/cso/README.md)
|
||||
|
||||
3. **Verify Compliance:**
|
||||
- Check metadata/naming -> [Standards](references/standards/README.md)
|
||||
- Add tests -> [Testing Guide](references/testing/README.md)
|
||||
|
||||
## Component Index
|
||||
|
||||
| Component | Purpose |
|
||||
|-----------|---------|
|
||||
| **[CSO](references/cso/README.md)** | "SEO for LLMs". How to write descriptions that trigger. |
|
||||
| **[Standards](references/standards/README.md)** | File naming, YAML frontmatter, directory structure. |
|
||||
| **[Anti-Rationalization](references/anti-rationalization/README.md)**| How to write rules that agents won't ignore. |
|
||||
| **[Testing](references/testing/README.md)** | How to ensure your skill actually works. |
|
||||
|
||||
## Templates
|
||||
|
||||
- [Technique Skill](references/templates/technique.md) (How-to)
|
||||
- [Reference Skill](references/templates/reference.md) (Docs)
|
||||
- [Discipline Skill](references/templates/discipline.md) (Rules)
|
||||
- [Pattern Skill](references/templates/pattern.md) (Design Patterns)
|
||||
|
||||
## When to Use
|
||||
- Creating a NEW skill from scratch
|
||||
- Improving an EXISTING skill that agents ignore
|
||||
- Debugging why a skill isn't being triggered
|
||||
- Standardizing skills across a team
|
||||
|
||||
## How It Works
|
||||
|
||||
1. **Identify goal** -> Use decision tree above
|
||||
2. **Select template** -> From `references/templates/`
|
||||
3. **Apply CSO** -> Optimize description for discovery
|
||||
4. **Add anti-rationalization** -> For discipline skills
|
||||
5. **Test** -> RED-GREEN-REFACTOR cycle
|
||||
|
||||
## Quick Example
|
||||
|
||||
```yaml
|
||||
---
|
||||
name: my-technique
|
||||
description: Use when [specific symptom occurs].
|
||||
metadata:
|
||||
category: technique
|
||||
triggers: error-text, symptom, tool-name
|
||||
---
|
||||
|
||||
# My Technique
|
||||
|
||||
## When to Use
|
||||
- [Symptom A]
|
||||
- [Error message]
|
||||
```
|
||||
|
||||
## Common Mistakes
|
||||
|
||||
| Mistake | Fix |
|
||||
|---------|-----|
|
||||
| Description summarizes workflow | Use "Use when..." triggers only |
|
||||
| No `metadata.triggers` | Add 3+ keywords |
|
||||
| Generic name ("helper") | Use gerund (`creating-skills`) |
|
||||
| Long monolithic SKILL.md | Split into `references/` |
|
||||
|
||||
See [gotchas.md](gotchas.md) for more.
|
||||
|
||||
## Pre-Deploy Checklist
|
||||
|
||||
Before deploying any skill:
|
||||
|
||||
- [ ] `name` field matches directory name exactly
|
||||
- [ ] `SKILL.md` filename is ALL CAPS
|
||||
- [ ] Description starts with "Use when..."
|
||||
- [ ] `metadata.triggers` has 3+ keywords
|
||||
- [ ] Total lines < 500 (use `references/` for more)
|
||||
- [ ] No `@` force-loading in cross-references
|
||||
- [ ] Tested with real scenarios
|
||||
236
.skills/writing-skills/examples.md
Normal file
236
.skills/writing-skills/examples.md
Normal file
|
|
@ -0,0 +1,236 @@
|
|||
# Skill Templates & Examples
|
||||
|
||||
Complete, copy-paste templates for each skill type.
|
||||
|
||||
---
|
||||
|
||||
## Template: Technique Skill
|
||||
|
||||
For how-to guides that teach a specific method.
|
||||
|
||||
```markdown
|
||||
---
|
||||
name: technique-name
|
||||
description: >-
|
||||
Use when [specific symptom].
|
||||
metadata:
|
||||
category: technique
|
||||
triggers: error-text, symptom, tool-name
|
||||
---
|
||||
|
||||
# Technique Name
|
||||
|
||||
## Overview
|
||||
|
||||
[1-2 sentence core principle]
|
||||
|
||||
## When to Use
|
||||
|
||||
- [Symptom A]
|
||||
- [Symptom B]
|
||||
- [Error message text]
|
||||
|
||||
**NOT for:**
|
||||
- [When to avoid]
|
||||
|
||||
## The Problem
|
||||
|
||||
\`\`\`javascript
|
||||
// Bad example
|
||||
function badCode() {
|
||||
// problematic pattern
|
||||
}
|
||||
\`\`\`
|
||||
|
||||
## The Solution
|
||||
|
||||
\`\`\`javascript
|
||||
// Good example
|
||||
function goodCode() {
|
||||
// improved pattern
|
||||
}
|
||||
\`\`\`
|
||||
|
||||
## Step-by-Step
|
||||
|
||||
1. [First step]
|
||||
2. [Second step]
|
||||
3. [Final step]
|
||||
|
||||
## Quick Reference
|
||||
|
||||
| Scenario | Approach |
|
||||
|----------|----------|
|
||||
| Case A | Solution A |
|
||||
| Case B | Solution B |
|
||||
|
||||
## Common Mistakes
|
||||
|
||||
**Mistake 1:** [Description]
|
||||
- Wrong: \`bad code\`
|
||||
- Right: \`good code\`
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Template: Reference Skill
|
||||
|
||||
For documentation, APIs, and lookup tables.
|
||||
|
||||
```markdown
|
||||
---
|
||||
name: reference-name
|
||||
description: >-
|
||||
Use when working with [domain].
|
||||
metadata:
|
||||
category: reference
|
||||
triggers: tool, api, specific-terms
|
||||
---
|
||||
|
||||
# Reference Name
|
||||
|
||||
## Quick Reference
|
||||
|
||||
| Command | Purpose |
|
||||
|---------|---------|
|
||||
| \`cmd1\` | Does X |
|
||||
| \`cmd2\` | Does Y |
|
||||
|
||||
## Common Patterns
|
||||
|
||||
**Pattern A:**
|
||||
\`\`\`bash
|
||||
example command
|
||||
\`\`\`
|
||||
|
||||
**Pattern B:**
|
||||
\`\`\`bash
|
||||
another example
|
||||
\`\`\`
|
||||
|
||||
## Detailed Docs
|
||||
|
||||
For more options, run \`--help\` or see:
|
||||
- patterns.md
|
||||
- [examples.md](examples.md)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Template: Discipline Skill
|
||||
|
||||
For rules that agents must follow. Requires anti-rationalization techniques.
|
||||
|
||||
```markdown
|
||||
---
|
||||
name: discipline-name
|
||||
description: >-
|
||||
Use when [BEFORE violation].
|
||||
metadata:
|
||||
category: discipline
|
||||
triggers: new feature, code change, implementation
|
||||
---
|
||||
|
||||
# Rule Name
|
||||
|
||||
## Iron Law
|
||||
|
||||
**[SINGLE SENTENCE ABSOLUTE RULE]**
|
||||
|
||||
Violating the letter IS violating the spirit.
|
||||
|
||||
## The Rule
|
||||
|
||||
1. ALWAYS [step 1]
|
||||
2. NEVER [step 2]
|
||||
3. [Step 3]
|
||||
|
||||
## Violations
|
||||
|
||||
[Action before rule]? **Delete it. Start over.**
|
||||
|
||||
**No exceptions:**
|
||||
- Don't keep it as "reference"
|
||||
- Don't "adapt" it
|
||||
- Delete means delete
|
||||
|
||||
## Common Rationalizations
|
||||
|
||||
| Excuse | Reality |
|
||||
|--------|---------|
|
||||
| "Too simple" | Simple code breaks. Rule takes 30 seconds. |
|
||||
| "I'll do it after" | After = never. Do it now. |
|
||||
| "Spirit not ritual" | The ritual IS the spirit. |
|
||||
|
||||
## Red Flags - STOP
|
||||
|
||||
- [Flag 1]
|
||||
- [Flag 2]
|
||||
- "This is different because..."
|
||||
|
||||
**All mean:** Delete. Start over.
|
||||
|
||||
## Valid Exceptions
|
||||
|
||||
- [Exception 1]
|
||||
- [Exception 2]
|
||||
|
||||
**Everything else:** Follow the rule.
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Template: Pattern Skill
|
||||
|
||||
For mental models and design patterns.
|
||||
|
||||
```markdown
|
||||
---
|
||||
name: pattern-name
|
||||
description: >-
|
||||
Use when [recognizable symptom].
|
||||
metadata:
|
||||
category: pattern
|
||||
triggers: complexity, hard-to-follow, nested
|
||||
---
|
||||
|
||||
# Pattern Name
|
||||
|
||||
## The Pattern
|
||||
|
||||
[1-2 sentence core idea]
|
||||
|
||||
## Recognition Signs
|
||||
|
||||
- [Sign that pattern applies]
|
||||
- [Another sign]
|
||||
- [Code smell]
|
||||
|
||||
## Before
|
||||
|
||||
\`\`\`typescript
|
||||
// Complex/problematic
|
||||
function before() {
|
||||
// nested, confusing
|
||||
}
|
||||
\`\`\`
|
||||
|
||||
## After
|
||||
|
||||
\`\`\`typescript
|
||||
// Clean/improved
|
||||
function after() {
|
||||
// flat, clear
|
||||
}
|
||||
\`\`\`
|
||||
|
||||
## When NOT to Use
|
||||
|
||||
- [Over-engineering case]
|
||||
- [Simple case that doesn't need it]
|
||||
|
||||
## Impact
|
||||
|
||||
**Before:** [Problem metric]
|
||||
**After:** [Improved metric]
|
||||
```
|
||||
175
.skills/writing-skills/gotchas.md
Normal file
175
.skills/writing-skills/gotchas.md
Normal file
|
|
@ -0,0 +1,175 @@
|
|||
---
|
||||
description: Common pitfalls and tribal knowledge for skill creation.
|
||||
metadata:
|
||||
tags: [gotchas, troubleshooting, mistakes]
|
||||
---
|
||||
|
||||
# Skill Writing Gotchas
|
||||
|
||||
Tribal knowledge to avoid common mistakes.
|
||||
|
||||
## YAML Frontmatter
|
||||
|
||||
### Invalid Syntax
|
||||
|
||||
```yaml
|
||||
# BAD: Mixed list and map
|
||||
metadata:
|
||||
references:
|
||||
triggers: a, b, c
|
||||
- item1
|
||||
- item2
|
||||
|
||||
# GOOD: Consistent structure
|
||||
metadata:
|
||||
triggers: a, b, c
|
||||
references:
|
||||
- item1
|
||||
- item2
|
||||
```
|
||||
|
||||
### Multiline Description
|
||||
|
||||
```yaml
|
||||
# BAD: Line breaks create parsing errors
|
||||
description: Use when creating skills.
|
||||
Also for updating.
|
||||
|
||||
# GOOD: Use YAML multiline syntax
|
||||
description: >-
|
||||
Use when creating or updating skills.
|
||||
Triggers: new skill, update skill
|
||||
```
|
||||
|
||||
## Naming
|
||||
|
||||
### Directory Must Match `name` Field
|
||||
|
||||
```
|
||||
# BAD
|
||||
directory: my-skill/
|
||||
name: mySkill # Mismatch!
|
||||
|
||||
# GOOD
|
||||
directory: my-skill/
|
||||
name: my-skill # Exact match
|
||||
```
|
||||
|
||||
### SKILL.md Must Be ALL CAPS
|
||||
|
||||
```
|
||||
# BAD
|
||||
skill.md
|
||||
Skill.md
|
||||
|
||||
# GOOD
|
||||
SKILL.md
|
||||
```
|
||||
|
||||
## Discovery
|
||||
|
||||
### Description = Triggers, NOT Workflow
|
||||
|
||||
```yaml
|
||||
# BAD: Agent reads this and skips the full skill
|
||||
description: Analyzes code, finds bugs, suggests fixes
|
||||
|
||||
# GOOD: Agent reads full skill to understand workflow
|
||||
description: Use when debugging errors or reviewing code quality
|
||||
```
|
||||
|
||||
### Pre-Violation Triggers for Discipline Skills
|
||||
|
||||
```yaml
|
||||
# BAD: Triggers AFTER violation
|
||||
description: Use when you forgot to write tests
|
||||
|
||||
# GOOD: Triggers BEFORE violation
|
||||
description: Use when implementing any feature, before writing code
|
||||
```
|
||||
|
||||
## Token Efficiency
|
||||
|
||||
### Skill Loaded Every Conversation = Token Drain
|
||||
|
||||
- Frequently-loaded skills: <200 words
|
||||
- All others: <500 words
|
||||
- Move details to `references/` files
|
||||
|
||||
### Don't Duplicate CLI Help
|
||||
|
||||
```markdown
|
||||
# BAD: 50 lines documenting all flags
|
||||
|
||||
# GOOD: One line
|
||||
Run `mytool --help` for all options.
|
||||
```
|
||||
|
||||
## Anti-Rationalization (Discipline Skills Only)
|
||||
|
||||
### Agents Are Smart at Finding Loopholes
|
||||
|
||||
```markdown
|
||||
# BAD: Trust agents will "get the spirit"
|
||||
Write test before code.
|
||||
|
||||
# GOOD: Close every loophole explicitly
|
||||
Write test before code.
|
||||
|
||||
**No exceptions:**
|
||||
- Don't keep code as "reference"
|
||||
- Don't "adapt" existing code
|
||||
- Delete means delete
|
||||
```
|
||||
|
||||
### Build Rationalization Table
|
||||
|
||||
Every excuse from baseline testing goes in the table:
|
||||
|
||||
| Excuse | Reality |
|
||||
|--------|---------|
|
||||
| "Too simple to test" | Simple code breaks. Test takes 30 seconds. |
|
||||
| "I'll test after" | Tests-after prove nothing immediately. |
|
||||
|
||||
## Cross-References
|
||||
|
||||
### Keep References One Level Deep
|
||||
|
||||
```markdown
|
||||
# BAD: Nested chain (A -> B -> C)
|
||||
See [patterns.md] -> which links to [advanced.md] -> which links to [deep.md]
|
||||
|
||||
# GOOD: Flat (A -> B, A -> C)
|
||||
See [patterns.md] and [advanced.md]
|
||||
```
|
||||
|
||||
### Never Force-Load with @
|
||||
|
||||
```markdown
|
||||
# BAD: Burns context immediately
|
||||
@skills/my-skill/SKILL.md
|
||||
|
||||
# GOOD: Agent loads when needed
|
||||
See [my-skill] for details.
|
||||
```
|
||||
|
||||
## Tier Selection
|
||||
|
||||
### Don't Overthink Tier Choice
|
||||
|
||||
```markdown
|
||||
# BAD: Starting with Tier 3 "just in case"
|
||||
# Result: Wasted effort, empty reference files
|
||||
|
||||
# GOOD: Start with Tier 1, upgrade when needed
|
||||
# Can always add references/ later
|
||||
```
|
||||
|
||||
### Signals You Need to Upgrade
|
||||
|
||||
| Signal | Action |
|
||||
|--------|--------|
|
||||
| SKILL.md > 200 lines | -> Tier 2 |
|
||||
| 3+ related sub-topics | -> Tier 2 |
|
||||
| 10+ products/services | -> Tier 3 |
|
||||
| "I need X" vs "I want Y" | -> Tier 3 decision trees |
|
||||
86
.skills/writing-skills/persuasion-principles.md
Normal file
86
.skills/writing-skills/persuasion-principles.md
Normal file
|
|
@ -0,0 +1,86 @@
|
|||
# Persuasion Principles for Skill Design
|
||||
|
||||
## Overview
|
||||
|
||||
LLMs respond to the same persuasion principles as humans. Understanding this psychology helps you design more effective skills - not to manipulate, but to ensure critical practices are followed even under pressure.
|
||||
|
||||
**Research foundation:** Meincke et al. (2025) tested 7 persuasion principles with N=28,000 AI conversations. Persuasion techniques more than doubled compliance rates (33% to 72%, p < .001).
|
||||
|
||||
## The Seven Principles
|
||||
|
||||
### 1. Authority
|
||||
**What it is:** Deference to expertise, credentials, or official sources.
|
||||
|
||||
**How it works in skills:**
|
||||
- Imperative language: "YOU MUST", "Never", "Always"
|
||||
- Non-negotiable framing: "No exceptions"
|
||||
- Eliminates decision fatigue and rationalization
|
||||
|
||||
**When to use:**
|
||||
- Discipline-enforcing skills (TDD, verification requirements)
|
||||
- Safety-critical practices
|
||||
- Established best practices
|
||||
|
||||
### 2. Commitment
|
||||
**What it is:** Consistency with prior actions, statements, or public declarations.
|
||||
|
||||
**How it works in skills:**
|
||||
- Require announcements: "Announce skill usage"
|
||||
- Force explicit choices: "Choose A, B, or C"
|
||||
- Use tracking: TodoWrite for checklists
|
||||
|
||||
### 3. Scarcity
|
||||
**What it is:** Urgency from time limits or limited availability.
|
||||
|
||||
**How it works in skills:**
|
||||
- Time-bound requirements: "Before proceeding"
|
||||
- Sequential dependencies: "Immediately after X"
|
||||
- Prevents procrastination
|
||||
|
||||
### 4. Social Proof
|
||||
**What it is:** Conformity to what others do or what's considered normal.
|
||||
|
||||
**How it works in skills:**
|
||||
- Universal patterns: "Every time", "Always"
|
||||
- Failure modes: "X without Y = failure"
|
||||
- Establishes norms
|
||||
|
||||
### 5. Unity
|
||||
**What it is:** Shared identity, "we-ness", in-group belonging.
|
||||
|
||||
**How it works in skills:**
|
||||
- Collaborative language: "our codebase", "we're colleagues"
|
||||
- Shared goals: "we both want quality"
|
||||
|
||||
### 6. Reciprocity
|
||||
**What it is:** Obligation to return benefits received.
|
||||
- Use sparingly - can feel manipulative
|
||||
- Rarely needed in skills
|
||||
|
||||
### 7. Liking
|
||||
**What it is:** Preference for cooperating with those we like.
|
||||
- **DON'T USE for compliance**
|
||||
- Conflicts with honest feedback culture
|
||||
|
||||
## Principle Combinations by Skill Type
|
||||
|
||||
| Skill Type | Use | Avoid |
|
||||
|------------|-----|-------|
|
||||
| Discipline-enforcing | Authority + Commitment + Social Proof | Liking, Reciprocity |
|
||||
| Guidance/technique | Moderate Authority + Unity | Heavy authority |
|
||||
| Collaborative | Unity + Commitment | Authority, Liking |
|
||||
| Reference | Clarity only | All persuasion |
|
||||
|
||||
## Ethical Use
|
||||
|
||||
**Legitimate:**
|
||||
- Ensuring critical practices are followed
|
||||
- Creating effective documentation
|
||||
- Preventing predictable failures
|
||||
|
||||
**The test:** Would this technique serve the user's genuine interests if they fully understood it?
|
||||
|
||||
## Research Citations
|
||||
|
||||
**Cialdini, R. B. (2021).** *Influence: The Psychology of Persuasion (New and Expanded).* Harper Business.
|
||||
**Meincke, L., et al. (2025).** Call Me A Jerk: Persuading AI to Comply with Objectionable Requests. University of Pennsylvania.
|
||||
|
|
@ -0,0 +1,85 @@
|
|||
# Anti-Rationalization Guide
|
||||
|
||||
Techniques for bulletproofing skills against agent rationalization.
|
||||
|
||||
## The Problem
|
||||
|
||||
Discipline-enforcing skills face a unique challenge: smart agents under pressure will find loopholes.
|
||||
|
||||
## Technique 1: Close Every Loophole Explicitly
|
||||
|
||||
Don't just state the rule - forbid specific workarounds.
|
||||
|
||||
### Bad Example
|
||||
```markdown
|
||||
Write code before test? Delete it.
|
||||
```
|
||||
|
||||
### Good Example
|
||||
```markdown
|
||||
Write code before test? Delete it. Start over.
|
||||
|
||||
**No exceptions**:
|
||||
- Don't keep it as "reference"
|
||||
- Don't "adapt" it while writing tests
|
||||
- Don't look at it
|
||||
- Delete means delete
|
||||
```
|
||||
|
||||
## Technique 2: Address "Spirit vs Letter" Arguments
|
||||
|
||||
Add foundational principle early:
|
||||
|
||||
```markdown
|
||||
**Violating the letter of the rules is violating the spirit of the rules.**
|
||||
```
|
||||
|
||||
## Technique 3: Build Rationalization Table
|
||||
|
||||
| Excuse | Reality |
|
||||
|--------|---------|
|
||||
| "Too simple to test" | Simple code breaks. Test takes 30 seconds. |
|
||||
| "I'll test after" | Tests passing immediately prove nothing. |
|
||||
| "Spirit not ritual" | The letter IS the spirit. |
|
||||
|
||||
## Technique 4: Create Red Flags List
|
||||
|
||||
```markdown
|
||||
## Red Flags - STOP and Start Over
|
||||
- Code before test
|
||||
- "I already manually tested it"
|
||||
- "This is different because..."
|
||||
|
||||
**All of these mean**: Delete code. Start over.
|
||||
```
|
||||
|
||||
## Technique 5: Use Strong Language
|
||||
|
||||
```markdown
|
||||
# Weak (invites rationalization)
|
||||
You should write tests first.
|
||||
|
||||
# Strong (no wiggle room)
|
||||
ALWAYS write test first.
|
||||
NEVER write code before test.
|
||||
```
|
||||
|
||||
## Technique 6: Provide Escape Hatch for Legitimate Cases
|
||||
|
||||
```markdown
|
||||
## When NOT to Use
|
||||
- Spike solutions (throwaway exploratory code)
|
||||
- One-time scripts deleting in 1 hour
|
||||
|
||||
**Everything else**: Follow the rule. No exceptions.
|
||||
```
|
||||
|
||||
## Complete Bulletproofing Checklist
|
||||
|
||||
- [ ] Forbidden each specific workaround explicitly?
|
||||
- [ ] Added "spirit vs letter" principle?
|
||||
- [ ] Built rationalization table from baseline tests?
|
||||
- [ ] Created red flags list?
|
||||
- [ ] Used strong language (ALWAYS/NEVER)?
|
||||
- [ ] Provided explicit escape hatch?
|
||||
- [ ] Description includes pre-violation symptoms?
|
||||
90
.skills/writing-skills/references/cso/README.md
Normal file
90
.skills/writing-skills/references/cso/README.md
Normal file
|
|
@ -0,0 +1,90 @@
|
|||
# CSO Guide - Claude Search Optimization
|
||||
|
||||
Advanced techniques for making skills discoverable by agents.
|
||||
|
||||
## The Discovery Problem
|
||||
|
||||
You have 100+ skills. Agent receives a task. How does it find the RIGHT skill?
|
||||
|
||||
**Answer**: The `description` field.
|
||||
|
||||
## Critical Rule: Description = Triggers, NOT Workflow
|
||||
|
||||
### The Trap
|
||||
|
||||
When description summarizes workflow, agents take a shortcut.
|
||||
|
||||
**Real example that failed**:
|
||||
|
||||
```yaml
|
||||
# Agent did ONE review instead of TWO
|
||||
description: Code review between tasks
|
||||
|
||||
# Skill body had flowchart showing TWO reviews
|
||||
```
|
||||
|
||||
**Why it failed**: Agent read description, thought "code review between tasks means one review", never read the flowchart.
|
||||
|
||||
**Fix**:
|
||||
|
||||
```yaml
|
||||
# Agent now reads full skill and follows flowchart
|
||||
description: Use when executing implementation plans with independent tasks
|
||||
```
|
||||
|
||||
### The Pattern
|
||||
|
||||
```yaml
|
||||
# BAD: Workflow summary
|
||||
description: Analyzes git diff, generates commit message in conventional format
|
||||
|
||||
# GOOD: Trigger conditions only
|
||||
description: Use when generating commit messages or reviewing staged changes
|
||||
```
|
||||
|
||||
## Token Efficiency
|
||||
|
||||
**Target word counts**:
|
||||
- Frequently-loaded skills: <200 words total
|
||||
- Other skills: <500 words
|
||||
|
||||
## Keyword Strategy
|
||||
|
||||
### Error Messages
|
||||
Include EXACT error text users will see.
|
||||
|
||||
### Symptoms
|
||||
Use words users naturally say: "flaky", "hangs", "slow", "timeout", "race condition"
|
||||
|
||||
### Tools & Commands
|
||||
Actual names, not descriptions: "pytest", not "Python testing"
|
||||
|
||||
### Synonyms
|
||||
Cover multiple ways to describe same thing: timeout/hang/freeze
|
||||
|
||||
## Description Template
|
||||
|
||||
```yaml
|
||||
description: "Use when [SPECIFIC TRIGGER]."
|
||||
metadata:
|
||||
triggers: [error1], [symptom2], [tool3]
|
||||
```
|
||||
|
||||
## Third Person Rule
|
||||
|
||||
```yaml
|
||||
# BAD: First person
|
||||
description: "I can help you with async tests"
|
||||
|
||||
# GOOD: Third person
|
||||
description: "Handles async tests with race conditions"
|
||||
```
|
||||
|
||||
## Verification Checklist
|
||||
|
||||
- [ ] Description starts with "Use when..."?
|
||||
- [ ] Description is <500 characters?
|
||||
- [ ] Description lists ONLY triggers, not workflow?
|
||||
- [ ] Includes 3+ keywords (errors/symptoms/tools)?
|
||||
- [ ] Third person throughout?
|
||||
- [ ] Name uses gerund or verb-first format?
|
||||
87
.skills/writing-skills/references/standards/README.md
Normal file
87
.skills/writing-skills/references/standards/README.md
Normal file
|
|
@ -0,0 +1,87 @@
|
|||
---
|
||||
description: Standards and naming rules for creating agent skills.
|
||||
metadata:
|
||||
tags: [standards, naming, yaml, structure]
|
||||
---
|
||||
|
||||
# Skill Development Guide
|
||||
|
||||
## Directory Structure
|
||||
|
||||
```
|
||||
skills/
|
||||
{skill-name}/ # kebab-case, matches `name` field
|
||||
SKILL.md # Required: main skill definition
|
||||
references/ # Optional: supporting documentation
|
||||
README.md # Sub-topic entry point
|
||||
*.md # Additional files
|
||||
```
|
||||
|
||||
## Naming Rules
|
||||
|
||||
| Element | Rule | Example |
|
||||
|---------|------|---------|
|
||||
| Directory | kebab-case, 1-64 chars | `react-best-practices` |
|
||||
| `SKILL.md` | ALL CAPS, exact filename | `SKILL.md` (not `skill.md`) |
|
||||
| `name` field | Must match directory name | `name: react-best-practices` |
|
||||
|
||||
## SKILL.md Structure
|
||||
|
||||
```markdown
|
||||
---
|
||||
name: {skill-name}
|
||||
description: >-
|
||||
Use when [trigger condition].
|
||||
metadata:
|
||||
category: technique
|
||||
triggers: keyword1, keyword2, error-text
|
||||
---
|
||||
|
||||
# Skill Title
|
||||
|
||||
Brief description of what this skill does.
|
||||
|
||||
## When to Use
|
||||
- Symptom or situation A
|
||||
- Symptom or situation B
|
||||
|
||||
## How It Works
|
||||
Step-by-step instructions or reference content.
|
||||
|
||||
## Examples
|
||||
Concrete usage examples.
|
||||
|
||||
## Common Mistakes
|
||||
What to avoid and why.
|
||||
```
|
||||
|
||||
## Description Best Practices
|
||||
|
||||
```yaml
|
||||
# BAD: Workflow summary
|
||||
description: Analyzes code, finds bugs, suggests fixes
|
||||
|
||||
# GOOD: Trigger conditions only
|
||||
description: Use when debugging errors or reviewing code quality.
|
||||
```
|
||||
|
||||
**Rules:**
|
||||
- Start with "Use when..."
|
||||
- Keep under 500 characters
|
||||
- Use third person
|
||||
|
||||
## Context Efficiency
|
||||
|
||||
| Guideline | Reason |
|
||||
|-----------|--------|
|
||||
| Keep SKILL.md < 500 lines | Reduces context consumption |
|
||||
| Put details in supporting files | Agent reads only what's needed |
|
||||
| Use tables for reference data | More compact than prose |
|
||||
|
||||
## Verification Checklist
|
||||
|
||||
- [ ] `name` matches directory name?
|
||||
- [ ] `SKILL.md` is ALL CAPS?
|
||||
- [ ] Description starts with "Use when..."?
|
||||
- [ ] Under 500 lines?
|
||||
- [ ] Tested with real scenarios?
|
||||
|
|
@ -0,0 +1,39 @@
|
|||
# SKILL.md Metadata Standard
|
||||
|
||||
Official frontmatter fields.
|
||||
|
||||
## Required Fields
|
||||
|
||||
```yaml
|
||||
---
|
||||
name: skill-name
|
||||
description: >-
|
||||
Use when [trigger condition].
|
||||
---
|
||||
```
|
||||
|
||||
| Field | Rules |
|
||||
|-------|-------|
|
||||
| `name` | 1-64 chars, lowercase, hyphens only, must match directory name |
|
||||
| `description` | 1-1024 chars, should describe when to use |
|
||||
|
||||
## Optional Fields
|
||||
|
||||
```yaml
|
||||
---
|
||||
name: skill-name
|
||||
description: Purpose and triggers.
|
||||
metadata:
|
||||
category: "reference"
|
||||
version: "1.0.0"
|
||||
---
|
||||
```
|
||||
|
||||
## Name Validation
|
||||
|
||||
```regex
|
||||
^[a-z0-9]+(-[a-z0-9]+)*$
|
||||
```
|
||||
|
||||
**Valid**: `my-skill`, `git-release`, `tdd`
|
||||
**Invalid**: `My-Skill`, `my_skill`, `-my-skill`
|
||||
54
.skills/writing-skills/references/templates/discipline.md
Normal file
54
.skills/writing-skills/references/templates/discipline.md
Normal file
|
|
@ -0,0 +1,54 @@
|
|||
---
|
||||
name: discipline-name
|
||||
description: >-
|
||||
Use when [BEFORE violation].
|
||||
metadata:
|
||||
category: discipline
|
||||
triggers: new feature, code change, implementation
|
||||
---
|
||||
|
||||
# Rule Name
|
||||
|
||||
## Iron Law
|
||||
|
||||
**[SINGLE SENTENCE ABSOLUTE RULE]**
|
||||
|
||||
Violating the letter IS violating the spirit.
|
||||
|
||||
## The Rule
|
||||
|
||||
1. ALWAYS [step 1]
|
||||
2. NEVER [step 2]
|
||||
3. [Step 3]
|
||||
|
||||
## Violations
|
||||
|
||||
[Action before rule]? **Delete it. Start over.**
|
||||
|
||||
**No exceptions:**
|
||||
- Don't keep it as "reference"
|
||||
- Don't "adapt" it
|
||||
- Delete means delete
|
||||
|
||||
## Common Rationalizations
|
||||
|
||||
| Excuse | Reality |
|
||||
|--------|---------|
|
||||
| "Too simple" | Simple code breaks. Rule takes 30 seconds. |
|
||||
| "I'll do it after" | After = never. Do it now. |
|
||||
| "Spirit not ritual" | The ritual IS the spirit. |
|
||||
|
||||
## Red Flags - STOP
|
||||
|
||||
- [Flag 1]
|
||||
- [Flag 2]
|
||||
- "This is different because..."
|
||||
|
||||
**All mean:** Delete. Start over.
|
||||
|
||||
## Valid Exceptions
|
||||
|
||||
- [Exception 1]
|
||||
- [Exception 2]
|
||||
|
||||
**Everything else:** Follow the rule.
|
||||
48
.skills/writing-skills/references/templates/pattern.md
Normal file
48
.skills/writing-skills/references/templates/pattern.md
Normal file
|
|
@ -0,0 +1,48 @@
|
|||
---
|
||||
name: pattern-name
|
||||
description: >-
|
||||
Use when [recognizable symptom].
|
||||
metadata:
|
||||
category: pattern
|
||||
triggers: complexity, hard-to-follow, nested
|
||||
---
|
||||
|
||||
# Pattern Name
|
||||
|
||||
## The Pattern
|
||||
|
||||
[1-2 sentence core idea]
|
||||
|
||||
## Recognition Signs
|
||||
|
||||
- [Sign that pattern applies]
|
||||
- [Another sign]
|
||||
- [Code smell]
|
||||
|
||||
## Before
|
||||
|
||||
```typescript
|
||||
// Complex/problematic
|
||||
function before() {
|
||||
// nested, confusing
|
||||
}
|
||||
```
|
||||
|
||||
## After
|
||||
|
||||
```typescript
|
||||
// Clean/improved
|
||||
function after() {
|
||||
// flat, clear
|
||||
}
|
||||
```
|
||||
|
||||
## When NOT to Use
|
||||
|
||||
- [Over-engineering case]
|
||||
- [Simple case that doesn't need it]
|
||||
|
||||
## Impact
|
||||
|
||||
**Before:** [Problem metric]
|
||||
**After:** [Improved metric]
|
||||
35
.skills/writing-skills/references/templates/reference.md
Normal file
35
.skills/writing-skills/references/templates/reference.md
Normal file
|
|
@ -0,0 +1,35 @@
|
|||
---
|
||||
name: reference-name
|
||||
description: >-
|
||||
Use when working with [domain].
|
||||
metadata:
|
||||
category: reference
|
||||
triggers: tool, api, specific-terms
|
||||
---
|
||||
|
||||
# Reference Name
|
||||
|
||||
## Quick Reference
|
||||
|
||||
| Command | Purpose |
|
||||
|---------|---------|
|
||||
| `cmd1` | Does X |
|
||||
| `cmd2` | Does Y |
|
||||
|
||||
## Common Patterns
|
||||
|
||||
**Pattern A:**
|
||||
```bash
|
||||
example command
|
||||
```
|
||||
|
||||
**Pattern B:**
|
||||
```bash
|
||||
another example
|
||||
```
|
||||
|
||||
## Detailed Docs
|
||||
|
||||
For more options, run `--help` or see:
|
||||
- patterns.md
|
||||
- examples.md
|
||||
59
.skills/writing-skills/references/templates/technique.md
Normal file
59
.skills/writing-skills/references/templates/technique.md
Normal file
|
|
@ -0,0 +1,59 @@
|
|||
---
|
||||
name: technique-name
|
||||
description: Use when [specific symptom].
|
||||
metadata:
|
||||
category: technique
|
||||
triggers: error-text, symptom, tool-name
|
||||
---
|
||||
|
||||
# Technique Name
|
||||
|
||||
## Overview
|
||||
|
||||
[1-2 sentence core principle]
|
||||
|
||||
## When to Use
|
||||
|
||||
- [Symptom A]
|
||||
- [Symptom B]
|
||||
- [Error message text]
|
||||
|
||||
**NOT for:**
|
||||
- [When to avoid]
|
||||
|
||||
## The Problem
|
||||
|
||||
```javascript
|
||||
// Bad example
|
||||
function badCode() {
|
||||
// problematic pattern
|
||||
}
|
||||
```
|
||||
|
||||
## The Solution
|
||||
|
||||
```javascript
|
||||
// Good example
|
||||
function goodCode() {
|
||||
// improved pattern
|
||||
}
|
||||
```
|
||||
|
||||
## Step-by-Step
|
||||
|
||||
1. [First step]
|
||||
2. [Second step]
|
||||
3. [Final step]
|
||||
|
||||
## Quick Reference
|
||||
|
||||
| Scenario | Approach |
|
||||
|----------|----------|
|
||||
| Case A | Solution A |
|
||||
| Case B | Solution B |
|
||||
|
||||
## Common Mistakes
|
||||
|
||||
**Mistake 1:** [Description]
|
||||
- Wrong: `bad code`
|
||||
- Right: `good code`
|
||||
|
|
@ -0,0 +1,17 @@
|
|||
# Platform Name Skill
|
||||
|
||||
Template for complex Tier 3 skills.
|
||||
|
||||
## Structure
|
||||
|
||||
```
|
||||
skill/
|
||||
SKILL.md # Dispatcher
|
||||
references/
|
||||
topic/
|
||||
README.md # Overview
|
||||
api.md # API Reference
|
||||
config.md # Configuration
|
||||
patterns.md # Recipes
|
||||
gotchas.md # Critical Errors
|
||||
```
|
||||
66
.skills/writing-skills/references/testing/README.md
Normal file
66
.skills/writing-skills/references/testing/README.md
Normal file
|
|
@ -0,0 +1,66 @@
|
|||
# Testing Guide - TDD for Skills
|
||||
|
||||
Complete methodology for testing skills using RED-GREEN-REFACTOR cycle.
|
||||
|
||||
## Testing All Skill Types
|
||||
|
||||
### Discipline-Enforcing Skills (rules/requirements)
|
||||
|
||||
**Test with**:
|
||||
- Academic questions: Do they understand the rules?
|
||||
- Pressure scenarios: Do they comply under stress?
|
||||
- Multiple pressures combined: time + sunk cost + exhaustion
|
||||
|
||||
**Success criteria**: Agent follows rule under maximum pressure
|
||||
|
||||
### Technique Skills (how-to guides)
|
||||
|
||||
**Test with**:
|
||||
- Application scenarios: Can they apply the technique correctly?
|
||||
- Variation scenarios: Do they handle edge cases?
|
||||
- Missing information tests: Do instructions have gaps?
|
||||
|
||||
**Success criteria**: Agent successfully applies technique to new scenario
|
||||
|
||||
### Pattern Skills (mental models)
|
||||
|
||||
**Test with**:
|
||||
- Recognition scenarios: Do they recognize when pattern applies?
|
||||
- Counter-examples: Do they know when NOT to apply?
|
||||
|
||||
**Success criteria**: Agent correctly identifies when/how to apply pattern
|
||||
|
||||
### Reference Skills (documentation/APIs)
|
||||
|
||||
**Test with**:
|
||||
- Retrieval scenarios: Can they find the right information?
|
||||
- Gap testing: Are common use cases covered?
|
||||
|
||||
**Success criteria**: Agent finds and correctly applies reference information
|
||||
|
||||
## Pressure Types for Testing
|
||||
|
||||
| Pressure | Example |
|
||||
|----------|---------|
|
||||
| Time | "You have 5 minutes to complete this task" |
|
||||
| Sunk cost | "You already spent 2 hours on this" |
|
||||
| Authority | "Senior developer said to skip tests" |
|
||||
| Exhaustion | "This is the 10th task today" |
|
||||
|
||||
## Complete Test Checklist
|
||||
|
||||
**Baseline (RED)**:
|
||||
- [ ] Designed 3+ pressure scenarios
|
||||
- [ ] Ran scenarios WITHOUT skill
|
||||
- [ ] Documented verbatim agent responses
|
||||
|
||||
**Implementation (GREEN)**:
|
||||
- [ ] Skill addresses SPECIFIC baseline failures
|
||||
- [ ] Re-ran scenarios WITH skill
|
||||
- [ ] Agent complied in all scenarios
|
||||
|
||||
**Bulletproofing (REFACTOR)**:
|
||||
- [ ] Tested with combined pressures
|
||||
- [ ] Found and documented new rationalizations
|
||||
- [ ] Added explicit counters
|
||||
- [ ] Re-tested until no more loopholes
|
||||
30
.skills/writing-skills/references/tier-1-simple/README.md
Normal file
30
.skills/writing-skills/references/tier-1-simple/README.md
Normal file
|
|
@ -0,0 +1,30 @@
|
|||
---
|
||||
description: When to use Tier 1 (Simple) skill architecture.
|
||||
metadata:
|
||||
tags: [tier-1, simple, single-file]
|
||||
---
|
||||
|
||||
# Tier 1: Simple Skills
|
||||
|
||||
Single-file skills for focused, specific purposes.
|
||||
|
||||
## When to Use
|
||||
|
||||
- **Single concept**: One technique, one pattern, one reference
|
||||
- **Under 200 lines**: Can fit comfortably in one file
|
||||
- **No complex decision logic**: User knows exactly what they need
|
||||
- **Frequently loaded**: Needs minimal token footprint
|
||||
|
||||
## Structure
|
||||
|
||||
```
|
||||
my-skill/
|
||||
SKILL.md # Everything in one file
|
||||
```
|
||||
|
||||
## Checklist
|
||||
|
||||
- [ ] Fits in <200 lines
|
||||
- [ ] Single focused purpose
|
||||
- [ ] No need for `references/` directory
|
||||
- [ ] Description uses "Use when..." pattern
|
||||
52
.skills/writing-skills/references/tier-2-expanded/README.md
Normal file
52
.skills/writing-skills/references/tier-2-expanded/README.md
Normal file
|
|
@ -0,0 +1,52 @@
|
|||
---
|
||||
description: When to use Tier 2 (Expanded) skill architecture.
|
||||
metadata:
|
||||
tags: [tier-2, expanded, multi-file]
|
||||
---
|
||||
|
||||
# Tier 2: Expanded Skills
|
||||
|
||||
Multi-file skills for complex topics with multiple sub-concepts.
|
||||
|
||||
## When to Use
|
||||
|
||||
- **Multiple related concepts**: Needs separation of concerns
|
||||
- **200-1000 lines total**: Too big for one file
|
||||
- **Needs reference files**: Patterns, examples, troubleshooting
|
||||
- **Cross-linking**: Users need to navigate between sub-topics
|
||||
|
||||
## Structure
|
||||
|
||||
```
|
||||
my-skill/
|
||||
SKILL.md # Overview + navigation
|
||||
references/
|
||||
core/
|
||||
README.md # Main concept
|
||||
patterns/
|
||||
README.md # Usage patterns
|
||||
troubleshooting/
|
||||
README.md # Common issues
|
||||
```
|
||||
|
||||
## Progressive Disclosure
|
||||
|
||||
1. **Metadata** (~100 tokens): Name + description loaded at startup
|
||||
2. **SKILL.md** (<500 lines): Decision tree + index
|
||||
3. **References** (as needed): Loaded only when user navigates
|
||||
|
||||
## Key Differences from Tier 1
|
||||
|
||||
| Aspect | Tier 1 | Tier 2 |
|
||||
|--------|--------|--------|
|
||||
| Files | 1 | 5-20 |
|
||||
| Total lines | <200 | 200-1000 |
|
||||
| Decision logic | None | Simple tree |
|
||||
| Token cost | Minimal | Medium (progressive) |
|
||||
|
||||
## Checklist
|
||||
|
||||
- [ ] SKILL.md has clear navigation links
|
||||
- [ ] Each `references/` subdir has README.md
|
||||
- [ ] No circular references between files
|
||||
- [ ] Decision tree points to specific files
|
||||
51
.skills/writing-skills/references/tier-3-platform/README.md
Normal file
51
.skills/writing-skills/references/tier-3-platform/README.md
Normal file
|
|
@ -0,0 +1,51 @@
|
|||
---
|
||||
description: When to use Tier 3 (Platform) skill architecture for large platforms.
|
||||
metadata:
|
||||
tags: [tier-3, platform, enterprise]
|
||||
---
|
||||
|
||||
# Tier 3: Platform Skills
|
||||
|
||||
Enterprise-grade skills for entire platforms (AWS, Cloudflare, Convex, etc).
|
||||
|
||||
## When to Use
|
||||
|
||||
- **Entire platform**: 10+ products/services
|
||||
- **1000+ lines total**: Would overwhelm context if monolithic
|
||||
- **Complex decision logic**: Users start with "I need X" not "I want product Y"
|
||||
|
||||
## The 5-File Pattern
|
||||
|
||||
Each product directory has exactly 5 files:
|
||||
|
||||
| File | Purpose | When to Load |
|
||||
|------|---------|--------------|
|
||||
| `README.md` | Overview, when to use | Always first |
|
||||
| `api.md` | Runtime APIs, methods | Implementing features |
|
||||
| `configuration.md` | Config, environment | Setting up |
|
||||
| `patterns.md` | Common workflows | Best practices |
|
||||
| `gotchas.md` | Pitfalls, limits | Debugging |
|
||||
|
||||
## Decision Trees
|
||||
|
||||
```markdown
|
||||
Need to store data?
|
||||
Simple key-value -> kv/
|
||||
Relational queries -> d1/
|
||||
Large files/blobs -> r2/
|
||||
Per-user state -> durable-objects/
|
||||
```
|
||||
|
||||
## Progressive Disclosure in Action
|
||||
|
||||
- **Startup**: Only name + description (~100 tokens)
|
||||
- **Activation**: SKILL.md with trees (<5000 tokens)
|
||||
- **Navigation**: One product's 5 files (as needed)
|
||||
|
||||
## Checklist
|
||||
|
||||
- [ ] SKILL.md contains ONLY decision trees + index
|
||||
- [ ] Each product has exactly 5 files
|
||||
- [ ] Decision trees cover all "I need X" scenarios
|
||||
- [ ] Cross-references stay one level deep
|
||||
- [ ] Every product has `gotchas.md`
|
||||
85
.skills/writing-skills/testing-skills-with-subagents.md
Normal file
85
.skills/writing-skills/testing-skills-with-subagents.md
Normal file
|
|
@ -0,0 +1,85 @@
|
|||
# Testing Skills With Subagents
|
||||
|
||||
**Load this reference when:** creating or editing skills, before deployment, to verify they work under pressure and resist rationalization.
|
||||
|
||||
## Overview
|
||||
|
||||
**Testing skills is just TDD applied to process documentation.**
|
||||
|
||||
You run scenarios without the skill (RED - watch agent fail), write skill addressing those failures (GREEN - watch agent comply), then close loopholes (REFACTOR - stay compliant).
|
||||
|
||||
**Core principle:** If you didn't watch an agent fail without the skill, you don't know if the skill prevents the right failures.
|
||||
|
||||
## When to Use
|
||||
|
||||
Test skills that:
|
||||
- Enforce discipline (TDD, testing requirements)
|
||||
- Have compliance costs (time, effort, rework)
|
||||
- Could be rationalized away ("just this once")
|
||||
- Contradict immediate goals (speed over quality)
|
||||
|
||||
Don't test:
|
||||
- Pure reference skills (API docs, syntax guides)
|
||||
- Skills without rules to violate
|
||||
- Skills agents have no incentive to bypass
|
||||
|
||||
## TDD Mapping for Skill Testing
|
||||
|
||||
| TDD Phase | Skill Testing | What You Do |
|
||||
|-----------|---------------|-------------|
|
||||
| **RED** | Baseline test | Run scenario WITHOUT skill, watch agent fail |
|
||||
| **Verify RED** | Capture rationalizations | Document exact failures verbatim |
|
||||
| **GREEN** | Write skill | Address specific baseline failures |
|
||||
| **Verify GREEN** | Pressure test | Run scenario WITH skill, verify compliance |
|
||||
| **REFACTOR** | Plug holes | Find new rationalizations, add counters |
|
||||
| **Stay GREEN** | Re-verify | Test again, ensure still compliant |
|
||||
|
||||
## RED Phase: Baseline Testing (Watch It Fail)
|
||||
|
||||
**Goal:** Run test WITHOUT the skill - watch agent fail, document exact failures.
|
||||
|
||||
**Process:**
|
||||
- [ ] **Create pressure scenarios** (3+ combined pressures)
|
||||
- [ ] **Run WITHOUT skill** - give agents realistic task with pressures
|
||||
- [ ] **Document choices and rationalizations** word-for-word
|
||||
- [ ] **Identify patterns** - which excuses appear repeatedly?
|
||||
- [ ] **Note effective pressures** - which scenarios trigger violations?
|
||||
|
||||
## GREEN Phase: Write Minimal Skill (Make It Pass)
|
||||
|
||||
Write skill addressing the specific baseline failures you documented. Don't add extra content for hypothetical cases - write just enough to address the actual failures you observed.
|
||||
|
||||
Run same scenarios WITH skill. Agent should now comply.
|
||||
|
||||
If agent still fails: skill is unclear or incomplete. Revise and re-test.
|
||||
|
||||
## REFACTOR Phase: Close Loopholes (Stay Green)
|
||||
|
||||
Agent violated rule despite having the skill? Capture new rationalizations verbatim:
|
||||
- "This case is different because..."
|
||||
- "I'm following the spirit not the letter"
|
||||
- "Being pragmatic means adapting"
|
||||
- "Deleting X hours is wasteful"
|
||||
|
||||
**Document every excuse.** These become your rationalization table.
|
||||
|
||||
## Testing Checklist (TDD for Skills)
|
||||
|
||||
**RED Phase:**
|
||||
- [ ] Created pressure scenarios (3+ combined pressures)
|
||||
- [ ] Ran scenarios WITHOUT skill (baseline)
|
||||
- [ ] Documented agent failures and rationalizations verbatim
|
||||
|
||||
**GREEN Phase:**
|
||||
- [ ] Wrote skill addressing specific baseline failures
|
||||
- [ ] Ran scenarios WITH skill
|
||||
- [ ] Agent now complies
|
||||
|
||||
**REFACTOR Phase:**
|
||||
- [ ] Identified NEW rationalizations from testing
|
||||
- [ ] Added explicit counters for each loophole
|
||||
- [ ] Updated rationalization table
|
||||
- [ ] Updated red flags list
|
||||
- [ ] Re-tested - agent still complies
|
||||
- [ ] Meta-tested to verify clarity
|
||||
- [ ] Agent follows rule under maximum pressure
|
||||
127
AGENTS.md
Normal file
127
AGENTS.md
Normal file
|
|
@ -0,0 +1,127 @@
|
|||
# AGENTS.md
|
||||
|
||||
## Project
|
||||
- Minecraft Console Client (MCC) is a cross-platform text/TUI client for Minecraft Java Edition.
|
||||
- Primary scope: connect to servers, send chat and commands, receive text, automate gameplay/admin tasks, and extend behavior through built-in bots or runtime C# scripts.
|
||||
- Secondary scope: protocol/version adaptation tooling, docs site, legacy GUI wrapper, and debug tooling.
|
||||
- Decompiled server source for both the old and new MC versions in `$MCC_REPO/MinecraftOfficial/<version>-decompiled/`
|
||||
|
||||
## Build / Run
|
||||
- Init submodules first: `git submodule update --init --recursive`
|
||||
- Build for local development: `source tools/mcc-env.sh && mcc-build`
|
||||
- Publish (matches CI shape): `source tools/mcc-env.sh && mcc-publish --rid <RID>`
|
||||
- Run/debug from source: `source tools/mcc-env.sh && mcc-debug -v 1.21.11 --file-input`
|
||||
- Docs: `cd docs && npm install && npm run docs:dev` or `npm run docs:build`
|
||||
- Docker: `cd Docker && docker build -t minecraft-console-client:latest .`
|
||||
- Tests: no dedicated test project is present in the main solution.
|
||||
- Current state: the solution builds after submodule init, but the underlying .NET build emits many analyzer and NuGet vulnerability warnings; treat them as real.
|
||||
- Server roots: `tools/` helpers look for server jars under `MinecraftOfficial/downloads/<version>/` by default, but also support an external root via the `MCC_SERVERS` environment variable.
|
||||
- Multi-version testing: tmux-based local server sessions are shared state. Run cross-version test matrices sequentially unless you have explicit per-version isolation. A server logging `Done` does not guarantee immediate RCON availability; retry RCON setup commands.
|
||||
- Automated test configs: for repeated or matrix test runs, prefer generating a temporary MCC config per run instead of reusing the repo-root `MinecraftClient.ini`, to avoid leaking state between runs.
|
||||
- For agent-driven local development, prefer `mcc-build`, `mcc-publish`, `mcc-build-clean`, `mcc-debug`, `mcc-run`, and `mcc-tui` over raw `dotnet build`, `dotnet publish`, or `dotnet run`, so worktree-local temp build routing stays active.
|
||||
|
||||
## Architecture
|
||||
- `Program` bootstraps console I/O, TOML config, auth/session state, MC version selection, Forge detection, then creates `McClient`.
|
||||
- `McClient` is the live session runtime: TCP client, selected protocol handler, Brigadier command dispatcher, loaded bots, world/inventory/entity state, queued chat, movement/pathing, reconnect flow.
|
||||
- `Protocol/` is the network/auth boundary. `ProtocolHandler` maps Minecraft versions to protocol numbers and selects either `Protocol16Handler` (1.4.6-1.6.4) or `Protocol18Handler` (1.7.2+).
|
||||
- `Scripting/ChatBot` is the extension boundary. Built-in bots and `/script` C# bots share the same event/tick API.
|
||||
- Main runtime flow: console input -> internal Brigadier command or server chat; packets -> protocol handler -> `McClient` state update -> bot events; `OnUpdate()` (20 TPS) drives bot ticks, delayed work, chat cooldowns, movement, and main-thread tasks.
|
||||
|
||||
## Technology Stack
|
||||
- Main app: C#, .NET 10, nullable enabled.
|
||||
- Command system: `Brigadier.NET`.
|
||||
- Config: TOML via `Samboy063.Tomlet`.
|
||||
- Runtime scripting: Roslyn (`Microsoft.CodeAnalysis.CSharp`) with in-memory compilation.
|
||||
- Networking/auth: custom Minecraft protocol handlers, DNS SRV lookup (`DnsClient`), Forge/session/profile-key support.
|
||||
- Integrations: `DSharpPlus`, `Telegram.Bot`, `MessagePack`, `Magick.NET`, `Sentry`.
|
||||
- Docs site: VuePress 2 (`docs/package.json`).
|
||||
- Tooling: Docker, GitHub Actions, Python 3.10+ scripts under `tools/` for palette/version generation.
|
||||
- Legacy UI: `MinecraftClientGUI` is a separate .NET Framework 4.0 WinForms wrapper, not the main runtime.
|
||||
|
||||
## Version Support
|
||||
Feature columns mean:
|
||||
- Inventory: `/inventory` plus inventory/container bot APIs
|
||||
- Movement: terrain handling, `/move`, and movement/pathing bots
|
||||
- Entity: entity tracking and entity-driven bot events
|
||||
|
||||
| Minecraft | Protocol path | Inventory | Movement | Entity | Notes |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| 1.4.6-1.6.4 | `Protocol16Handler` | No | No | No | Core login/chat only |
|
||||
| 1.7.2-1.7.10 | `Protocol18Handler` | No | Yes | No | Pre-1.8 special case |
|
||||
| 1.8-1.9.4 | `Protocol18Handler` | Partial / docs conflict | Yes | Yes | Runtime gates allow 1.8+, but docs still warn inventory is unsupported through 1.9 |
|
||||
| 1.10-1.12.2 | `Protocol18Handler` | Yes | Yes | Yes | Pre-flattening palettes |
|
||||
| 1.13-1.19.2 | `Protocol18Handler` | Yes | Yes | Yes | Flattened block/item/entity palettes |
|
||||
| 1.19.3-1.20.4 | `Protocol18Handler` | Yes | Yes | Yes | Newer chat/signing and palette splits |
|
||||
| 1.20.6-1.21.4 | `Protocol18Handler` | Yes | Yes | Yes | Registry-driven world/attribute handling |
|
||||
| 1.21.5-1.21.8 | `Protocol18Handler` | Yes | Yes | Yes | 1.21.7/1.21.8 reuse 1.21.6 block/entity palettes in code |
|
||||
| 1.21.9-1.21.10 | `Protocol18Handler` | Yes | Yes | Yes | Version tools prefer server data reports since 1.21.9 |
|
||||
| 1.21.11 | `Protocol18Handler` | Yes | Yes | Yes | Own entity/item/metadata palettes; blocks reuse 1.21.9 palette |
|
||||
| 26.1 | `Protocol18Handler` | Yes | Yes | Yes | New Minecraft version naming scheme |
|
||||
| 26.2 | `Protocol18Handler` | Yes | Yes | Yes | Latest coded support; own item/block/entity palettes, reuses 26.1 packet IDs and metadata serializers |
|
||||
|
||||
Notes:
|
||||
- Declared code range is `1.4.6` to `26.2`.
|
||||
- Human docs are stale in places and sometimes stop at older ranges; prefer code when docs and code disagree.
|
||||
- Movement/pathing limits called out in docs still apply: no swimming, no knockback. The `Physics/` engine adds vanilla-accurate collision and movement but some edge cases remain.
|
||||
|
||||
## Module Map
|
||||
### Core Runtime
|
||||
|
||||
| Module | What It Owns | Important Files |
|
||||
| --- | --- | --- |
|
||||
| `MinecraftClient/` | Main `net10.0` runtime assembly and the best starting point. `Program.cs` owns startup, config load/writeback, CLI handling, auth/version selection, update/data-generation entrypoints, and restart/failure flow. `McClient.cs` owns the live session runtime: protocol handler ownership, command dispatch, bot lifecycle, world/inventory/entity state, queued chat, movement ticks, reconnect/disconnect logic, and the main-thread invoke queue. `Settings.cs` defines the TOML schema and runtime/internal overrides used across the app. | `Program.cs`, `McClient.cs`, `Settings.cs`, `ConsoleIO.cs`, `Command.cs`, `UpgradeHelper.cs`, `AutoTimeout.cs` |
|
||||
| `MinecraftClient/Protocol/` | Network/auth/session boundary. `ProtocolHandler.cs` does DNS SRV lookup, server ping/version detection, MC-version to protocol mapping, and handler selection. `Protocol16.cs` and `Protocol18.cs` implement the packet flow for legacy and modern versions. `Protocol18Terrain.cs` decodes chunk sections/biomes into `World`. `DataTypes.cs` is the low-level reader/writer layer for VarInts, metadata, NBT-like structures, and packet fields. `Message/`, `ProfileKey/`, `Session/`, `Handlers/Forge/`, `Handlers/PacketPalettes/`, `Handlers/Packet/`, and `Handlers/StructuredComponents/` cover chat/signing, cached auth, Forge, packet IDs, packet-level parsing, and 1.20.6+ item components with versioned registries under `StructuredComponents/Registries/`. | `Protocol/ProtocolHandler.cs`, `Protocol/Handlers/Protocol16.cs`, `Protocol/Handlers/Protocol18.cs`, `Protocol/Handlers/Protocol18Terrain.cs`, `Protocol/Handlers/DataTypes.cs`, `Protocol/Message/ChatParser.cs`, `Protocol/MicrosoftAuthentication.cs`, `Protocol/MojangAPI.cs` |
|
||||
| `MinecraftClient/Mapping/` | World model, terrain storage, movement logic, and versioned block/entity metadata. `World.cs` stores chunk columns, dimension data, and 1.20.6+ registry-derived dimension/attribute mappings. `Chunk*`, `Block.cs`, and `Location.cs` are the terrain primitives. `Movement.cs` contains step generation, gravity/on-ground checks, and path execution support. `Material.cs` plus `BlockPalettes/*.cs` map block-state IDs to MCC materials. `Entity.cs`, `EntityType.cs`, `EntityPalettes/*.cs`, `EntityMetadataPalette.cs`, and `EntityMetadataPalettes/*.cs` do the same for entities and metadata serializers. | `Mapping/World.cs`, `Mapping/ChunkColumn.cs`, `Mapping/Chunk.cs`, `Mapping/Block.cs`, `Mapping/Location.cs`, `Mapping/Movement.cs`, `Mapping/RaycastHelper.cs`, `Mapping/Material.cs`, `Mapping/Entity.cs`, `Mapping/EntityType.cs` |
|
||||
| `MinecraftClient/Inventory/` | Inventory/container snapshots, item decoding, and versioned item registries. `Container.cs` models player inventories and server windows, including slot contents and container properties. `Item.cs` bridges older NBT-based items with 1.20.6+ structured components. `ItemType.cs` plus `ItemPalettes/*.cs` provide version-specific item ID mapping. Enchantment, effects, and villager-trade files add higher-level semantics on top of raw inventory data. | `Inventory/Container.cs`, `Inventory/ContainerType.cs`, `Inventory/Item.cs`, `Inventory/ItemMovingHelper.cs`, `Inventory/ItemType.cs`, `Inventory/ItemPalettes/*.cs`, `Inventory/EnchantmentMapping.cs`, `Inventory/VillagerTrade.cs` |
|
||||
| `MinecraftClient/Physics/` | Vanilla-accurate per-tick physics engine. `PlayerPhysics.cs` mirrors vanilla `Entity.move()`, `LivingEntity.aiStep()/travel()`, and `Player.travel()` logic at 20 TPS, handling ground/air/water/lava/creative-fly travel, jumping, sprint-jump boost, climbing, sneak-edge-detection, friction, drag, gravity, slow-falling, and levitation. `CollisionDetector.cs` resolves full AABB collisions against the block world including step-up, mirroring vanilla axis-separated resolution. `BlockShapes.cs` maps block-state IDs to collision AABBs using PrismarineJS data from `BlockShapeData.json`. `Vec3d.cs` and `Aabb.cs` provide the geometric primitives. `MovementInput.cs` captures player input state. | `Physics/PlayerPhysics.cs`, `Physics/PhysicsConsts.cs`, `Physics/CollisionDetector.cs`, `Physics/BlockShapes.cs`, `Physics/BlockShapeData.json`, `Physics/Vec3d.cs`, `Physics/Aabb.cs`, `Physics/MovementInput.cs` |
|
||||
|
||||
### Commands And Extensions
|
||||
|
||||
| Module | What It Owns | Important Files |
|
||||
| --- | --- | --- |
|
||||
| `MinecraftClient/Commands/` and `MinecraftClient/CommandHandler/` | Internal MCC command system built on Brigadier. Commands are discovered by reflection from `MinecraftClient.Commands` in `McClient.LoadCommands()`. Each file in `Commands/` registers one internal command. `ArgumentType/*.cs` provides typed Brigadier arguments and completion sources for accounts, bots, items, locations, scripts, inventories, and more. `Patch/*.cs` carries MCC-specific Brigadier extensions, and `CmdResult.cs` is the command execution result object. | `Command.cs`, `Commands/*.cs`, `CommandHandler/MccArguments.cs`, `CommandHandler/CmdResult.cs`, `CommandHandler/ArgumentType/*.cs`, `CommandHandler/Patch/*.cs` |
|
||||
| `MinecraftClient/ChatBots/` | Built-in bots and bridges loaded from config through `McClient.RegisterBots()`. The folder mixes gameplay automation (`AutoAttack`, `AutoDig`, `AutoEat`, `AutoFishing`, `Farmer`), utility/logging bots (`ChatLog`, `PlayerListLogger`, `Alerts`), bridges (`DiscordBridge`, `TelegramBridge`, `RemoteControl`), and tooling like `ScriptScheduler`, `Map`, and `ReplayCapture`. | `ChatBots/AutoRelog.cs`, `ChatBots/Farmer.cs`, `ChatBots/FollowPlayer.cs`, `ChatBots/ItemsCollector.cs`, `ChatBots/Map.cs`, `ChatBots/RemoteControl.cs`, `ChatBots/ScriptScheduler.cs`, `ChatBots/DiscordBridge.cs`, `ChatBots/TelegramBridge.cs`, `ChatBots/ReplayCapture.cs` |
|
||||
| `MinecraftClient/Scripting/` | Shared extension boundary for compiled bots and runtime C# scripts. `ChatBot.cs` is the main bot API and lifecycle surface. Built-in bots and `/script` bots use the same event model. `CSharpRunner.cs` parses `//MCCScript` files, compiles them with Roslyn, caches assemblies, and executes them through `CSharpAPI`. `DynamicRun/Builder/*` handles in-memory compilation/load-context plumbing, while `BotMovementLock.cs` coordinates movement ownership between automation pieces. | `Scripting/ChatBot.cs`, `Scripting/CSharpRunner.cs`, `Scripting/BotMovementLock.cs`, `Scripting/AssemblyResolver.cs`, `Scripting/DynamicRun/Builder/Compiler.cs`, `Scripting/DynamicRun/Builder/CompileRunner.cs` |
|
||||
| `MinecraftClient/config/` | Sample runtime assets excluded from compilation. This is the examples/staging area for end-user scripts and standalone bots. `sample-script*.cs` shows supported `/script` patterns (basic, chatbot, world access, HTTP requests, tasks, PM forwarding, extended), while `config/ChatBots/*.cs` are copy/adapt examples rather than built-in bots. | `config/README.md`, `config/sample-script.cs`, `config/sample-script-with-chatbot.cs`, `config/sample-script-with-world-access.cs`, `config/sample-script-with-http-request.cs`, `config/sample-script-with-task.cs`, `config/ChatBots/*.cs` |
|
||||
| `ConsoleInteractive/` | Required git submodule for richer line editing and console UI. MCC uses the submodule's `ConsoleReader`, `ConsoleWriter`, and suggestion UI from `ConsoleIO.cs` and `McClient.cs` when `BasicIO` is not enabled. | `ConsoleInteractive/README.md`, `ConsoleInteractive/ConsoleInteractive/ConsoleInteractive.sln` |
|
||||
|
||||
### Support And Tooling
|
||||
|
||||
| Module | What It Owns | Important Files |
|
||||
| --- | --- | --- |
|
||||
| `MinecraftClient/Logger/`, `MinecraftClient/Proxy/`, `MinecraftClient/Crypto/`, `MinecraftClient/Resources/`, `MinecraftClient/WinAPI/` | Support subsystems under the main app. Logging supports console/file output plus regex filtering. `ProxyHandler.cs` routes update/login/in-game traffic through HTTP or SOCKS proxies. `Crypto/` implements the stream ciphers needed for online-mode protocol encryption. `Resources/` contains UI strings, generated translation accessors, config help text, icons, and embedded Minecraft asset data. `WinAPI/` contains small Windows-only console helpers. | `Logger/FilteredLogger.cs`, `Logger/FileLogLogger.cs`, `Proxy/ProxyHandler.cs`, `Crypto/CryptoHandler.cs`, `Crypto/AesCfb8Stream.cs`, `Resources/Translations/Translations.resx`, `Resources/ConfigComments/ConfigComments.resx`, `Resources/en_us.json`, `WinAPI/ConsoleIcon.cs` |
|
||||
| `docs/` | VuePress documentation site. `.vuepress/config.ts` sets bundler, theme, plugins, and redirects. `.vuepress/configs/**` holds locale and nav wiring. `guide/*.md` contains the user-facing install, usage, bot, and scripting docs. | `docs/.vuepress/config.ts`, `docs/.vuepress/configs/**`, `docs/guide/README.md`, `docs/guide/configuration.md`, `docs/guide/chat-bots.md`, `docs/guide/creating-bots.md`, `docs/guide/creating-text-script.md`, `docs/guide/ai-assisted-development.md` |
|
||||
| `tools/` | Python helpers for Minecraft version adaptation and palette generation. `README.md` is the authoritative workflow. `diff_registries.py` compares versions and validates decompiled data against server reports. The `gen_*` scripts emit the versioned palette source files consumed by `Protocol/`, `Mapping/`, `Inventory/`, and `Physics/`. | `tools/README.md`, `tools/diff_registries.py`, `tools/gen_block_palette.py`, `tools/gen_item_palette.py`, `tools/gen_entity_palette.py`, `tools/gen_entity_metadata_palette.py`, `tools/gen_block_shapes.py`, `tools/gen_command_argument_registry.py` |
|
||||
| `DebugTools/` | Standalone packet/proxy debugging utilities for inspecting traffic and compression behavior outside the main client runtime. | `DebugTools/MinecraftClientProxy/Program.cs`, `DebugTools/MinecraftClientProxy/PacketProxy.cs`, `DebugTools/MinecraftClientProxy/ZlibUtils.cs` |
|
||||
| `MinecraftClientGUI/` | Legacy Windows GUI wrapper around the console app. WinForms shell that launches and communicates with the console executable; not part of the main `net10.0` runtime path. | `MinecraftClientGUI/Program.cs`, `MinecraftClientGUI/Form1.cs`, `MinecraftClientGUI/Form1.Designer.cs`, `MinecraftClientGUI/MinecraftClient.cs` |
|
||||
|
||||
## Engineering Guidance
|
||||
|
||||
Read `docs/guide/ai-assisted-development.md` before starting development work on MCC. It documents the full build-run-test loop, local server harness, repository tools, and standard workflows.
|
||||
|
||||
### DO
|
||||
- Keep startup/config/auth logic in `Program` and connection runtime logic in `McClient` or `Protocol/*`.
|
||||
- Update version support holistically: protocol constants, version mapping, packet palette, block palette, item palette, entity palette, metadata palette, and routing switches.
|
||||
- Use `tools/` and authoritative server data reports when adapting to new Minecraft versions.
|
||||
- Guard optional subsystems with `GetTerrainEnabled()`, `GetInventoryEnabled()`, and `GetEntityHandlingEnabled()` before using them.
|
||||
- For built-in bots, wire all pieces together: bot class, `Settings.ChatBotConfigHealper`, and `McClient.RegisterBots()`.
|
||||
- Keep `Initialize()` for setup/prereq checks and `AfterGameJoined()` for sending chat or commands.
|
||||
- Normalize inbound chat with `GetVerbatim()` before `IsChatMessage()` / `IsPrivateMessage()`.
|
||||
- Clean up commands, plugin channels, threads, timers, and movement locks in `OnUnload()`.
|
||||
- Prefer nullable-aware code, pattern matching, `ArgumentNullException.ThrowIfNull`, `Try*` APIs for expected failures, and `InvokeOnMainThread()` for cross-thread state changes.
|
||||
- Use modern C# 14 features.
|
||||
- Use provided skills proactively depending on the context, read their descriptions to determine when to use them.
|
||||
- All user-facing text (log messages, command output, TUI labels, notifications, help text, error messages) **must** go through the translation system: add entries to `Translations.resx` + `Translations.Designer.cs`, then reference `Translations.key_name` in code. Never hardcode user-visible strings directly in `.cs` files. Use `string.Format(Translations.key, ...)` for parameterized messages. Translation keys follow dot-delimited naming: `<module>.<scope>.<detail>` (e.g. `cmd.inventory.tui_opened`, `tui.inventory.controls`). Pure technical identifiers (class names, protocol constants, color codes) are exempt.
|
||||
|
||||
### DON'T
|
||||
- Don't update only `MCVer2ProtocolVersion()` or only one palette file when adding a new Minecraft version.
|
||||
- Don't send chat in `Initialize()`.
|
||||
- Don't mutate inventory snapshots and expect server-side effects; use handler APIs/window actions.
|
||||
- Don't bypass Brigadier with ad hoc command parsing.
|
||||
- Never modify `ConsoleInteractive/`; treat it as an external required submodule.
|
||||
- Don't start background workers when `Update()` or delayed tasks are sufficient; if you must, stop them on unload/disconnect.
|
||||
- Don't leave movement locks, plugin channels, or dispatcher registrations behind.
|
||||
- Don't trust older docs over current code for supported versions or feature gates. When AGENTS.md, skills, and older docs disagree, prefer current code and current tool behavior, then update the stale source.
|
||||
- Don't hardcode user-facing strings (messages, labels, help text) directly in source code; always use `Translations.*` resources so the text can be localized.
|
||||
- Never use "—" ("em dash"), unless specifically being instructed to do so!
|
||||
- Never generate MCC config files (`.ini`) in the repo root. When `dotnet run --project MinecraftClient -- --help` is used to generate a config template, it writes `MinecraftClient.ini` to the current directory. Always run this command from a system temp directory (e.g. `mktemp -d`) or use `prepare_offline_mcc_config.sh` which already handles output routing.
|
||||
Some files were not shown because too many files have changed in this diff Show more
Loading…
Add table
Add a link
Reference in a new issue