flex: drop scripts/benchmarks/results/stats JSONs

Remove the raw benchmark output JSONs from the PR diff. The writeup
markdowns under scripts/benchmarks/results/*.md keep their filename
anchors as a record of what each table was measured from; anyone who
wants the raw numbers can re-run the benchmark scripts.

41 files removed, 6068 lines deleted.
This commit is contained in:
Daniel Han 2026-04-21 05:52:11 +00:00
commit 8792e5da7b
41 changed files with 0 additions and 6068 deletions

View file

@ -1,25 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": null,
"n_prompts": 128,
"n_decoded_tokens": 56415,
"wall_times_s": [
12.628456736041699,
11.507176020997576,
10.393205604981631,
10.137171553040389,
11.370744699030183
],
"median_wall_s": 11.370744699030183,
"best_wall_s": 10.137171553040389,
"decode_tps_median": 4961.416467719275,
"decode_tps_best": 5565.161811144426,
"max_new_tokens": 512,
"peak_memory_gb": 80.88226366043091,
"sample_completions": [
"Let $h$ be the height of the tetrahedron. Then, the volume of the tetrahedron is $\\frac{1}{3} \\cdot 120 \\cdot h = 40h$.<end_working_out><SOLUTION>400</SOLUTION>",
" \nTo solve this problem, we will use the concept of mass points and the properties of similar triangles. \n\nFirst, let's assign masses to the points based on the given information. Since $M$ is the midpoint of $BC$, we can assign a mass of 1 to both $B$ and $C$. This means that the mass at $M$ is 2 (since $",
" To solve this problem, we need to find the value of \\( n \\) that minimizes the sum \\( \\sum_{i=1}^{n} f(i) \\) under the given conditions. Let's break down the problem step by step.\n\n1. **Understanding the Constraints:**\n - \\( f \\) is a non-negative valued function on \\( \\{1, 2"
]
}

View file

@ -1,15 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": null,
"n_prompts": 128,
"n_decoded_tokens": 58443,
"wall_times_s": [
11.44752091699047,
8.989795534987934
],
"median_wall_s": 11.44752091699047,
"decode_tps": 5105.297507101174,
"max_new_tokens": 512,
"peak_memory_gb": 80.88137865066528
}

View file

@ -1,25 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": null,
"n_prompts": 16,
"n_decoded_tokens": 7446,
"wall_times_s": [
6.069258489005733,
4.580210059997626,
5.2929606509860605,
5.689690829021856,
5.988061942043714
],
"median_wall_s": 5.689690829021856,
"best_wall_s": 4.580210059997626,
"decode_tps_median": 1308.6827076823927,
"decode_tps_best": 1625.6896304891004,
"max_new_tokens": 512,
"peak_memory_gb": 43.700210094451904,
"sample_completions": [
" To solve this problem, we need to determine how many ways we can divide a \\(20 \\times 24\\) rectangle into \\(4 \\times 5\\) rectangles. We will consider rotations and reflections as distinct.\n\nFirst, let's calculate the area of the \\(20 \\times 24\\) rectangle:\n\\[\n20 \\times 24 = 480\n",
" To solve this problem, we need to find the area of the region inside the larger circle \\( C \\) with radius 30 and outside the six smaller congruent circles that form a ring and are each internally tangent to \\( C \\).\n\nFirst, let's denote the radius of each of the six smaller circles as \\( r \\). Since the six smaller circles form a ring and are each externally",
" \nA 10-digit palindrome has the form \\( \\overline{abcdefghij} \\) where \\( a = j \\), \\( b = i \\), \\( c = h \\), \\( d = g \\), \\( e = f \\), and \\( f = e \\). This means the number can be written as \\( \\overline{abcdeedcba} \\).\n\nTo determine"
]
}

View file

@ -1,25 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": null,
"n_prompts": 32,
"n_decoded_tokens": 15164,
"wall_times_s": [
7.840532291971613,
5.729173633968458,
6.219996218045708,
5.283988946001045,
4.839393497037236
],
"median_wall_s": 5.729173633968458,
"best_wall_s": 4.839393497037236,
"decode_tps_median": 2646.804053920124,
"decode_tps_best": 3133.4505055816758,
"max_new_tokens": 512,
"peak_memory_gb": 43.90812540054321,
"sample_completions": [
" To solve this problem, we need to find the maximum value of \\(a\\) such that the line \\(y = mx + 2\\) does not pass through any lattice points for \\(0 < x \\leq 100\\) when \\(\\frac{1}{2} < m < a\\).\n\nFirst, let's consider the condition that the line \\(y = mx + 2",
" \nTo find the area of the smaller square, we need to determine its side length. Let's denote the side length of the smaller square as \\( s \\).\n\nFrom the diagram, we can see that the larger square has a side length of 6. The smaller square is inscribed within the larger square such that its vertices touch the midpoints of the sides of the larger square. This means",
" \nTo solve this problem, we need to consider all possible pairs of special fractions \\(\\frac{a}{b}\\) and \\(\\frac{c}{d}\\) where \\(a + b = 15\\) and \\(c + d = 15\\). We will then find the distinct integers that can be written as the sum of these two fractions.\n\nFirst, let's list"
]
}

View file

@ -1,15 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": null,
"n_prompts": 32,
"n_decoded_tokens": 14777,
"wall_times_s": [
9.13400061900029,
6.742950548010413
],
"median_wall_s": 9.13400061900029,
"decode_tps": 1617.8015106832052,
"max_new_tokens": 512,
"peak_memory_gb": 43.90812540054321
}

View file

@ -1,15 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": "outputs/lora_rank32_fresh",
"n_prompts": 32,
"n_decoded_tokens": 14893,
"wall_times_s": [
8.811616113001946,
6.381639341008849
],
"median_wall_s": 8.811616113001946,
"decode_tps": 1690.1553368881664,
"max_new_tokens": 512,
"peak_memory_gb": 43.90812540054321
}

View file

@ -1,25 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": "outputs/lora_rank32_fresh",
"n_prompts": 64,
"n_decoded_tokens": 29566,
"wall_times_s": [
8.612908595998306,
6.458285094005987,
6.398469406005461,
6.125680990982801,
6.644617859972641
],
"median_wall_s": 6.458285094005987,
"best_wall_s": 6.125680990982801,
"decode_tps_median": 4577.99548481385,
"decode_tps_best": 4826.565412649157,
"max_new_tokens": 512,
"peak_memory_gb": 44.21115064620972,
"sample_completions": [
"Let the roots of the equation be $r, r^2, r^3, r^4, r^5$ in geometric progression. By Vieta's formulas, the sum of the roots is $r + r^2 + r^3 + r^4 + r^5 = 180$. Dividing both sides by $r^5$, we get $1",
"First, let's represent the given number in a more manageable form. The number \\(1\\underbrace{00\\ldots 0}_{100\\text{ zeros}}1\\underbrace{00\\ldots 0}_{100\\text{ zeros}}1\\) can be written as \\(10^{201} + 10^{10",
" \nTo solve this problem, we need to find the number of integers \\( n \\) in the range \\( 1 \\leq n \\leq 2016 \\) such that the remainder when \\( n \\) is divided by 20 is smaller than the remainder when \\( n \\) is divided by 16. Let's denote the remainder when \\( n \\)"
]
}

View file

@ -1,25 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": "outputs/lora_rank32_fresh",
"n_prompts": 64,
"n_decoded_tokens": 29566,
"wall_times_s": [
8.423556671012193,
6.48478461895138,
6.419383131025825,
6.145251678011846,
6.663184700999409
],
"median_wall_s": 6.48478461895138,
"best_wall_s": 6.145251678011846,
"decode_tps_median": 4559.28789270737,
"decode_tps_best": 4811.1943251713,
"max_new_tokens": 512,
"peak_memory_gb": 44.21115064620972,
"sample_completions": [
"Let the roots of the equation be $r, r^2, r^3, r^4, r^5$ in geometric progression. By Vieta's formulas, the sum of the roots is $r + r^2 + r^3 + r^4 + r^5 = 180$. Dividing both sides by $r^5$, we get $1",
"First, let's represent the given number in a more manageable form. The number \\(1\\underbrace{00\\ldots 0}_{100\\text{ zeros}}1\\underbrace{00\\ldots 0}_{100\\text{ zeros}}1\\) can be written as \\(10^{201} + 10^{10",
" \nTo solve this problem, we need to find the number of integers \\( n \\) in the range \\( 1 \\leq n \\leq 2016 \\) such that the remainder when \\( n \\) is divided by 20 is smaller than the remainder when \\( n \\) is divided by 16. Let's denote the remainder when \\( n \\)"
]
}

View file

@ -1,25 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": "outputs/lora_rank32_fresh",
"n_prompts": 64,
"n_decoded_tokens": 29601,
"wall_times_s": [
11.575495404948015,
6.411582662025467,
6.848826738016214,
7.4589985449565575,
6.9681332929758355
],
"median_wall_s": 6.9681332929758355,
"best_wall_s": 6.411582662025467,
"decode_tps_median": 4248.053066068501,
"decode_tps_best": 4616.800805723188,
"max_new_tokens": 512,
"peak_memory_gb": 44.22274446487427,
"sample_completions": [
"First, let's find the sum of the numbers in Amanda's list. The sum of the first n even numbers is given by the formula n(n+1). In this case, n = 50 (since there are 50 even numbers from 2 to 100). So, the sum of Amanda's list is 50(50+1) = ",
" \nTo find the area of the smaller square, we need to determine its side length. Let's denote the side length of the smaller square as \\( s \\).\n\nFrom the diagram, we can see that the larger square has a side length of 6. The smaller square is inscribed within the larger square such that its vertices touch the midpoints of the sides of the larger square. This means",
" To solve the problem, we need to find the number of ordered pairs \\((x, y)\\) of positive integers that satisfy the inequalities \\(x \\le 2y \\le 60\\) and \\(y \\le 2x \\le 60\\).\n\nFirst, let's rewrite the inequalities in a more convenient form:\n1. \\(x \\le 2y \\le"
]
}

View file

@ -1,25 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": "outputs/lora_rank32_fresh",
"n_prompts": 64,
"n_decoded_tokens": 29553,
"wall_times_s": [
8.755648983002175,
6.662199873011559,
6.631406380969565,
6.668323302001227,
5.8994951589847915
],
"median_wall_s": 6.662199873011559,
"best_wall_s": 5.8994951589847915,
"decode_tps_median": 4435.922152338692,
"decode_tps_best": 5009.411687539311,
"max_new_tokens": 512,
"peak_memory_gb": 44.21115064620972,
"sample_completions": [
" \nTo find the minimum sum of the labels of the eight chosen squares, we need to consider the arrangement of the numbers on the chessboard. The answer is 10.",
"Let's denote the angles $\\angle BAP = \\angle PAQ = \\angle QAC = \\theta$. Since $AP$ and $AQ$ trisect $\\angle A$, we have $\\angle BAC = 3\\theta$.\n\nWe will use the Angle Bisector Theorem and the Law of Sines to find the ratio $\\frac{SOLUTION}</SOLUTION>",
"Let's denote the number of pages in the first volume as $x$. Then, the number of pages in the second volume is $x + 50$, and the number of pages in the third volume is $1.5(x + 50)$.\n\nThe sum of the page numbers on the first pages of the three volumes is $1 + (x + 1) + ("
]
}

View file

@ -1,30 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": "outputs/lora_rank32_fresh",
"n_prompts": 64,
"n_decoded_tokens": 28565,
"wall_times_s": [
8.299892835028004,
6.869692277978174,
6.813798578979913,
5.0465612940024585,
5.072170660016127,
5.903178220964037,
6.141771270020399,
5.922227941977326,
7.205250932951458,
6.960845094989054
],
"median_wall_s": 6.813798578979913,
"best_wall_s": 5.0465612940024585,
"decode_tps_median": 4192.228412521762,
"decode_tps_best": 5660.289915421779,
"max_new_tokens": 512,
"peak_memory_gb": 44.21115064620972,
"sample_completions": [
"First, let's analyze the problem. We are given a natural number $a$ and we need to find the number of elements $b$ in the set $\\{ b \\in \\mathbb{N} \\mid a + b \\text{ is a divisor of } ab \\}$. We need to find the maximum value of $M(a)$ for $a \\leq 1",
"Let the roots of the equation be $r, r^2, r^3, r^4, r^5$ in geometric progression. By Vieta's formulas, the sum of the roots is $r + r^2 + r^3 + r^4 + r^5 = 180$. Dividing both sides by $r^5$, we get $1",
"First, we need to find the total number of possible triples of positive integers $(a, b, c)$ with $1 \\leq a, b, c \\leq 5$. Since each of $a$, $b$, and $c$ can take on 5 different values, the total number of possible triples is $5 \\times 5 \\times 5 = 1"
]
}

View file

@ -1,25 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": "outputs/lora_rank32_fresh",
"n_prompts": 64,
"n_decoded_tokens": 29019,
"wall_times_s": [
14.41304371797014,
6.878086202952545,
6.826645737979561,
5.051277688005939,
5.077490734984167
],
"median_wall_s": 6.826645737979561,
"best_wall_s": 5.051277688005939,
"decode_tps_median": 4250.843110043757,
"decode_tps_best": 5744.883134994633,
"max_new_tokens": 512,
"peak_memory_gb": 44.21115064620972,
"sample_completions": [
"First, we need to find the length of segment $ABCD is 17.8 units. The length of segment $DB$ is 12.8 units.",
" \nTo solve this problem, we need to consider the different ways people can stand or sit around the table without having two adjacent people standing. Let's denote standing as S and sitting as T. We have 8 people, so there are 2^8 = 256 possible outcomes when flipping the coins.\n\nWe want to find the number of valid configurations where no two adjacent people stand.",
"Let the roots of the equation be $r, r^2, r^3, r^4, r^5$ in geometric progression. By Vieta's formulas, the sum of the roots is $r + r^2 + r^3 + r^4 + r^5 = 180$. Dividing both sides by $r^5$, we get $1"
]
}

View file

@ -1,25 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": "outputs/lora_rank32_fresh",
"n_prompts": 64,
"n_decoded_tokens": 29019,
"wall_times_s": [
8.298801471013576,
6.888721746974625,
6.2715082070208155,
4.575723551039118,
4.594870681001339
],
"median_wall_s": 6.2715082070208155,
"best_wall_s": 4.575723551039118,
"decode_tps_median": 4627.116642774041,
"decode_tps_best": 6341.947820123435,
"max_new_tokens": 512,
"peak_memory_gb": 44.21115064620972,
"sample_completions": [
"First, we need to find the length of segment $ABCD is 17.8 units. The length of segment $DB$ is 12.8 units.",
" \nTo solve this problem, we need to consider the different ways people can stand or sit around the table without having two adjacent people standing. Let's denote standing as S and sitting as T. We have 8 people, so there are 2^8 = 256 possible outcomes when flipping the coins.\n\nWe want to find the number of valid configurations where no two adjacent people stand.",
"Let the roots of the equation be $r, r^2, r^3, r^4, r^5$ in geometric progression. By Vieta's formulas, the sum of the roots is $r + r^2 + r^3 + r^4 + r^5 = 180$. Dividing both sides by $r^5$, we get $1"
]
}

View file

@ -1,25 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": "outputs/lora_rank32_fresh",
"n_prompts": 64,
"n_decoded_tokens": 28495,
"wall_times_s": [
8.62534567998955,
6.866325525043067,
6.2769941369770095,
5.074044068984222,
6.802140125015285
],
"median_wall_s": 6.802140125015285,
"best_wall_s": 5.074044068984222,
"decode_tps_median": 4189.122757881435,
"decode_tps_best": 5615.836128460044,
"max_new_tokens": 512,
"peak_memory_gb": 44.21115064620972,
"sample_completions": [
"First, let's find the sum of the numbers in Amanda's list. The sum of the first n even numbers is given by the formula n(n+1). In this case, n = 50 (since there are 50 even numbers from 2 to 100). So, the sum of Amanda's list is 50(50+1) = ",
"Let's denote the number of pages in the first volume as $x$. Then, the number of pages in the second volume is $x + 50$, and the number of pages in the third volume is $1.5(x + 50)$.\n\nThe sum of the page numbers on the first pages of the three volumes is $1 + (x + 1) + (",
"Let the roots of the equation be $r, r^2, r^3, r^4, r^5$ in geometric progression. By Vieta's formulas, the sum of the roots is $r + r^2 + r^3 + r^4 + r^5 = 180$. Dividing both sides by $r^5$, we get $1"
]
}

View file

@ -1,25 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": null,
"n_prompts": 64,
"n_decoded_tokens": 27985,
"wall_times_s": [
8.14494420203846,
6.53967270400608,
5.112081662984565,
5.535065989010036,
6.894235474988818
],
"median_wall_s": 6.53967270400608,
"best_wall_s": 5.112081662984565,
"decode_tps_median": 4279.266144750168,
"decode_tps_best": 5474.286571482827,
"max_new_tokens": 512,
"peak_memory_gb": 44.21115064620972,
"sample_completions": [
"First, let's find the sum of the numbers in Amanda's list. The sum of the first n even numbers is given by the formula n(n+1). In this case, n = 50 (since there are 50 even numbers from 2 to 100). So, the sum of Amanda's list is 50(50+1) = ",
"Let the roots of the equation be $r, r^2, r^3, r^4, r^5$ in geometric progression. By Vieta's formulas, the sum of the roots is $r + r^2 + r^3 + r^4 + r^5 = 180$. Dividing both sides by $r^5$, we get $1",
"First, let's find the angle \\( \\angle AOB into three equal parts. The area of each smaller triangle is:\n\\[ \\frac{\\sqrt{3}}{12} \\text{ triangle} = \\frac{\\sqrt{3}/4 \\]\n\nNow, let's find the value of \\( k + m + n \\). We have:\n\\[ k = 1 \\]\n\\["
]
}

View file

@ -1,23 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": "outputs/lora_rank32_fresh",
"n_prompts": 64,
"n_decoded_tokens": 30184,
"wall_times_s": [
9.522195567958988,
7.053195148007944,
7.056460196967237
],
"median_wall_s": 7.056460196967237,
"best_wall_s": 7.053195148007944,
"decode_tps_median": 4277.498796489016,
"decode_tps_best": 4279.478926444416,
"max_new_tokens": 512,
"peak_memory_gb": 44.21115064620972,
"sample_completions": [
"First, we need to determine how many $4 \\times 5$ rectangles can fit into a $20 \\times 24$ rectangle. We can do this by dividing the dimensions of the larger rectangle by the dimensions of the smaller rectangle.\n\nFor the width, we have $20 \\div 4 = 5$ rectangles that can fit.\nFor the height, we have $",
"First, we need to find the total number of letters in the word \"FLUFFY\". There are 6 letters in total. \n\nNext, we need to find the number of distinct arrangements of these 6 letters. Since there are 6 letters, the total number of arrangements is 6! (6 factorial), which is equal to 6 x 5 x 4 x 3",
"Let the common ratio of the geometric sequence be $r$. Then the second term is $\\frac{3}{4}r=15$, so $r=20$. The $n$th term of the sequence is $\\frac{3}{4}(20)^{n-1}$. We want to find the smallest $n$ such that $\\frac{3}{4}("
]
}

View file

@ -1,25 +0,0 @@
{
"backend": "qwen3_flex",
"capture_cudagraph": true,
"lora_adapter": null,
"n_prompts": 8,
"n_decoded_tokens": 4096,
"wall_times_s": [
8.54062199202599,
8.53521726100007,
7.605857895978261,
6.020388883014675,
8.549159363028593
],
"median_wall_s": 8.53521726100007,
"best_wall_s": 6.020388883014675,
"decode_tps_median": 479.89405245907847,
"decode_tps_best": 680.3547211968393,
"max_new_tokens": 512,
"peak_memory_gb": 43.68726634979248,
"sample_completions": [
"First, we need to find the length of the legs of the trapezoid. Since the trapezoid is isosceles, the legs are equal in length. Let's call the length of each leg $x$. We can use the Pythagorean theorem to find $x$.\n\nThe height of the trapezoid is 3, and the difference between the lengths",
"Let $P(x)$ be a monic polynomial of degree $2023$ such that $P(k) = k^{2023}P(1-\\frac{1}{k})$ for every positive integer $1 \\leq k \\leq 2023$. We want to find $P(-1)$ in the form $\\frac{a}{b",
" To solve this problem, we need to determine the maximum value of \\(a\\) such that the line \\(y = mx + 2\\) does not pass through any lattice points for \\(0 < x \\leq 100\\) when \\(\\frac{1}{2} < m < a\\).\n\nFirst, let's consider the condition for the line \\(y = mx + 2"
]
}

View file

@ -1,352 +0,0 @@
[
{
"step": 1,
"loss": -0.0862,
"grad_norm": 716.0,
"learning_rate": 0.0,
"num_tokens": 4262.0,
"completions/mean_length": 953.5,
"completions/min_length": 824.0,
"completions/max_length": 1092.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 953.5,
"completions/min_terminated_length": 824.0,
"completions/max_terminated_length": 1092.0,
"rewards/match_format_exactly/mean": 2.25,
"rewards/match_format_exactly/std": 1.5,
"rewards/match_format_approximately/mean": 1.125,
"rewards/match_format_approximately/std": 0.75,
"rewards/check_answer/mean": -1.25,
"rewards/check_answer/std": 2.1794495582580566,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": 0.625,
"reward_std": 3.4731109142303467,
"frac_reward_zero_std": 0.0,
"entropy": 0.1351587027311325,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 7.868439688409789e-05,
"time_ms": 56644.44096497027,
"memory_mb": 57451.41162109375,
"memory_gb": 56.104894161224365
},
{
"step": 2,
"loss": 0.041,
"grad_norm": 186.0,
"learning_rate": 5e-06,
"num_tokens": 6762.0,
"completions/mean_length": 536.0,
"completions/min_length": 492.0,
"completions/max_length": 603.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 536.0,
"completions/min_terminated_length": 492.0,
"completions/max_terminated_length": 603.0,
"rewards/match_format_exactly/mean": 0.75,
"rewards/match_format_exactly/std": 1.5,
"rewards/match_format_approximately/mean": 0.375,
"rewards/match_format_approximately/std": 0.75,
"rewards/check_answer/mean": -2.125,
"rewards/check_answer/std": 0.25,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": -2.5,
"reward_std": 2.0,
"frac_reward_zero_std": 0.0,
"entropy": 0.05973631516098976,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.00015736879376819577,
"time_ms": 29505.720576969907,
"memory_mb": 53805.20556640625,
"memory_gb": 52.5441460609436
},
{
"step": 3,
"loss": 0.0,
"grad_norm": 0.0,
"learning_rate": 4.444444444444444e-06,
"num_tokens": 10099.0,
"completions/mean_length": 657.25,
"completions/min_length": 436.0,
"completions/max_length": 1302.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 657.25,
"completions/min_terminated_length": 436.0,
"completions/max_terminated_length": 1302.0,
"rewards/match_format_exactly/mean": 3.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": 1.5,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -2.5,
"rewards/check_answer/std": 0.0,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": 0.5,
"reward_std": 0.0,
"frac_reward_zero_std": 1.0,
"entropy": 0.06822667270898819,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.00023605319065229366,
"time_ms": 62939.08593803644,
"memory_mb": 59249.8505859375,
"memory_gb": 57.86118221282959
},
{
"step": 4,
"loss": 0.0855,
"grad_norm": 274.0,
"learning_rate": 3.88888888888889e-06,
"num_tokens": 13644.0,
"completions/mean_length": 721.25,
"completions/min_length": 441.0,
"completions/max_length": 988.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 721.25,
"completions/min_terminated_length": 441.0,
"completions/max_terminated_length": 988.0,
"rewards/match_format_exactly/mean": 1.5,
"rewards/match_format_exactly/std": 1.7320507764816284,
"rewards/match_format_approximately/mean": 0.75,
"rewards/match_format_approximately/std": 0.8660253882408142,
"rewards/check_answer/mean": -2.25,
"rewards/check_answer/std": 0.28867512941360474,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": -1.5,
"reward_std": 2.309401035308838,
"frac_reward_zero_std": 0.0,
"entropy": 0.25085046887397766,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.00031473758753639155,
"time_ms": 49076.82833302533,
"memory_mb": 56826.26611328125,
"memory_gb": 55.49440050125122
},
{
"step": 5,
"loss": 0.0162,
"grad_norm": 143.0,
"learning_rate": 3.3333333333333333e-06,
"num_tokens": 15442.0,
"completions/mean_length": 293.5,
"completions/min_length": 246.0,
"completions/max_length": 365.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 293.5,
"completions/min_terminated_length": 246.0,
"completions/max_terminated_length": 365.0,
"rewards/match_format_exactly/mean": 3.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": 1.5,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -3.0,
"rewards/check_answer/std": 1.0,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": 0.0,
"reward_std": 1.0,
"frac_reward_zero_std": 0.0,
"entropy": 0.0809403508901596,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.00039342198442048943,
"time_ms": 18123.881562962197,
"memory_mb": 52261.25830078125,
"memory_gb": 51.03638505935669
},
{
"step": 6,
"loss": 0.0,
"grad_norm": 0.0,
"learning_rate": 2.7777777777777783e-06,
"num_tokens": 21877.0,
"completions/mean_length": 1511.75,
"completions/min_length": 1112.0,
"completions/max_length": 1846.0,
"completions/clipped_ratio": 0.5,
"completions/mean_terminated_length": 1177.5,
"completions/min_terminated_length": 1112.0,
"completions/max_terminated_length": 1243.0,
"rewards/match_format_exactly/mean": 0.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": -3.0,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -2.0,
"rewards/check_answer/std": 0.0,
"rewards/check_numbers/mean": -2.5,
"rewards/check_numbers/std": 0.0,
"reward": -7.5,
"reward_std": 0.0,
"frac_reward_zero_std": 1.0,
"entropy": 0.2572544813156128,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0004721063813045873,
"time_ms": 92771.2398529984,
"memory_mb": 63378.49365234375,
"memory_gb": 61.89306020736694
},
{
"step": 7,
"loss": 0.0736,
"grad_norm": 74.0,
"learning_rate": 2.222222222222222e-06,
"num_tokens": 24635.0,
"completions/mean_length": 546.5,
"completions/min_length": 466.0,
"completions/max_length": 585.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 546.5,
"completions/min_terminated_length": 466.0,
"completions/max_terminated_length": 585.0,
"rewards/match_format_exactly/mean": 3.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": 1.5,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -1.5,
"rewards/check_answer/std": 2.0,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": 1.5,
"reward_std": 2.0,
"frac_reward_zero_std": 0.0,
"entropy": 0.05001620948314667,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0005507907781886852,
"time_ms": 30032.14380296413,
"memory_mb": 53704.75830078125,
"memory_gb": 52.44605302810669
},
{
"step": 8,
"loss": -0.0475,
"grad_norm": 97.0,
"learning_rate": 1.6666666666666667e-06,
"num_tokens": 27393.0,
"completions/mean_length": 622.5,
"completions/min_length": 533.0,
"completions/max_length": 696.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 622.5,
"completions/min_terminated_length": 533.0,
"completions/max_terminated_length": 696.0,
"rewards/match_format_exactly/mean": 0.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": -1.5,
"rewards/match_format_approximately/std": 1.7320507764816284,
"rewards/check_answer/mean": -2.0,
"rewards/check_answer/std": 0.0,
"rewards/check_numbers/mean": -0.75,
"rewards/check_numbers/std": 2.872281312942505,
"reward": -4.25,
"reward_std": 4.27200174331665,
"frac_reward_zero_std": 0.0,
"entropy": 0.08264704048633575,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0006294751750727831,
"time_ms": 36016.123837034684,
"memory_mb": 54504.00634765625,
"memory_gb": 53.22656869888306
},
{
"step": 9,
"loss": -0.1881,
"grad_norm": 274.0,
"learning_rate": 1.111111111111111e-06,
"num_tokens": 29461.0,
"completions/mean_length": 420.0,
"completions/min_length": 327.0,
"completions/max_length": 647.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 420.0,
"completions/min_terminated_length": 327.0,
"completions/max_terminated_length": 647.0,
"rewards/match_format_exactly/mean": 1.5,
"rewards/match_format_exactly/std": 1.7320507764816284,
"rewards/match_format_approximately/mean": 0.0,
"rewards/match_format_approximately/std": 2.1213202476501465,
"rewards/check_answer/mean": -1.25,
"rewards/check_answer/std": 1.8484227657318115,
"rewards/check_numbers/mean": -1.75,
"rewards/check_numbers/std": 0.5,
"reward": -1.5,
"reward_std": 5.16397762298584,
"frac_reward_zero_std": 0.0,
"entropy": 0.22891533374786377,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0007081595719568809,
"time_ms": 35705.52668598248,
"memory_mb": 54149.77783203125,
"memory_gb": 52.88064241409302
},
{
"step": 10,
"loss": 0.0444,
"grad_norm": 236.0,
"learning_rate": 5.555555555555555e-07,
"num_tokens": 33785.0,
"completions/mean_length": 913.0,
"completions/min_length": 832.0,
"completions/max_length": 998.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 913.0,
"completions/min_terminated_length": 832.0,
"completions/max_terminated_length": 998.0,
"rewards/match_format_exactly/mean": 0.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": -2.25,
"rewards/match_format_approximately/std": 1.5,
"rewards/check_answer/mean": -2.0,
"rewards/check_answer/std": 0.0,
"rewards/check_numbers/mean": -2.25,
"rewards/check_numbers/std": 0.5,
"reward": -6.5,
"reward_std": 2.0,
"frac_reward_zero_std": 0.0,
"entropy": 0.1578676998615265,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0007868439688409789,
"time_ms": 53879.347916983534,
"memory_mb": 56904.7578125,
"memory_gb": 55.57105255126953
}
]

View file

@ -1,64 +0,0 @@
{
"backend": "cb_paged",
"max_steps": 10,
"train_wall_s": 466.01091928296955,
"median_step_ms_post_warmup": 36016.123837034684,
"n_logged_steps": 10,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
},
"losses": [
-0.0862,
0.041,
0.0,
0.0855,
0.0162,
0.0,
0.0736,
-0.0475,
-0.1881,
0.0444
],
"rewards": [
0.625,
-2.5,
0.5,
-1.5,
0.0,
-7.5,
1.5,
-4.25,
-1.5,
-6.5
],
"kls": [],
"grad_norms": [
716.0,
186.0,
0.0,
274.0,
143.0,
0.0,
74.0,
97.0,
274.0,
236.0
],
"step_times_ms": [
56644.44096497027,
29505.720576969907,
62939.08593803644,
49076.82833302533,
18123.881562962197,
92771.2398529984,
30032.14380296413,
36016.123837034684,
35705.52668598248,
53879.347916983534
],
"peak_memory_gb": 55.57105255126953,
"logs_path": "logs/grpo_cb_paged_10.json"
}

File diff suppressed because it is too large Load diff

View file

@ -1,144 +0,0 @@
{
"backend": "cb_paged",
"max_steps": 30,
"train_wall_s": 1564.461075181025,
"median_step_ms_post_warmup": 39820.741517003626,
"n_logged_steps": 30,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
},
"losses": [
-0.0862,
0.041,
0.1442,
0.4175,
0.0106,
0.0,
-0.053,
-0.0433,
0.0257,
0.1615,
0.2728,
0.5823,
0.0,
0.2309,
0.2529,
-0.0039,
0.0829,
0.3114,
0.0053,
0.0803,
0.0498,
-0.0079,
-0.0559,
0.0286,
0.1726,
0.4091,
-0.102,
-0.0477,
0.1184,
0.2766
],
"rewards": [
0.625,
-2.5,
-1.5,
-5.5,
0.0,
-7.5,
-4.5,
-3.0,
-4.5,
-3.5,
-5.25,
-5.5,
-3.5,
-2.0,
-1.375,
6.75,
-0.375,
-6.5,
-1.125,
-1.5,
5.125,
9.875,
2.625,
-4.5,
9.875,
-3.625,
-6.5,
-5.0,
-3.5,
-0.625
],
"kls": [],
"grad_norms": [
716.0,
184.0,
664.0,
632.0,
36.25,
0.0,
78.5,
418.0,
51.75,
330.0,
244.0,
920.0,
0.0,
380.0,
368.0,
57.0,
185.0,
296.0,
134.0,
252.0,
326.0,
101.0,
213.0,
128.0,
78.0,
276.0,
290.0,
752.0,
576.0,
800.0
],
"step_times_ms": [
55554.10714598838,
29354.5744830044,
38962.64252299443,
91278.88684801292,
12430.784016032703,
90694.74714196986,
30823.07043799665,
37861.39360797824,
18964.377576019615,
55770.05596697563,
63336.94338303758,
88045.46197201125,
25137.27058301447,
32668.75728900777,
89534.06670497498,
30537.121773988474,
90275.84768499946,
91312.23277695244,
22925.990092975553,
39714.41951999441,
35041.47868498694,
22683.228761015926,
29075.48440602841,
39820.741517003626,
25694.13814501604,
90294.3110250053,
56216.15647501312,
47257.37831299193,
90669.61020795861,
91291.07107501477
],
"peak_memory_gb": 61.89242887496948,
"logs_path": "logs/grpo_cb_paged_30.json"
}

File diff suppressed because it is too large Load diff

View file

@ -1,175 +0,0 @@
{
"backend": "unsloth_fi_false",
"max_steps": 30,
"train_wall_s": 1165.377411015972,
"median_step_ms_post_warmup": 41297.174014966,
"n_logged_steps": 30,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
},
"losses": [
0.0,
-0.0893,
-0.1912,
0.4302,
-0.0144,
0.0,
0.0036,
0.0,
-0.1558,
0.018,
0.1468,
0.0,
-0.2741,
0.096,
0.0103,
0.0398,
0.0139,
0.3858,
0.9674,
0.2627,
0.0582,
0.0,
0.6477,
-0.109,
0.0296,
0.0185,
0.0,
0.014,
-0.0851,
0.0898
],
"rewards": [
0.5,
-6.5,
-4.5,
-4.5,
9.375,
-7.5,
-3.5,
-7.5,
-2.625,
-4.5,
0.625,
-3.5,
-6.5,
4.125,
9.875,
0.0,
-5.5,
-3.5,
-2.375,
-2.5,
8.0,
11.5,
-1.0,
4.75,
10.125,
0.5,
-7.5,
-2.5,
-1.125,
5.0
],
"kls": [
0.0,
0.0,
0.006437055766582489,
0.007001329679042101,
0.0032435881439596415,
0.00288483127951622,
0.0062899235635995865,
0.0008946225862018764,
0.0027981880120933056,
0.0029807849787175655,
0.00796814076602459,
0.001598043367266655,
0.003610937623307109,
0.0050869532860815525,
0.0019661628175526857,
0.0062008751556277275,
0.003240604419261217,
0.00907122902572155,
0.00969112291932106,
0.006743168458342552,
0.00867636501789093,
0.004581788554787636,
0.007497473154217005,
0.007104712072759867,
0.00291788624599576,
0.011633609421551228,
0.0014683930203318596,
0.006795317865908146,
0.005157058592885733,
0.003586029401049018
],
"grad_norms": [
0.0,
0.6121569275856018,
0.5873263478279114,
0.4428107738494873,
0.9299039244651794,
0.0014747647801414132,
0.6682185530662537,
0.00014817823830526322,
0.2690228223800659,
0.4899609088897705,
0.46429336071014404,
0.0002485642035026103,
0.4754463732242584,
0.7229195237159729,
0.49645838141441345,
0.3055652379989624,
0.3895750939846039,
0.5219303369522095,
0.3664180636405945,
0.453957200050354,
0.5753984451293945,
0.001453780336305499,
0.45908382534980774,
1.2539762258529663,
0.6490684747695923,
0.6853195428848267,
0.0011842504609376192,
0.6820011734962463,
0.42553478479385376,
0.259000688791275
],
"step_times_ms": [
47513.93520901911,
26899.73210898461,
41262.48180796392,
66495.43262599036,
10969.204296008684,
60708.51903402945,
19134.767919022124,
41297.174014966,
61718.76015001908,
46218.60432100948,
63510.61685796594,
10268.670362012926,
43201.109810965136,
22460.022343031596,
21824.53149399953,
27588.505985040683,
44275.010473967995,
60585.87563998299,
60958.746705029625,
60960.08875203552,
20987.195259018335,
13986.805958964396,
60834.94512201287,
27714.821267989464,
16654.144487984013,
23337.049510038923,
29055.362954968587,
28389.59298102418,
43584.85947694862,
60590.078279026784
],
"peak_memory_gb": 10.659695148468018,
"logs_path": "logs/grpo_fi_false_30.json"
}

View file

@ -1,362 +0,0 @@
[
{
"step": 1,
"loss": 0.0,
"grad_norm": 0.0,
"learning_rate": 0.0,
"num_tokens": 3693.0,
"completions/mean_length": 811.25,
"completions/min_length": 779.0,
"completions/max_length": 856.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 811.25,
"completions/min_terminated_length": 779.0,
"completions/max_terminated_length": 856.0,
"rewards/match_format_exactly/mean": 3.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": 1.5,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -2.5,
"rewards/check_answer/std": 0.0,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": 0.5,
"reward_std": 0.0,
"frac_reward_zero_std": 1.0,
"completion_length": 811.25,
"kl": 0.0,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 7.868439688409789e-05,
"time_ms": 77897.29859499494,
"memory_mb": 9442.14013671875,
"memory_gb": 9.220839977264404
},
{
"step": 2,
"loss": -0.0893,
"grad_norm": 0.6125104427337646,
"learning_rate": 5e-06,
"num_tokens": 6238.0,
"completions/mean_length": 547.25,
"completions/min_length": 487.0,
"completions/max_length": 645.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 547.25,
"completions/min_terminated_length": 487.0,
"completions/max_terminated_length": 645.0,
"rewards/match_format_exactly/mean": 0.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": -2.25,
"rewards/match_format_approximately/std": 1.5,
"rewards/check_answer/mean": -2.0,
"rewards/check_answer/std": 0.0,
"rewards/check_numbers/mean": -2.25,
"rewards/check_numbers/std": 0.5,
"reward": -6.5,
"reward_std": 2.0,
"frac_reward_zero_std": 0.0,
"completion_length": 547.25,
"kl": 0.0,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.00015736879376819577,
"time_ms": 21634.983669035137,
"memory_mb": 9076.26025390625,
"memory_gb": 8.863535404205322
},
{
"step": 3,
"loss": -0.1386,
"grad_norm": 0.5993297696113586,
"learning_rate": 4.444444444444444e-06,
"num_tokens": 10165.0,
"completions/mean_length": 804.75,
"completions/min_length": 605.0,
"completions/max_length": 1214.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 804.75,
"completions/min_terminated_length": 605.0,
"completions/max_terminated_length": 1214.0,
"rewards/match_format_exactly/mean": 1.5,
"rewards/match_format_exactly/std": 1.7320507764816284,
"rewards/match_format_approximately/mean": -0.75,
"rewards/match_format_approximately/std": 2.598076105117798,
"rewards/check_answer/mean": -1.25,
"rewards/check_answer/std": 1.8484227657318115,
"rewards/check_numbers/mean": -2.0,
"rewards/check_numbers/std": 0.5773502588272095,
"reward": -2.5,
"reward_std": 6.0,
"frac_reward_zero_std": 0.0,
"completion_length": 804.75,
"kl": 0.008573448285460472,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.00023605319065229366,
"time_ms": 40298.14357904252,
"memory_mb": 9952.80322265625,
"memory_gb": 9.719534397125244
},
{
"step": 4,
"loss": -0.1236,
"grad_norm": 0.5647851228713989,
"learning_rate": 3.88888888888889e-06,
"num_tokens": 13320.0,
"completions/mean_length": 623.75,
"completions/min_length": 421.0,
"completions/max_length": 789.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 623.75,
"completions/min_terminated_length": 421.0,
"completions/max_terminated_length": 789.0,
"rewards/match_format_exactly/mean": 0.75,
"rewards/match_format_exactly/std": 1.5,
"rewards/match_format_approximately/mean": 0.375,
"rewards/match_format_approximately/std": 0.75,
"rewards/check_answer/mean": -2.125,
"rewards/check_answer/std": 0.25,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": -2.5,
"reward_std": 2.0,
"frac_reward_zero_std": 0.0,
"completion_length": 623.75,
"kl": 0.009312103502452374,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.00031473758753639155,
"time_ms": 26193.765547999647,
"memory_mb": 9306.09326171875,
"memory_gb": 9.087981700897217
},
{
"step": 5,
"loss": 0.0,
"grad_norm": 0.0010538548231124878,
"learning_rate": 3.3333333333333333e-06,
"num_tokens": 14970.0,
"completions/mean_length": 256.5,
"completions/min_length": 246.0,
"completions/max_length": 260.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 256.5,
"completions/min_terminated_length": 246.0,
"completions/max_terminated_length": 260.0,
"rewards/match_format_exactly/mean": 3.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": 1.5,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -2.5,
"rewards/check_answer/std": 0.0,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": 0.5,
"reward_std": 0.0,
"frac_reward_zero_std": 1.0,
"completion_length": 256.5,
"kl": 0.002130241831764579,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.00039342198442048943,
"time_ms": 9324.433026020415,
"memory_mb": 8777.982421875,
"memory_gb": 8.572248458862305
},
{
"step": 6,
"loss": 0.0,
"grad_norm": 0.00012166703527327627,
"learning_rate": 2.7777777777777783e-06,
"num_tokens": 21841.0,
"completions/mean_length": 1620.75,
"completions/min_length": 1214.0,
"completions/max_length": 1846.0,
"completions/clipped_ratio": 0.5,
"completions/mean_terminated_length": 1395.5,
"completions/min_terminated_length": 1214.0,
"completions/max_terminated_length": 1577.0,
"rewards/match_format_exactly/mean": 0.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": -3.0,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -2.0,
"rewards/check_answer/std": 0.0,
"rewards/check_numbers/mean": -2.5,
"rewards/check_numbers/std": 0.0,
"reward": -7.5,
"reward_std": 0.0,
"frac_reward_zero_std": 1.0,
"completion_length": 1620.75,
"kl": 0.0007094849133864045,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0004721063813045873,
"time_ms": 61652.94101298787,
"memory_mb": 10914.11279296875,
"memory_gb": 10.658313274383545
},
{
"step": 7,
"loss": -0.0097,
"grad_norm": 0.9194015860557556,
"learning_rate": 2.222222222222222e-06,
"num_tokens": 24587.0,
"completions/mean_length": 543.5,
"completions/min_length": 511.0,
"completions/max_length": 562.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 543.5,
"completions/min_terminated_length": 511.0,
"completions/max_terminated_length": 562.0,
"rewards/match_format_exactly/mean": 0.75,
"rewards/match_format_exactly/std": 1.5,
"rewards/match_format_approximately/mean": -1.875,
"rewards/match_format_approximately/std": 2.25,
"rewards/check_answer/mean": -2.125,
"rewards/check_answer/std": 0.25,
"rewards/check_numbers/mean": -2.25,
"rewards/check_numbers/std": 0.5,
"reward": -5.5,
"reward_std": 4.0,
"frac_reward_zero_std": 0.0,
"completion_length": 543.5,
"kl": 0.005215016193687916,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0005507907781886852,
"time_ms": 18906.7719859886,
"memory_mb": 8996.00390625,
"memory_gb": 8.785160064697266
},
{
"step": 8,
"loss": 0.0338,
"grad_norm": 0.5346357822418213,
"learning_rate": 1.6666666666666667e-06,
"num_tokens": 27519.0,
"completions/mean_length": 666.0,
"completions/min_length": 615.0,
"completions/max_length": 714.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 666.0,
"completions/min_terminated_length": 615.0,
"completions/max_terminated_length": 714.0,
"rewards/match_format_exactly/mean": 0.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": -2.25,
"rewards/match_format_approximately/std": 1.5,
"rewards/check_answer/mean": -2.0,
"rewards/check_answer/std": 0.0,
"rewards/check_numbers/mean": -1.0,
"rewards/check_numbers/std": 3.0,
"reward": -5.25,
"reward_std": 4.5,
"frac_reward_zero_std": 0.0,
"completion_length": 666.0,
"kl": 0.0001442090724594891,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0006294751750727831,
"time_ms": 23948.13355000224,
"memory_mb": 9202.39306640625,
"memory_gb": 8.986711978912354
},
{
"step": 9,
"loss": 0.0,
"grad_norm": 0.0033828848972916603,
"learning_rate": 1.111111111111111e-06,
"num_tokens": 29207.0,
"completions/mean_length": 325.0,
"completions/min_length": 310.0,
"completions/max_length": 334.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 325.0,
"completions/min_terminated_length": 310.0,
"completions/max_terminated_length": 334.0,
"rewards/match_format_exactly/mean": 3.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": 1.5,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -2.5,
"rewards/check_answer/std": 0.0,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": 0.5,
"reward_std": 0.0,
"frac_reward_zero_std": 1.0,
"completion_length": 325.0,
"kl": 0.010901343077421188,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0007081595719568809,
"time_ms": 11388.401818985585,
"memory_mb": 8705.7236328125,
"memory_gb": 8.501683235168457
},
{
"step": 10,
"loss": 0.2036,
"grad_norm": 0.22472381591796875,
"learning_rate": 5.555555555555555e-07,
"num_tokens": 34891.0,
"completions/mean_length": 1253.0,
"completions/min_length": 1044.0,
"completions/max_length": 1846.0,
"completions/clipped_ratio": 0.25,
"completions/mean_terminated_length": 1055.3333740234375,
"completions/min_terminated_length": 1044.0,
"completions/max_terminated_length": 1067.0,
"rewards/match_format_exactly/mean": 1.5,
"rewards/match_format_exactly/std": 1.7320507764816284,
"rewards/match_format_approximately/mean": 0.0,
"rewards/match_format_approximately/std": 2.1213202476501465,
"rewards/check_answer/mean": -2.25,
"rewards/check_answer/std": 0.28867512941360474,
"rewards/check_numbers/mean": -1.75,
"rewards/check_numbers/std": 0.5,
"reward": -2.5,
"reward_std": 3.8297085762023926,
"frac_reward_zero_std": 0.0,
"completion_length": 1253.0,
"kl": 0.00215436820872128,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0007868439688409789,
"time_ms": 61737.33338096645,
"memory_mb": 10919.5498046875,
"memory_gb": 10.663622856140137
}
]

View file

@ -1,75 +0,0 @@
{
"backend": "unsloth_fi_false",
"max_steps": 10,
"train_wall_s": 355.41431403998286,
"median_step_ms_post_warmup": 23948.13355000224,
"n_logged_steps": 10,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
},
"losses": [
0.0,
-0.0893,
-0.1386,
-0.1236,
0.0,
0.0,
-0.0097,
0.0338,
0.0,
0.2036
],
"rewards": [
0.5,
-6.5,
-2.5,
-2.5,
0.5,
-7.5,
-5.5,
-5.25,
0.5,
-2.5
],
"kls": [
0.0,
0.0,
0.008573448285460472,
0.009312103502452374,
0.002130241831764579,
0.0007094849133864045,
0.005215016193687916,
0.0001442090724594891,
0.010901343077421188,
0.00215436820872128
],
"grad_norms": [
0.0,
0.6125104427337646,
0.5993297696113586,
0.5647851228713989,
0.0010538548231124878,
0.00012166703527327627,
0.9194015860557556,
0.5346357822418213,
0.0033828848972916603,
0.22472381591796875
],
"step_times_ms": [
77897.29859499494,
21634.983669035137,
40298.14357904252,
26193.765547999647,
9324.433026020415,
61652.94101298787,
18906.7719859886,
23948.13355000224,
11388.401818985585,
61737.33338096645
],
"peak_memory_gb": 10.663622856140137,
"logs_path": "logs/grpo_unsloth_fi_false_10.json"
}

View file

@ -1,362 +0,0 @@
[
{
"step": 1,
"loss": 0.0305,
"grad_norm": 0.4133029878139496,
"learning_rate": 0.0,
"num_tokens": 3705.0,
"completions/mean_length": 814.25,
"completions/min_length": 781.0,
"completions/max_length": 864.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 814.25,
"completions/min_terminated_length": 781.0,
"completions/max_terminated_length": 864.0,
"rewards/match_format_exactly/mean": 3.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": 1.5,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -3.0,
"rewards/check_answer/std": 1.0,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": 0.0,
"reward_std": 1.0,
"frac_reward_zero_std": 0.0,
"completion_length": 814.25,
"kl": 0.0,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 7.868439688409789e-05,
"time_ms": 17984.56621397054,
"memory_mb": 161170.2099609375,
"memory_gb": 157.39278316497803
},
{
"step": 2,
"loss": -0.1941,
"grad_norm": 0.8333088159561157,
"learning_rate": 5e-06,
"num_tokens": 7167.0,
"completions/mean_length": 776.5,
"completions/min_length": 525.0,
"completions/max_length": 1078.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 776.5,
"completions/min_terminated_length": 525.0,
"completions/max_terminated_length": 1078.0,
"rewards/match_format_exactly/mean": 0.75,
"rewards/match_format_exactly/std": 1.5,
"rewards/match_format_approximately/mean": 0.375,
"rewards/match_format_approximately/std": 0.75,
"rewards/check_answer/mean": -2.125,
"rewards/check_answer/std": 0.25,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": -2.5,
"reward_std": 2.0,
"frac_reward_zero_std": 0.0,
"completion_length": 776.5,
"kl": 0.0,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.00015736879376819577,
"time_ms": 6704.717919987161,
"memory_mb": 161638.7705078125,
"memory_gb": 157.85036182403564
},
{
"step": 3,
"loss": 0.2632,
"grad_norm": 0.4677680730819702,
"learning_rate": 4.444444444444444e-06,
"num_tokens": 11596.0,
"completions/mean_length": 930.25,
"completions/min_length": 445.0,
"completions/max_length": 1846.0,
"completions/clipped_ratio": 0.25,
"completions/mean_terminated_length": 625.0,
"completions/min_terminated_length": 445.0,
"completions/max_terminated_length": 863.0,
"rewards/match_format_exactly/mean": 1.5,
"rewards/match_format_exactly/std": 1.7320507764816284,
"rewards/match_format_approximately/mean": -0.75,
"rewards/match_format_approximately/std": 2.598076105117798,
"rewards/check_answer/mean": -2.75,
"rewards/check_answer/std": 1.190238118171692,
"rewards/check_numbers/mean": -1.625,
"rewards/check_numbers/std": 1.1814539432525635,
"reward": -3.625,
"reward_std": 4.479118347167969,
"frac_reward_zero_std": 0.0,
"completion_length": 930.25,
"kl": 0.011923530139029026,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.00023605319065229366,
"time_ms": 12108.133931003977,
"memory_mb": 162822.73876953125,
"memory_gb": 159.00658082962036
},
{
"step": 4,
"loss": -0.2013,
"grad_norm": 0.5163940191268921,
"learning_rate": 3.88888888888889e-06,
"num_tokens": 14365.0,
"completions/mean_length": 527.25,
"completions/min_length": 315.0,
"completions/max_length": 598.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 527.25,
"completions/min_terminated_length": 315.0,
"completions/max_terminated_length": 598.0,
"rewards/match_format_exactly/mean": 3.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": 1.5,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -3.0,
"rewards/check_answer/std": 1.0,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": 0.0,
"reward_std": 1.0,
"frac_reward_zero_std": 0.0,
"completion_length": 527.25,
"kl": 0.004221913404762745,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.00031473758753639155,
"time_ms": 4019.3081409670413,
"memory_mb": 160943.57275390625,
"memory_gb": 157.17145776748657
},
{
"step": 5,
"loss": 0.2093,
"grad_norm": 1.160618782043457,
"learning_rate": 3.3333333333333333e-06,
"num_tokens": 16176.0,
"completions/mean_length": 296.75,
"completions/min_length": 246.0,
"completions/max_length": 421.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 296.75,
"completions/min_terminated_length": 246.0,
"completions/max_terminated_length": 421.0,
"rewards/match_format_exactly/mean": 3.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": 1.5,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -3.0,
"rewards/check_answer/std": 1.0,
"rewards/check_numbers/mean": -1.125,
"rewards/check_numbers/std": 0.75,
"reward": 0.375,
"reward_std": 0.25,
"frac_reward_zero_std": 0.0,
"completion_length": 296.75,
"kl": 0.003692739875987172,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.00039342198442048943,
"time_ms": 3241.851194994524,
"memory_mb": 160665.36376953125,
"memory_gb": 156.89976930618286
},
{
"step": 6,
"loss": 0.0,
"grad_norm": 0.0007444396032951772,
"learning_rate": 2.7777777777777783e-06,
"num_tokens": 23948.0,
"completions/mean_length": 1846.0,
"completions/min_length": 1846.0,
"completions/max_length": 1846.0,
"completions/clipped_ratio": 1.0,
"completions/mean_terminated_length": 0.0,
"completions/min_terminated_length": 0.0,
"completions/max_terminated_length": 0.0,
"rewards/match_format_exactly/mean": 0.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": -3.0,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -2.0,
"rewards/check_answer/std": 0.0,
"rewards/check_numbers/mean": -2.5,
"rewards/check_numbers/std": 0.0,
"reward": -7.5,
"reward_std": 0.0,
"frac_reward_zero_std": 1.0,
"completion_length": 1846.0,
"kl": 0.0025038770399987698,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0004721063813045873,
"time_ms": 10860.668059962336,
"memory_mb": 162817.607421875,
"memory_gb": 159.0015697479248
},
{
"step": 7,
"loss": 0.0371,
"grad_norm": 2.262518882751465,
"learning_rate": 2.222222222222222e-06,
"num_tokens": 26728.0,
"completions/mean_length": 552.0,
"completions/min_length": 506.0,
"completions/max_length": 627.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 552.0,
"completions/min_terminated_length": 506.0,
"completions/max_terminated_length": 627.0,
"rewards/match_format_exactly/mean": 3.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": 1.5,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -1.5,
"rewards/check_answer/std": 2.0,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": 1.5,
"reward_std": 2.0,
"frac_reward_zero_std": 0.0,
"completion_length": 552.0,
"kl": 0.006138760130852461,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0005507907781886852,
"time_ms": 4138.088690000586,
"memory_mb": 160989.1611328125,
"memory_gb": 157.2159776687622
},
{
"step": 8,
"loss": 0.0,
"grad_norm": 0.001562082557938993,
"learning_rate": 1.6666666666666667e-06,
"num_tokens": 29420.0,
"completions/mean_length": 606.0,
"completions/min_length": 560.0,
"completions/max_length": 636.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 606.0,
"completions/min_terminated_length": 560.0,
"completions/max_terminated_length": 636.0,
"rewards/match_format_exactly/mean": 0.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": -3.0,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -2.0,
"rewards/check_answer/std": 0.0,
"rewards/check_numbers/mean": -2.5,
"rewards/check_numbers/std": 0.0,
"reward": -7.5,
"reward_std": 0.0,
"frac_reward_zero_std": 1.0,
"completion_length": 606.0,
"kl": 0.004652692936360836,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0006294751750727831,
"time_ms": 4177.52773797838,
"memory_mb": 160996.86376953125,
"memory_gb": 157.22349977493286
},
{
"step": 9,
"loss": 0.0,
"grad_norm": 0.00027447607135400176,
"learning_rate": 1.111111111111111e-06,
"num_tokens": 31353.0,
"completions/mean_length": 386.25,
"completions/min_length": 334.0,
"completions/max_length": 464.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 386.25,
"completions/min_terminated_length": 334.0,
"completions/max_terminated_length": 464.0,
"rewards/match_format_exactly/mean": 3.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": 1.5,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -2.5,
"rewards/check_answer/std": 0.0,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": 0.5,
"reward_std": 0.0,
"frac_reward_zero_std": 1.0,
"completion_length": 386.25,
"kl": 0.0017617446137592196,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0007081595719568809,
"time_ms": 3263.2006779895164,
"memory_mb": 160745.5380859375,
"memory_gb": 156.97806453704834
},
{
"step": 10,
"loss": 0.2052,
"grad_norm": 0.4320540428161621,
"learning_rate": 5.555555555555555e-07,
"num_tokens": 35257.0,
"completions/mean_length": 808.0,
"completions/min_length": 650.0,
"completions/max_length": 1119.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 808.0,
"completions/min_terminated_length": 650.0,
"completions/max_terminated_length": 1119.0,
"rewards/match_format_exactly/mean": 1.5,
"rewards/match_format_exactly/std": 1.7320507764816284,
"rewards/match_format_approximately/mean": 0.0,
"rewards/match_format_approximately/std": 2.1213202476501465,
"rewards/check_answer/mean": -3.25,
"rewards/check_answer/std": 1.4433757066726685,
"rewards/check_numbers/mean": -1.75,
"rewards/check_numbers/std": 0.5,
"reward": -3.5,
"reward_std": 2.8284270763397217,
"frac_reward_zero_std": 0.0,
"completion_length": 808.0,
"kl": 0.001998987514525652,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0007868439688409789,
"time_ms": 6874.866564990953,
"memory_mb": 161736.77734375,
"memory_gb": 157.94607162475586
}
]

View file

@ -1,75 +0,0 @@
{
"backend": "vllm",
"max_steps": 10,
"train_wall_s": 74.41919421299826,
"median_step_ms_post_warmup": 4138.088690000586,
"n_logged_steps": 10,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
},
"losses": [
0.0305,
-0.1941,
0.2632,
-0.2013,
0.2093,
0.0,
0.0371,
0.0,
0.0,
0.2052
],
"rewards": [
0.0,
-2.5,
-3.625,
0.0,
0.375,
-7.5,
1.5,
-7.5,
0.5,
-3.5
],
"kls": [
0.0,
0.0,
0.011923530139029026,
0.004221913404762745,
0.003692739875987172,
0.0025038770399987698,
0.006138760130852461,
0.004652692936360836,
0.0017617446137592196,
0.001998987514525652
],
"grad_norms": [
0.4133029878139496,
0.8333088159561157,
0.4677680730819702,
0.5163940191268921,
1.160618782043457,
0.0007444396032951772,
2.262518882751465,
0.001562082557938993,
0.00027447607135400176,
0.4320540428161621
],
"step_times_ms": [
17984.56621397054,
6704.717919987161,
12108.133931003977,
4019.3081409670413,
3241.851194994524,
10860.668059962336,
4138.088690000586,
4177.52773797838,
3263.2006779895164,
6874.866564990953
],
"peak_memory_gb": 157.94607162475586,
"logs_path": "logs/grpo_vllm_10.json"
}

File diff suppressed because it is too large Load diff

View file

@ -1,175 +0,0 @@
{
"backend": "vllm",
"max_steps": 30,
"train_wall_s": 215.93619061401114,
"median_step_ms_post_warmup": 5140.9110790118575,
"n_logged_steps": 30,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
},
"losses": [
0.0305,
-0.1941,
0.2006,
0.2437,
0.0,
0.0,
0.0,
0.2403,
0.0251,
0.0,
0.0,
0.331,
0.0298,
0.1561,
0.4028,
0.0184,
0.1807,
0.2383,
0.8338,
0.0,
-0.0037,
0.0379,
0.3886,
0.1431,
0.1063,
-0.0821,
0.0,
0.0,
0.0181,
0.2357
],
"rewards": [
0.0,
-2.5,
0.0,
-6.5,
0.5,
-7.5,
-7.5,
-5.5,
-3.5,
-7.5,
-7.5,
-5.5,
-6.5,
2.875,
4.75,
0.625,
2.75,
-3.5,
-3.375,
-3.5,
6.0,
8.25,
-0.5,
-4.5,
3.625,
0.25,
-7.5,
-7.5,
-1.0,
-0.75
],
"kls": [
0.0,
0.0,
0.0024282929953187704,
0.005713047459721565,
0.011252232827246189,
0.0008392990566790104,
0.0047083343379199505,
0.004438905976712704,
0.002922436688095331,
0.001744209323078394,
0.002685483079403639,
0.007562238723039627,
0.01840771734714508,
0.00663342559710145,
0.004423078149557114,
0.004664436914026737,
0.002878781408071518,
0.00984956230968237,
0.0006688942667096853,
0.013835551217198372,
0.0076245637610554695,
0.011878136545419693,
0.010346058756113052,
0.009395054541528225,
0.0016164245316758752,
0.005659917835146189,
7.657324022147804e-05,
0.001502353698015213,
0.009886534884572029,
0.0031997335609048605
],
"grad_norms": [
0.41349539160728455,
0.8339279294013977,
0.6402159929275513,
0.2846885919570923,
0.00401803245767951,
0.00014282428310252726,
0.0015061397571116686,
0.44619685411453247,
0.7223323583602905,
0.0002955764648504555,
0.0008108518086373806,
0.9407532215118408,
0.6642693281173706,
0.5970175266265869,
0.2463085651397705,
0.4361814856529236,
0.25606873631477356,
0.38782942295074463,
0.2885834872722626,
0.0024749308358877897,
0.6102232336997986,
0.8350751996040344,
0.28949517011642456,
0.5029579401016235,
0.3104912340641022,
0.4499339461326599,
7.777348946547136e-05,
0.00013634964125230908,
0.419629842042923,
0.22457966208457947
],
"step_times_ms": [
17866.63037497783,
6304.458727012388,
4510.530841012951,
11239.92098000599,
2776.7173860338517,
10157.209870987572,
4554.909924976528,
10537.410682998598,
2670.8698750007898,
6595.109536021482,
10160.135700018145,
3498.259258980397,
3411.464748030994,
3980.731577030383,
10481.332287017722,
5105.93488701852,
10356.32998100482,
10494.073983980343,
10833.778033033013,
3094.477139005903,
3840.9026580047794,
2903.038158954587,
10495.372234028764,
10606.634334020782,
3563.0593819660135,
4147.492960037198,
5140.9110790118575,
10219.65475397883,
4818.701309966855,
10382.734156039078
],
"peak_memory_gb": 159.00153064727783,
"logs_path": "logs/grpo_vllm_30.json"
}

View file

@ -1,29 +0,0 @@
{
"backend": "tpaged",
"lora_adapter": "outputs/lora_rank32_fresh",
"attn_impl": "paged_attention",
"persistent_cb": true,
"n_prompts": 32,
"n_prompt_tokens": 4847,
"n_decoded_tokens": 14750,
"wall_times_s": [
34.991439365025144,
33.211731680028606
],
"median_wall_s": 34.991439365025144,
"prompt_tps": 138.5195947339252,
"decode_tps": 421.53167367967745,
"max_new_tokens": 512,
"sample_completions": [
"First, we can factor the quadratic expression $n^2-3n+2$ as $(n-1)(n-2)$. For this expression to be a prime number, one of the factors must be equal to 1 and the other factor must be a prime number. \n",
"First, we need to determine how many $4 \\times 5$ rectangles can fit into a $20 \\times 24$ rectangle. To do this, we divide the dimensions of the larger rectangle by the dimensions of the smaller rect",
" \nTo find the area of the smaller square, we need to determine its side length. Let's denote the side length of the smaller square as \\( s \\).\n\nFrom the diagram, we can see that the larger square has "
],
"peak_memory_gb": 103.81339406967163,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
}
}

View file

@ -1,29 +0,0 @@
{
"backend": "tpaged",
"lora_adapter": "outputs/lora_rank32_fresh",
"attn_impl": "sdpa_paged",
"persistent_cb": true,
"n_prompts": 32,
"n_prompt_tokens": 4847,
"n_decoded_tokens": 14785,
"wall_times_s": [
33.532237556006294,
34.068173120962456
],
"median_wall_s": 34.068173120962456,
"prompt_tps": 142.27355199793783,
"decode_tps": 433.9827658942667,
"max_new_tokens": 512,
"sample_completions": [
"First, we need to understand the structure of a cube. A cube has 12 edges and 8 vertices. Each vertex is connected to 3 edges. \n\nNow, let's consider the pairs of parallel edges. Since a cube has 12 ed",
"First, we need to determine how many $4 \\times 5$ rectangles can fit into a $20 \\times 24$ rectangle. We can do this by dividing the dimensions of the larger rectangle by the dimensions of the smaller",
"First, let's count the total number of letters in the word \"FLUFFY\". There are 6 letters in total.\n\nNext, we need to determine how many of these letters are repeated. In this case, the letter \"F\" appe"
],
"peak_memory_gb": 111.93839406967163,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
}
}

View file

@ -1,27 +0,0 @@
{
"backend": "unsloth_fi_false",
"lora_adapter": "outputs/lora_rank32_fresh",
"n_prompts": 32,
"n_prompt_tokens": 4847,
"n_decoded_tokens": 16384,
"wall_times_s": [
25.54396249598358,
25.480999241000973
],
"median_wall_s": 25.54396249598358,
"prompt_tps": 189.75129644674828,
"decode_tps": 641.4040109311994,
"max_new_tokens": 512,
"sample_completions": [
"First, we need to find the length of the legs of the trapezoid. Since the trapezoid is isosceles, the legs are equal in length. Let's call the length of each leg $x$. We can use the Pythagorean theore",
"Let $Q(x) = P(x) - x^{2023}P(1-\\frac{1}{x})$. Then $Q(k) = 0$ for every positive integer $1 \\leq k \\leq 2023$. Since $P(x)$ is a monic polynomial of degree $2023$, $Q(x)$ is also a monic polynomial of",
" To solve this problem, we need to determine the maximum value of \\(a\\) such that the line \\(y = mx + 2\\) does not pass through any lattice points for \\(0 < x \\leq 100\\) when \\(\\frac{1}{2} < m < a\\).\n"
],
"peak_memory_gb": 15.8363037109375,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
}
}

View file

@ -1,27 +0,0 @@
{
"backend": "vllm",
"lora_adapter": "outputs/lora_rank32_fresh",
"n_prompts": 32,
"n_prompt_tokens": 4847,
"n_decoded_tokens": 15140,
"wall_times_s": [
3.304712440993171,
3.2573561430326663
],
"median_wall_s": 3.304712440993171,
"prompt_tps": 1466.6934223612275,
"decode_tps": 4581.336582329066,
"max_new_tokens": 512,
"sample_completions": [
"First, we need to find the length of the legs of the trapezoid. Since the trapezoid is isosceles, the legs are equal in length. Let's call the length of each leg $x$. We can use the Pythagorean theore",
"Let $Q(x) = P(x) - x^{2023}P(1-\\frac{1}{x})$. Then $Q(k) = 0$ for every positive integer $1 \\leq k \\leq 2023$. Since $P(x)$ is a monic polynomial of degree $2023$, $Q(x)$ is also a monic polynomial of",
" To solve this problem, we need to find the maximum value of \\(a\\) such that the line \\(y = mx + 2\\) does not pass through any lattice points for \\(0 < x \\leq 100\\) when \\(\\frac{1}{2} < m < a\\).\n\nFirs"
],
"peak_memory_gb": 156.21798133850098,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
}
}

View file

@ -1,362 +0,0 @@
[
{
"step": 1,
"loss": 0.2423,
"grad_norm": 0.24541568756103516,
"learning_rate": 0.0,
"num_tokens": 5422.0,
"completions/mean_length": 1243.5,
"completions/min_length": 1019.0,
"completions/max_length": 1846.0,
"completions/clipped_ratio": 0.25,
"completions/mean_terminated_length": 1042.666748046875,
"completions/min_terminated_length": 1019.0,
"completions/max_terminated_length": 1089.0,
"rewards/match_format_exactly/mean": 2.25,
"rewards/match_format_exactly/std": 1.5,
"rewards/match_format_approximately/mean": 0.75,
"rewards/match_format_approximately/std": 1.5,
"rewards/check_answer/mean": -2.375,
"rewards/check_answer/std": 0.25,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": -0.875,
"reward_std": 2.75,
"frac_reward_zero_std": 0.0,
"completion_length": 1243.5,
"kl": 0.0,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 7.868439688409789e-05,
"time_ms": 60626.65366800502,
"memory_mb": 162696.66357421875,
"memory_gb": 158.883460521698
},
{
"step": 2,
"loss": 0.1559,
"grad_norm": 0.7674608826637268,
"learning_rate": 5e-06,
"num_tokens": 7690.0,
"completions/mean_length": 478.0,
"completions/min_length": 329.0,
"completions/max_length": 553.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 478.0,
"completions/min_terminated_length": 329.0,
"completions/max_terminated_length": 553.0,
"rewards/match_format_exactly/mean": 0.75,
"rewards/match_format_exactly/std": 1.5,
"rewards/match_format_approximately/mean": -1.875,
"rewards/match_format_approximately/std": 2.25,
"rewards/check_answer/mean": -2.125,
"rewards/check_answer/std": 0.25,
"rewards/check_numbers/mean": -2.25,
"rewards/check_numbers/std": 0.5,
"reward": -5.5,
"reward_std": 4.0,
"frac_reward_zero_std": 0.0,
"completion_length": 478.0,
"kl": 0.0,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.00015736879376819577,
"time_ms": 3893.7715340289287,
"memory_mb": 160786.34130859375,
"memory_gb": 157.01791143417358
},
{
"step": 3,
"loss": -0.165,
"grad_norm": 0.47987309098243713,
"learning_rate": 4.444444444444444e-06,
"num_tokens": 12437.0,
"completions/mean_length": 1009.75,
"completions/min_length": 745.0,
"completions/max_length": 1319.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 1009.75,
"completions/min_terminated_length": 745.0,
"completions/max_terminated_length": 1319.0,
"rewards/match_format_exactly/mean": 2.25,
"rewards/match_format_exactly/std": 1.5,
"rewards/match_format_approximately/mean": 0.375,
"rewards/match_format_approximately/std": 2.25,
"rewards/check_answer/mean": -1.375,
"rewards/check_answer/std": 1.9311050176620483,
"rewards/check_numbers/mean": -1.75,
"rewards/check_numbers/std": 0.5,
"reward": -0.5,
"reward_std": 5.0332231521606445,
"frac_reward_zero_std": 0.0,
"completion_length": 1009.75,
"kl": 0.0038291513919830322,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.00023605319065229366,
"time_ms": 7929.127738985699,
"memory_mb": 161959.54833984375,
"memory_gb": 158.16362142562866
},
{
"step": 4,
"loss": 0.3177,
"grad_norm": 0.37925368547439575,
"learning_rate": 3.88888888888889e-06,
"num_tokens": 17403.0,
"completions/mean_length": 1076.5,
"completions/min_length": 546.0,
"completions/max_length": 1846.0,
"completions/clipped_ratio": 0.25,
"completions/mean_terminated_length": 820.0,
"completions/min_terminated_length": 546.0,
"completions/max_terminated_length": 1192.0,
"rewards/match_format_exactly/mean": 0.75,
"rewards/match_format_exactly/std": 1.5,
"rewards/match_format_approximately/mean": -0.75,
"rewards/match_format_approximately/std": 1.9364917278289795,
"rewards/check_answer/mean": -2.125,
"rewards/check_answer/std": 0.25,
"rewards/check_numbers/mean": -2.0,
"rewards/check_numbers/std": 0.5773502588272095,
"reward": -4.125,
"reward_std": 3.4970226287841797,
"frac_reward_zero_std": 0.0,
"completion_length": 1076.5,
"kl": 0.005913741886615753,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.00031473758753639155,
"time_ms": 11084.477900003549,
"memory_mb": 162749.5458984375,
"memory_gb": 158.93510341644287
},
{
"step": 5,
"loss": -0.02,
"grad_norm": 0.6089861989021301,
"learning_rate": 3.3333333333333333e-06,
"num_tokens": 19235.0,
"completions/mean_length": 302.0,
"completions/min_length": 256.0,
"completions/max_length": 334.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 302.0,
"completions/min_terminated_length": 256.0,
"completions/max_terminated_length": 334.0,
"rewards/match_format_exactly/mean": 3.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": 1.5,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -1.125,
"rewards/check_answer/std": 4.190763473510742,
"rewards/check_numbers/mean": -0.25,
"rewards/check_numbers/std": 2.5,
"reward": 3.125,
"reward_std": 6.650501251220703,
"frac_reward_zero_std": 0.0,
"completion_length": 302.0,
"kl": 0.015983864665031433,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.00039342198442048943,
"time_ms": 2679.084858042188,
"memory_mb": 160465.05859375,
"memory_gb": 156.70415878295898
},
{
"step": 6,
"loss": 0.0,
"grad_norm": 0.002889552852138877,
"learning_rate": 2.7777777777777783e-06,
"num_tokens": 26958.0,
"completions/mean_length": 1833.75,
"completions/min_length": 1797.0,
"completions/max_length": 1846.0,
"completions/clipped_ratio": 0.75,
"completions/mean_terminated_length": 1797.0,
"completions/min_terminated_length": 1797.0,
"completions/max_terminated_length": 1797.0,
"rewards/match_format_exactly/mean": 0.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": -3.0,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -2.0,
"rewards/check_answer/std": 0.0,
"rewards/check_numbers/mean": -2.5,
"rewards/check_numbers/std": 0.0,
"reward": -7.5,
"reward_std": 0.0,
"frac_reward_zero_std": 1.0,
"completion_length": 1833.75,
"kl": 0.003961368463933468,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0004721063813045873,
"time_ms": 10652.5042289868,
"memory_mb": 162744.833984375,
"memory_gb": 158.9305019378662
},
{
"step": 7,
"loss": 0.0,
"grad_norm": 0.003148352960124612,
"learning_rate": 2.222222222222222e-06,
"num_tokens": 29616.0,
"completions/mean_length": 521.5,
"completions/min_length": 452.0,
"completions/max_length": 675.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 521.5,
"completions/min_terminated_length": 452.0,
"completions/max_terminated_length": 675.0,
"rewards/match_format_exactly/mean": 0.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": -3.0,
"rewards/match_format_approximately/std": 0.0,
"rewards/check_answer/mean": -2.0,
"rewards/check_answer/std": 0.0,
"rewards/check_numbers/mean": -2.5,
"rewards/check_numbers/std": 0.0,
"reward": -7.5,
"reward_std": 0.0,
"frac_reward_zero_std": 1.0,
"completion_length": 521.5,
"kl": 0.009647021070122719,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0005507907781886852,
"time_ms": 4468.553012993652,
"memory_mb": 160983.70361328125,
"memory_gb": 157.21064805984497
},
{
"step": 8,
"loss": 0.0613,
"grad_norm": 0.4861072301864624,
"learning_rate": 1.6666666666666667e-06,
"num_tokens": 32984.0,
"completions/mean_length": 775.0,
"completions/min_length": 671.0,
"completions/max_length": 933.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 775.0,
"completions/min_terminated_length": 671.0,
"completions/max_terminated_length": 933.0,
"rewards/match_format_exactly/mean": 0.0,
"rewards/match_format_exactly/std": 0.0,
"rewards/match_format_approximately/mean": -2.25,
"rewards/match_format_approximately/std": 1.5,
"rewards/check_answer/mean": -2.0,
"rewards/check_answer/std": 0.0,
"rewards/check_numbers/mean": -2.25,
"rewards/check_numbers/std": 0.5,
"reward": -6.5,
"reward_std": 2.0,
"frac_reward_zero_std": 0.0,
"completion_length": 775.0,
"kl": 0.003189136739820242,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0006294751750727831,
"time_ms": 5805.1959190052,
"memory_mb": 161361.3447265625,
"memory_gb": 157.5794382095337
},
{
"step": 9,
"loss": 0.006,
"grad_norm": 0.5726504921913147,
"learning_rate": 1.111111111111111e-06,
"num_tokens": 35116.0,
"completions/mean_length": 436.0,
"completions/min_length": 237.0,
"completions/max_length": 635.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 436.0,
"completions/min_terminated_length": 237.0,
"completions/max_terminated_length": 635.0,
"rewards/match_format_exactly/mean": 1.5,
"rewards/match_format_exactly/std": 1.7320507764816284,
"rewards/match_format_approximately/mean": 0.75,
"rewards/match_format_approximately/std": 0.8660253882408142,
"rewards/check_answer/mean": -1.25,
"rewards/check_answer/std": 1.8484227657318115,
"rewards/check_numbers/mean": -1.5,
"rewards/check_numbers/std": 0.0,
"reward": -0.5,
"reward_std": 3.8297085762023926,
"frac_reward_zero_std": 0.0,
"completion_length": 436.0,
"kl": 0.0023976736702024937,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0007081595719568809,
"time_ms": 4304.089896031655,
"memory_mb": 160920.1044921875,
"memory_gb": 157.14853954315186
},
{
"step": 10,
"loss": 0.1582,
"grad_norm": 0.3651980459690094,
"learning_rate": 5.555555555555555e-07,
"num_tokens": 39155.0,
"completions/mean_length": 841.75,
"completions/min_length": 729.0,
"completions/max_length": 1108.0,
"completions/clipped_ratio": 0.0,
"completions/mean_terminated_length": 841.75,
"completions/min_terminated_length": 729.0,
"completions/max_terminated_length": 1108.0,
"rewards/match_format_exactly/mean": 2.25,
"rewards/match_format_exactly/std": 1.5,
"rewards/match_format_approximately/mean": 0.375,
"rewards/match_format_approximately/std": 2.25,
"rewards/check_answer/mean": -2.375,
"rewards/check_answer/std": 0.25,
"rewards/check_numbers/mean": -1.75,
"rewards/check_numbers/std": 0.5,
"reward": -1.5,
"reward_std": 4.0,
"frac_reward_zero_std": 0.0,
"completion_length": 841.75,
"kl": 0.00485160993412137,
"clip_ratio/low_mean": 0.0,
"clip_ratio/low_min": 0.0,
"clip_ratio/high_mean": 0.0,
"clip_ratio/high_max": 0.0,
"clip_ratio/region_mean": 0.0,
"epoch": 0.0007868439688409789,
"time_ms": 6777.490795007907,
"memory_mb": 161641.44970703125,
"memory_gb": 157.8529782295227
}
]

View file

@ -1,28 +0,0 @@
{
"backend": "vllm",
"lora_adapter": null,
"n_prompts": 128,
"n_prompt_tokens": 18551,
"n_decoded_tokens": 60123,
"wall_times_s": [
4.089218033012003,
4.009195051970892,
3.9967994149774313
],
"median_wall_s": 4.009195051970892,
"prompt_tps": 4627.1133630379145,
"decode_tps": 14996.277113143688,
"max_new_tokens": 512,
"sample_completions": [
"First, we need to find the length of the legs of the trapezoid. Since the trapezoid is isosceles, the legs are equal in length. Let's call the length of each leg $x$. We can use the Pythagorean theore",
" \nTo solve this problem, we need to analyze the given conditions and derive the form of the polynomial \\( P(x) \\). The key condition is that \\( P(k) = k^{2023} P\\left(1 - \\frac{1}{k}\\right) \\) for eve",
" To solve this problem, we need to find the maximum value of \\(a\\) such that the line \\(y = mx + 2\\) does not pass through any lattice points for \\(0 < x \\leq 100\\) when \\(\\frac{1}{2} < m < a\\).\n\nFirs"
],
"peak_memory_gb": 156.63964891433716,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
}
}

View file

@ -1,28 +0,0 @@
{
"backend": "vllm",
"lora_adapter": null,
"n_prompts": 16,
"n_prompt_tokens": 2061,
"n_decoded_tokens": 7259,
"wall_times_s": [
1.9610779809881933,
1.9720804590033367,
1.962739369017072
],
"median_wall_s": 1.962739369017072,
"prompt_tps": 1050.0630050703758,
"decode_tps": 3698.4024035933326,
"max_new_tokens": 512,
"sample_completions": [
"First, we need to find the length of the legs of the trapezoid. Since the trapezoid is isosceles, the legs are equal in length. Let's call the length of each leg $x$. We can use the Pythagorean theore",
" \nTo solve this problem, we need to analyze the given conditions and derive the form of the polynomial \\( P(x) \\). The key condition is that \\( P(k) = k^{2023} P\\left(1 - \\frac{1}{k}\\right) \\) for eve",
" To solve this problem, we need to find the maximum value of \\(a\\) such that the line \\(y = mx + 2\\) does not pass through any lattice points for \\(0 < x \\leq 100\\) when \\(\\frac{1}{2} < m < a\\).\n\nFirs"
],
"peak_memory_gb": 156.21798133850098,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
}
}

View file

@ -1,28 +0,0 @@
{
"backend": "vllm",
"lora_adapter": null,
"n_prompts": 256,
"n_prompt_tokens": 36963,
"n_decoded_tokens": 120774,
"wall_times_s": [
5.908074813021813,
5.705008256016299,
5.693635127041489
],
"median_wall_s": 5.705008256016299,
"prompt_tps": 6479.044085697884,
"decode_tps": 21169.82037188746,
"max_new_tokens": 512,
"sample_completions": [
"First, we need to find the length of the legs of the trapezoid. Since the trapezoid is isosceles, the legs are equal in length. Let's call the length of each leg $x$. We can use the Pythagorean theore",
" \nTo solve this problem, we need to analyze the given conditions and derive the form of the polynomial \\( P(x) \\). The key condition is that \\( P(k) = k^{2023} P\\left(1 - \\frac{1}{k}\\right) \\) for eve",
" To solve this problem, we need to find the maximum value of \\(a\\) such that the line \\(y = mx + 2\\) does not pass through any lattice points for \\(0 < x \\leq 100\\) when \\(\\frac{1}{2} < m < a\\).\n\nFirs"
],
"peak_memory_gb": 157.12996101379395,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
}
}

View file

@ -1,28 +0,0 @@
{
"backend": "vllm",
"lora_adapter": null,
"n_prompts": 32,
"n_prompt_tokens": 4847,
"n_decoded_tokens": 15097,
"wall_times_s": [
2.422758528031409,
2.3895113189937547,
2.388265542977024
],
"median_wall_s": 2.3895113189937547,
"prompt_tps": 2028.4482276656954,
"decode_tps": 6318.028242844854,
"max_new_tokens": 512,
"sample_completions": [
"First, we need to find the length of the legs of the trapezoid. Since the trapezoid is isosceles, the legs are equal in length. Let's call the length of each leg $x$. We can use the Pythagorean theore",
" \nTo solve this problem, we need to analyze the given conditions and derive the form of the polynomial \\( P(x) \\). The key condition is that \\( P(k) = k^{2023} P\\left(1 - \\frac{1}{k}\\right) \\) for eve",
" To solve this problem, we need to find the maximum value of \\(a\\) such that the line \\(y = mx + 2\\) does not pass through any lattice points for \\(0 < x \\leq 100\\) when \\(\\frac{1}{2} < m < a\\).\n\nFirs"
],
"peak_memory_gb": 156.21798133850098,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
}
}

View file

@ -1,28 +0,0 @@
{
"backend": "vllm",
"lora_adapter": null,
"n_prompts": 64,
"n_prompt_tokens": 9129,
"n_decoded_tokens": 30300,
"wall_times_s": [
2.916221586987376,
2.8970531829982065,
2.8907751629594713
],
"median_wall_s": 2.8970531829982065,
"prompt_tps": 3151.13303876329,
"decode_tps": 10458.9036120635,
"max_new_tokens": 512,
"sample_completions": [
"First, we need to find the length of the legs of the trapezoid. Since the trapezoid is isosceles, the legs are equal in length. Let's call the length of each leg $x$. We can use the Pythagorean theore",
" \nTo solve this problem, we need to analyze the given conditions and derive the form of the polynomial \\( P(x) \\). The key condition is that \\( P(k) = k^{2023} P\\left(1 - \\frac{1}{k}\\right) \\) for eve",
" To solve this problem, we need to find the maximum value of \\(a\\) such that the line \\(y = mx + 2\\) does not pass through any lattice points for \\(0 < x \\leq 100\\) when \\(\\frac{1}{2} < m < a\\).\n\nFirs"
],
"peak_memory_gb": 156.2349009513855,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
}
}

View file

@ -1,28 +0,0 @@
{
"backend": "vllm",
"lora_adapter": "outputs/lora_rank32_fresh",
"n_prompts": 64,
"n_prompt_tokens": 9129,
"n_decoded_tokens": 30163,
"wall_times_s": [
3.911774954001885,
3.8794285799958743,
3.8696357629960403
],
"median_wall_s": 3.8794285799958743,
"prompt_tps": 2353.181612125389,
"decode_tps": 7775.114138080634,
"max_new_tokens": 512,
"sample_completions": [
"First, we need to find the length of the legs of the trapezoid. Since the trapezoid is isosceles, the legs are equal in length. Let's call the length of each leg $x$. We can use the Pythagorean theore",
"Let $Q(x) = P(x) - x^{2023}P(1-\\frac{1}{x})$. Then $Q(k) = 0$ for every positive integer $1 \\leq k \\leq 2023$. Since $P(x)$ is a monic polynomial of degree $2023$, $Q(x)$ is also a monic polynomial of",
" To solve this problem, we need to find the maximum value of \\(a\\) such that the line \\(y = mx + 2\\) does not pass through any lattice points for \\(0 < x \\leq 100\\) when \\(\\frac{1}{2} < m < a\\).\n\nFirs"
],
"peak_memory_gb": 156.2349009513855,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
}
}

View file

@ -1,28 +0,0 @@
{
"backend": "vllm",
"lora_adapter": null,
"n_prompts": 8,
"n_prompt_tokens": 1009,
"n_decoded_tokens": 3961,
"wall_times_s": [
2.087421328993514,
2.0852316500386223,
2.084826519014314
],
"median_wall_s": 2.0852316500386223,
"prompt_tps": 483.87909323230895,
"decode_tps": 1899.5491459793616,
"max_new_tokens": 512,
"sample_completions": [
"First, we need to find the length of the legs of the trapezoid. Since the trapezoid is isosceles, the legs are equal in length. Let's call the length of each leg $x$.\n\nWe can use the Pythagorean theor",
" \nTo solve this problem, we need to analyze the given conditions and derive the form of the polynomial \\( P(x) \\). Let's start by examining the functional equation provided:\n\n\\[ P(k) = k^{2023} P\\left",
" To solve this problem, we need to find the maximum value of \\(a\\) such that the line \\(y = mx + 2\\) does not pass through any lattice points for \\(0 < x \\leq 100\\) when \\(\\frac{1}{2} < m < a\\).\n\nFirs"
],
"peak_memory_gb": 156.21798133850098,
"sampling": {
"temperature": 0.1,
"top_p": 0.97,
"min_p": 0.5,
"top_k": 5
}
}