Improve the precision of the FusedAddRMSNormKernel function #587

Abatom · 2024-11-06T06:56:33Z

When sizeof(T) == 2, the sum of the read input and residual (float x) is split into two parts, high and low 16 bits, and saved to input and residual respectively. Later, input and residual are read out and combined to x, with the aim of improving the precision of the subsequent x * rms_rcp operation.

Increase precision from 1e-2 to 1e-3.

Abatom · 2024-11-06T07:02:13Z

def fused_add_rms_norm(x, residual, weight, eps):
    orig_dtype = x.dtype
    x = x.to(torch.float32)
    x = x + residual.to(torch.float32)
    residual = x.to(orig_dtype)

    variance = x.pow(2).mean(dim=-1, keepdim=True)
    x = x * torch.rsqrt(variance + eps)
    x = x.to(orig_dtype) * weight
    return x, residual

If the function is modified as follows, the output result of the fused_add_rms_norm function will be almost the same as that of FusedAddRMSNormKernel, the precision can reach 1e-20.

def fused_add_rms_norm(x, residual, weight, eps):
    orig_dtype = x.dtype
    x = x.to(torch.float32)
    x = x + residual.to(torch.float32)
    residual = x.to(orig_dtype)

    variance = x.pow(2).mean(dim=-1, keepdim=True)
    x = x * torch.rsqrt(variance + eps) * weight.to(orig_dtype)
    return x.to(orig_dtype), residual

yzh119

Nice contribution, thank you @Abatom !
Left some comments for discussion.

include/flashinfer/norm.cuh

zhyncs · 2024-11-06T09:05:06Z

It's better to add the benchmark result for the new one @Abatom

yzh119 · 2024-11-06T09:19:01Z

@zhyncs we haven't set up a standard benchmark for normalization kernels so I think we can leave it for further work.

One interesting feature to have in flashinfer is to add benchmarking class that returns bandwidth and FLOP utilization like proton. Ideally we can port nvbench to python but I don't have a concrete idea about the amount of work.

Abatom · 2024-11-06T11:17:06Z

@yzh119 The shared memory has already been used in place of global memory, and an global memory read has also been reduced.

yzh119

LGTM, I think this PR is ready to be merged.

Brief note (to remind myself what this PR is doing): keep residual in fp32 in shared memory to increase the numerical accuracy of rmsnorm.

gemma-style rmsnorm kernels (introduced in #477 ) are similar to original rmsnorm kernel, and we should use the same kernel for them. This PR cleans up duplicate code and unifies the kernels for gemma-style and original rmsnorm kernels. The precision improvements (#587, #592) are kept in this PR.

Abatom added 2 commits November 6, 2024 14:30

Improve Precision

bcaea47

1e-2 -> 1e-3

26c0119

yzh119 reviewed Nov 6, 2024

View reviewed changes

include/flashinfer/norm.cuh Outdated Show resolved Hide resolved

Abatom added 2 commits November 6, 2024 19:01

use shared memory

6e384c3

use shared memory

bef7a5a

Abatom requested a review from yzh119 November 6, 2024 11:21

yzh119 approved these changes Nov 6, 2024

View reviewed changes

yzh119 merged commit c7dc921 into flashinfer-ai:main Nov 6, 2024

github-actions bot mentioned this pull request Nov 6, 2024

chore(main): release 0.2.0 #476

Open

Abatom mentioned this pull request Nov 8, 2024

perf: reduce the read and write of shared memory in the FusedAddRMSNormKernel #592

Merged

yzh119 mentioned this pull request Nov 24, 2024

misc: remove duplicate norm cuda kernels #631

Merged

yzh119 mentioned this pull request Dec 1, 2024

[Bug] flashinfer's RMSNorm implementation causes precision differences in model outputs compared to the HuggingFace implementation sgl-project/sglang#2258

Closed

5 tasks

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Improve the precision of the FusedAddRMSNormKernel function #587

Improve the precision of the FusedAddRMSNormKernel function #587

Abatom commented Nov 6, 2024 •

edited

Loading

Abatom commented Nov 6, 2024

yzh119 left a comment •

edited

Loading

zhyncs commented Nov 6, 2024

yzh119 commented Nov 6, 2024

Abatom commented Nov 6, 2024

yzh119 left a comment

Improve the precision of the FusedAddRMSNormKernel function #587

Improve the precision of the FusedAddRMSNormKernel function #587

Conversation

Abatom commented Nov 6, 2024 • edited Loading

Abatom commented Nov 6, 2024

yzh119 left a comment • edited Loading

Choose a reason for hiding this comment

zhyncs commented Nov 6, 2024

yzh119 commented Nov 6, 2024

Abatom commented Nov 6, 2024

yzh119 left a comment

Choose a reason for hiding this comment

Abatom commented Nov 6, 2024 •

edited

Loading

yzh119 left a comment •

edited

Loading