Achieve Up to 9x Faster & 13x Smaller Model Serving Compared to Naive Setups
Great job, thanks!
Have you benchmarked this implementation on NVIDIA GPUs?
Hey Bogdan, glad you enjoyed this. Unfortunately, I never got around to benchmarking this on NVIDIA hardware
Great job, thanks!
Have you benchmarked this implementation on NVIDIA GPUs?
Hey Bogdan, glad you enjoyed this. Unfortunately, I never got around to benchmarking this on NVIDIA hardware