AI coding assistance discussion

igor_kavinski · Mar 2, 2025

Very relevant post: https://forums.anandtech.com/threads/rdna4-cdna3-architectures-thread.2602668/post-41407076

igor_kavinski · Mar 3, 2025

moonshotai/Moonlight-16B-A3B · Hugging Face

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

huggingface.co

Pretty CRAZY model.

Don't believe me? Ask it something and watch it go. Like really, really go.

It doesn't stop until it runs against some sort of limit. Keeps going through different code possibilities.

C1 · Mar 3, 2025

In today's news (and hopefully not old news):

Google launches a free AI coding assistant – What to know

igor_kavinski · Mar 7, 2025

Speculative decoding now allows a larger 231B model to oversee the draft work of the smaller 13B model, resulting in improved response times.

igor_kavinski · Mar 7, 2025

igor_kavinski · Mar 7, 2025

WOW!

"Shut up and code!" and it actually complied!

igor_kavinski · Mar 7, 2025

YES!!!!!

igor_kavinski · Mar 7, 2025

Let me know if any CPU owners with AVX-512 support wish to run this.

igor_kavinski · Mar 7, 2025

It's LIVE!

Discussion - Rudi_Float_Bench v0.02a

v0.02 with an additional AVX-512 specific binary (it will either crash or exit unexpectedly if the CPU is lacking the necessary ISA extensions): https://drive.google.com/file/d/12RuZsWdhNueu7th2HCuzblBA8CUGFu9u/view?usp=sharing If it complains about missing vcruntime140 DLL something, install...

forums.anandtech.com

Kids can now create their own CPU benchmarks!

(yes, I'm a 44 year old kid...)

igor_kavinski · Mar 11, 2025

RAM latency checker: https://www.overclock.net/posts/29439133/

As described there, it's not the absolute latency but it seems fairly consistent.

Tested to work on Haswell and onwards. Don't think I can try it on my Epyc today so someone may wanna volunteer and test on their Ryzen? Thanks!

EDIT: Tested and working as intended on Tiger Lake. Average latency deviation isn't wild which means it can be useful.

Red Squirrel · Mar 19, 2025

Lol this is pretty funny and why it's rather important to understand how to code if you're going to use AI.

Red Squirrel · Mar 19, 2025

I'm starting to think this guy is trolling lol.

igor_kavinski · Mar 30, 2025

A disappointment to report, hoping it would dissuade someone else from investing in expensive hardware (good thing LLM wasn't the only thing I bought the laptop for).

So my Thinkpad now has 128GB RAM and RTX 5000 16GB dGPU. I was hoping I would be able to run Llama 3.3 70B. It loads, at a context length of 16384 and consumes 71GB system RAM and all of VRAM. Unfortunately, the calculations are not offloaded to the GPU, despite lowering the core count to 1 and using all 80 cores of the GPU. It stays at 0% utilization. The processing happens on the CPU and even when setting it to max 6 cores (HT not supported by LM Studio I guess), the CPU utilization does not go beyond 17%. It gives a response, at the most horrible speed of something like 0.05 tokens per second or even lower. Gave up on it and now downloading another 8B LLM at F16 and Q8, to take advantage of speculative decoding. If I still don't get any GPU utilization, I will need to troubleshoot (maybe driver issue?).

igor_kavinski · Apr 1, 2025

LM Studio can't use both GPUs in parallel so one of them is doing the hard work while the other is chilling and just holding some data in its VRAM.

igor_kavinski · Apr 2, 2025

Tried the same prompt with/without GPU offloading and in the CPU only scenario, it created twice as many tokens to arrive at the solution. May have to do that a number of times to verify if this behavior is consistent but it does beg the question why the "thinking" is better with the involvement of GPU.

igor_kavinski · Apr 5, 2025

Intel Core i9-10980XE vs. Intel Core Ultra 9 285K vs. AMD Ryzen 9 9950X3D vs. AMD Ryzen 7 9800X3D vs. AMD Ryzen 9 9950X Benchmarks - OpenBenchmarking.org

openbenchmarking.org

Search Llama, Whisper and OpenVINO and marvel at the total domination of Zen 5 in these benchmarks.

igor_kavinski · Apr 7, 2025

Meta-KI: Erste Llama-4-Modelle sind besonders effizient – aber nicht in der EU

Meta präsentiert die ersten Llama-4-Modelle. Die setzen sich noch nicht an die Leistungsspitze, sollen aber besonders effizient sein.

www-computerbase-de.translate.goog

Llama 4 Scout looks very nice for offline use on less powerful PCs.

igor_kavinski · Apr 14, 2025

Was planning to benchmark my 9950X3D using LG ExaOne Deep F16 model. On the Xeon 6248R, I got a speed of roughly 3.7 tokens per second. Tried it at home and first, it loads up only the bottom half of the threads in Task Manager. Second, it keeps processing and never gets to the "thinking" stage. Just wastes a whole lot of power for nothing. So I suspect:

1) LM Studio isn't optimized for 9950X3D or getting confused by the CCD crap.

2) LM Studio is secretly co-owned by Intel or AMD or both and it only works flawlessly on server CPUs.

Extremely annoyed since the Xeon was only 41% utilized with max speeds of 3.9 GHz on 24 cores while the 9950X3D was hitting 5+ GHz on 16 cores and still failed to progress to the thinking stage.

MS_AT · Apr 15, 2025

igor_kavinski said:
Extremely annoyed since the Xeon was only 41% utilized with max speeds of 3.9 GHz on 24 cores while the 9950X3D was hitting 5+ GHz on 16 cores and still failed to progress to the thinking stage.

The prompt processing part is compute heavy and during that on 16 threads your clocks should be sinking low if the code is well optimized. SMT will be by definition useless in this case.

The token generation part in single user case is dominated by memory BW, during that the clocks will be high, and you could get by with fewer than 16 threads even (it takes 2, pinned to different CCDs to maximize MemBW usage and if more threads help will depend on the model compute needs)

https://github.com/ikawrakow/ik_llama.cpp discussions in this repo are quite insightful, as well as in https://github.com/ggml-org/llama.cpp

igor_kavinski · Apr 20, 2025

Athene V2 Chat IQ4_XS 73B model

9950X3D

6200C52 FCLK 2133 UCLK 3100 CO -37

Pretty impressive that it's maintaining a solid 5.35 GHz speed, even with my less than stellar 240mm AIO cooler.

dank69 · Apr 22, 2025

I wouldn't trust AI for anything beyond scaffolding. The wettest of the wet jr. coders.

igor_kavinski · Apr 28, 2025

MS_AT said:
The token generation part in single user case is dominated by memory BW, during that the clocks will be high, and you could get by with fewer than 16 threads even (it takes 2, pinned to different CCDs to maximize MemBW usage and if more threads help will depend on the model compute needs)

I found out why I've been getting really low token/s speeds compared to what people say in Reddit threads. They are using low parameter models (in other words, almost crap and useless).

I tried StarCoder 10.7B.

Results:

9950X3D 7200C34 ~10.5 tokens/sec
9950X3D 7600C36 ~11.5 tokens/sec

A770 16GB ~25 tokens/sec

This is the only model where the A770 is able to shine so far in my testing (not that I've been able to try that many).

Seems large parameter count models make even GPUs crumble to their knees.

But the response quality of the Athene V2 Chat model was much higher and it made more of an effort with an elaborate solution to solve the problem (asked it to provide code for the best sort algorithm) whereas StarCoder kept its response short and not as comprehensive.

MS_AT · Apr 28, 2025

igor_kavinski said:
They are using low parameter models (in other words, almost crap and useless).

You don't tell us the quantization you are using, for example for programming AMD suggested Q6 in their materials. Also for programming try qwen 2.5 coder, it should do better in theory.

igor_kavinski · Apr 28, 2025

IQ4_XS because the higher ones don't necessarily give better results.

igor_kavinski · Apr 28, 2025

I found an extremely lovely (I don't say that lightly because it made me really happy) tiny model: tiny-llama-R1.

Just watched it go crazy fast using only the 9950X3D at 67.5 tokens/sec @ IQ4 quant!!!

I'm almost too giddy to try it on the GPU because it might break 100 tokens/sec there

AI coding assistance discussion

Lifer

Lifer

Platinum Member

Lifer

Lifer

Lifer

Lifer

Lifer

Lifer

Lifer

No Lifer

No Lifer

Lifer

Lifer

Lifer

Lifer

Lifer

Lifer

Senior member

Lifer

Lifer

Lifer

Senior member

Lifer

Lifer