• VNOJ
  • Trang chủ
  • Danh sách bài
  • Các bài nộp
  • Thành viên
  • Tổ chức
  • Các kỳ thi
  • Thông tin
    >
    • Máy chấm
    • Custom Checkers
    • Github
    • Giao diện
    • Ngôn ngữ VI EN
Đăng nhập  hoặc  Đăng ký

Blog - Trang 1

  • Thông tin
  • Thống kê
  • Blog

0

fix

yoshi_fp36 đã đăng vào 2, Tháng 9, 2026, 2:05

cortex x4 is actually 10 wide btw

also

we are so back

fujitsu monaka is a continuation of neoverse v1

i think it'll have higher perf than the ancient v1 (obv not as high as the modern x cores but on par with eg c1 pro is fine)

yoshi_fp36
o2, Tháng 9, 2026, 2:05 1

0

new chip 2026 btw

yoshi_fp36 đã đăng vào 31, Tháng 8, 2026, 15:58

near memory accelerator

1.1ghz

3072 riscv in order cores

capable of executing 3.3 tflops of dot product

integer rate not disclosed

vpu can do 128 flops (64 fmas?) per cycle

lowkey not fire chip vs the goated 3072 a55maxxing chip

yoshi_fp36
o31, Tháng 8, 2026, 15:58 0

0

this chip might be lowkey fire at this point

yoshi_fp36 đã đăng vào 31, Tháng 8, 2026, 15:52

hypothetical chip btw

tsmc n2

die area 400mm2

3072 A55 cores at 2 ghz

tdp about 150-250w

compute

24 tflops fp64

48 tflops fp32

96 tflops fp16

192 tops int8

advertisement

this chip has 96MB of ultra high bandwidth memory known as "L1 cache"

bandwidth up to 48tbps

end of advertisement

yoshi_fp36
o31, Tháng 8, 2026, 15:52 0

-2

4 wide alu in the beef 26 😭😭😭

yoshi_fp36 đã đăng vào 23, Tháng 8, 2026, 9:31

zen 4 is capable of over 2 spec17/ghz and 20 spec06/ghz with a 4 wide alu btw 😭😭😭

(haswell please say some)

highest performing zen4 implementations are on par with cortex-x2 (or slightly less) in perf per clock and much faster in perf per core...

also zen4 equalizes golden cove and skymont while lacking 4 WHOLE ALU ports btw (zen4 is faster than gracemont and s**lake)

😭😭😭

btw cortex a77 is the first arm core to have a 4 wide alu and cortex x2 is the last one in the cortex x line (x3/4 is 6 wide, x925/c1u is 8 wide)

-snapdragon 865 users (yes, the ONLY implementation of cortex-a77 that i know is kryo 5xx)

yoshi_fp36
o23, Tháng 8, 2026, 9:31 0

0

i reinvented lmh_gas btw 😭😭😭

yoshi_fp36 đã đăng vào 22, Tháng 8, 2026, 8:40

first time doing lmh_gas: talk about a 150 lines piece of segment tree code that i do not even understand after reading it again

second time doing lmh_gas (indirectly): reinvents the official solution

both are slope trick solutions btw 😭😭😭

yoshi_fp36
o22, Tháng 8, 2026, 8:40 0

-4

con gà màu xanh có mào màu cam

yoshi_fp36 đã đăng vào 16, Tháng 8, 2026, 9:35

gà này to lắm không biết có phải gà chọi không

yoshi_fp36
o16, Tháng 8, 2026, 9:35 1

-3

not the most accurate measurement of cortex-x2 (< 1% inaccuracy)

yoshi_fp36 đã đăng vào 16, Tháng 8, 2026, 1:16

Estimated clock speed> 2.56 GHz

Nops per clk> 10.69

Adds per clk> 3.88

XORs per clk> 3.88

CMPs per clk> 2.97

----Renamer Tests----

Indepdent movs per clk> 3.80

Dependent movs per clk> 1.36

eor -> 0 per clk> 1.00

mov -> 0 per clk> 5.72

sub -> 0 per clk> 1.00

----ALU Pipe Layout Tests----

Not taken jmps per clk> 1.91

Jump fusion test> 8.01

1:1 mixed not taken jmps / muls per clk> 3.76

1:2 mixed not taken jmps / muls per clk> 3.00

1:1 mixed not taken jmps / adds per clk> 3.76

1:2 mixed not taken jmps / adds per clk> 4.18

1:1 mixed add/mul per clk> 3.76

2:1 mixed add/mul per clk> 3.87

ror per clk> 3.76

1:1 mixed mul/ror per clk> 3.72

1:3 madd:add per clk> 3.59

32-bit mul per clk> 2.00

64-bit mul per clk> 2.00

64-bit multiply latency> 2.00 clocks

----ASIMD Crypto Tests----

aese per clk> 2.00

1:1 aese and vec 128 add per clk> 3.80

pmull per clk> 2.00

1:1 pmull and vec 128 add per clk> 4.00

----FP/ASIMD Tests----

scalar fp32 add per clk> 3.99

128-bit vec int32 add per clk> 3.99

128-bit vec int32 multiply per clk> 2.00

128-bit vec int32 mixed multiply and add per clk> 3.99

128-bit vec fp32 add per clk> 4.00

128-bit vec fp32 multiply per clk> 4.03

128-bit vec fp32 mixed multiply and add per clk> 3.99

1:1 mixed scalar adds and 128-bit vec int32 add per clk> 7.78

2:1 mixed scalar adds and 128-bit vec int32 add per clk> 5.86

3:1 mixed scalar adds and 128-bit vec int32 add per clk> 5.14

1:1 mixed scalar 32-bit multiply and 128-bit vec int32 multiply per clk> 4.00

1:1 mixed 128-bit vec fp32 multiply and 128-bit vec int32 multiply per clk> 4.01

1:1 mixed 128-bit vec fp32 add and 128-bit vec int32 add per clk> 3.99

1:2 mixed not taken jumps and 128-bit vec int32 add per clk> 5.69

1:1 mixed not taken jumps and 128-bit vec int32 mul per clk> 3.88

128-bit vec int32 add latency> 2.00 clocks

128-bit vec int32 mul latency> 4.00 clocks

Scalar FADD Latency> 2.00 clocks

128-bit vector FADD latency> 2.00 clocks

128-bit vector FMUL latency> 3.00 clocks

128-bit vector FMA per clk> 4.03

128-bit vector FMA latency> 4.00 clocks

Scalar FMA per clk> 4.03

Scalar FMA latency> 4.00 clocks

1:1 mixed 128-bit vector FMA/FADD per clk> 4.03

1:1 mixed 128-bit vector FMA/FMUL per clk> 4.01

----Load/Store Tests----

128-bit vec loads per clk> 3.00

128-bit vec stores per clk> 1.99

64-bit loads per clk> 2.98

1:1 mixed 64-bit loads/stores per clk> 2.65

2:1 mixed 64-bit loads/stores per clk> 2.73

there's still a lot of noise but it's within 1% tolerance.

yoshi_fp36
o16, Tháng 8, 2026, 1:16 0

-2

Cortex-A73

yoshi_fp36 đã đăng vào 15, Tháng 8, 2026, 12:27

Cortex-A73 is a dual issue out-of-order core. It doesn't use a ROB, and doesn't care about NOPs. At all!

The "slot-based" mechanism can reorder at least 76 instructions. This is low compared to modern cores, but within the range of its time.

Cortex-A73 has 7 execution ports. There are two 128-bit NEON ports for integer instructions and two ALU ports. Also included within is two 64-bit FP ports on top of the integer NEON ports.

Cortex-A73 can load up to 16 bytes per cycle, but real bandwidth measurements shows a lower result. (About 13.4 bytes per cycle.)

FMA latency is reduced from 8 to 7 cycles. It can do one 128-bit FMA per cycle, or two 64-bit FMAs per cycle.

There's an inherent limitation of 1.5 IPC for 128-bit vector operations. This is due to the register file being 6R3W, and the read paths are all 64-bit even for 128-bit instructions [speculation, citation needed].

Cortex-A73 can pull at a rate of 11GB/s from main memory, as mentioned in previous articles. Compared to high-performance cores of the 2020s, A73 falls short but is slightly better than the best A55 implementation (ref: Genio 1200).

yoshi_fp36
o15, Tháng 8, 2026, 12:27 0

-2

most accurate measurement of cortex-a73 (< 1% inaccuracy)

yoshi_fp36 đã đăng vào 15, Tháng 8, 2026, 11:47

Estimated clock speed> 2.45 GHz

Nops per clk> 1.88

Adds per clk> 1.88

XORs per clk> 1.88

CMPs per clk> 1.88

----Renamer Tests----

Indepdent movs per clk> 1.82

Dependent movs per clk> 1.82

eor -> 0 per clk> 1.00

mov -> 0 per clk> 1.82

sub -> 0 per clk> 1.00

----ALU Pipe Layout Tests----

Not taken jmps per clk> 0.95

Jump fusion test> 2.00

1:1 mixed not taken jmps / muls per clk> 1.88

1:2 mixed not taken jmps / muls per clk> 1.50

1:1 mixed not taken jmps / adds per clk> 1.88

1:2 mixed not taken jmps / adds per clk> 1.78

1:1 mixed add/mul per clk> 1.88

2:1 mixed add/mul per clk> 1.87

ror per clk> 1.88

1:1 mixed mul/ror per clk> 1.88

1:3 madd:add per clk> 1.52

32-bit mul per clk> 1.00

64-bit mul per clk> 1.00

64-bit multiply latency> 4.00 clocks

----ASIMD Crypto Tests----

aese per clk> 1.00

1:1 aese and vec 128 add per clk> 1.20

pmull per clk> 1.00

1:1 pmull and vec 128 add per clk> 1.20

----FP/ASIMD Tests----

scalar fp32 add per clk> 1.88

128-bit vec int32 add per clk> 1.21

128-bit vec int32 multiply per clk> 0.50

128-bit vec int32 mixed multiply and add per clk> 1.00

128-bit vec fp32 add per clk> 1.00

128-bit vec fp32 multiply per clk> 1.00

128-bit vec fp32 mixed multiply and add per clk> 1.00

1:1 mixed scalar adds and 128-bit vec int32 add per clk> 1.94

2:1 mixed scalar adds and 128-bit vec int32 add per clk> 1.92

3:1 mixed scalar adds and 128-bit vec int32 add per clk> 1.90

1:1 mixed scalar 32-bit multiply and 128-bit vec int32 multiply per clk> 1.00

1:1 mixed 128-bit vec fp32 multiply and 128-bit vec int32 multiply per clk> 0.67

1:1 mixed 128-bit vec fp32 add and 128-bit vec int32 add per clk> 1.00

1:2 mixed not taken jumps and 128-bit vec int32 add per clk> 1.92

1:1 mixed not taken jumps and 128-bit vec int32 mul per clk> 1.00

128-bit vec int32 add latency> 3.00 clocks

128-bit vec int32 mul latency> 4.00 clocks

Scalar FADD Latency> 4.00 clocks

128-bit vector FADD latency> 4.00 clocks

128-bit vector FMUL latency> 4.00 clocks

128-bit vector FMA per clk> 1.00

128-bit vector FMA latency> 7.00 clocks

Scalar FMA per clk> 1.87

Scalar FMA latency> 7.00 clocks

1:1 mixed 128-bit vector FMA/FADD per clk> 0.96

1:1 mixed 128-bit vector FMA/FMUL per clk> 1.00

----Load/Store Tests----

128-bit vec loads per clk> 0.84

128-bit vec stores per clk> 0.50

64-bit loads per clk> 1.00

1:1 mixed 64-bit loads/stores per clk> 1.82

2:1 mixed 64-bit loads/stores per clk> 1.41

reference clock is calculated using 128-bit vector integer multiply latency

yoshi_fp36
o15, Tháng 8, 2026, 11:47 0

1

me when i raged

yoshi_fp36 đã đăng vào 11, Tháng 8, 2026, 15:20

my temper clock frequency spikes to 6.0ghz and generating hundreds of watts of heat the moment somebody tried to play games with me

yoshi_fp36
o11, Tháng 8, 2026, 15:20 0
  • «
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • »

dựa trên nền tảng DMOJ | theo dõi VNOI trên Github và Facebook