MOVSB - Throughput and Uops
With unroll_count=500 and no inner loop
Code:
0: a4 movs BYTE PTR es:[rdi],BYTE PTR ds:[rsi]
Show nanoBench command
Results:
Instructions retired: 1.0
Core cycles: 3.97
Reference cycles: 3.32
UOPS_RETIRED.ANY: 5.0
RETIRE_SLOTS: 5.0
UOPS_MS: 1.0
UOPS_PORT_0: 0.65
UOPS_PORT_1: 0.65
UOPS_PORT_2: 1.0
UOPS_PORT_3: 1.0
UOPS_PORT_4: 1.01
UOPS_PORT_5: 0.0
DIV_CYCLES: 0.0
ILD_STALL.LCP: 0.0
INST_DECODED.DEC0: 1.0
With loop_count=1000 and unroll_count=10
Code:
0: a4 movs BYTE PTR es:[rdi],BYTE PTR ds:[rsi]
Show nanoBench command
Results:
Instructions retired: 1.2
Core cycles: 4.1
Reference cycles: 3.42
UOPS_RETIRED.ANY: 5.2
RETIRE_SLOTS: 5.2
UOPS_MS: 1.0
UOPS_PORT_0: 0.7
UOPS_PORT_1: 0.7
UOPS_PORT_2: 1.0
UOPS_PORT_3: 1.0
UOPS_PORT_4: 1.0
UOPS_PORT_5: 0.0
DIV_CYCLES: 0.0
ILD_STALL.LCP: 0.0
INST_DECODED.DEC0: 1.0
With loop_count=100 and unroll_count=100
Code:
0: a4 movs BYTE PTR es:[rdi],BYTE PTR ds:[rsi]
Show nanoBench command
Results:
Instructions retired: 1.02
Core cycles: 4.01
Reference cycles: 3.34
UOPS_RETIRED.ANY: 5.02
RETIRE_SLOTS: 5.02
UOPS_MS: 1.0
UOPS_PORT_0: 0.66
UOPS_PORT_1: 0.66
UOPS_PORT_2: 1.0
UOPS_PORT_3: 1.0
UOPS_PORT_4: 1.0
UOPS_PORT_5: 0.0
DIV_CYCLES: 0.0
ILD_STALL.LCP: 0.0
INST_DECODED.DEC0: 1.0