# A B C D E F G H I J K L M N O P Q R S T U V W X Y Z

Dell PowerEdge YX4X Server Memory RAS Whitepaper

Summary

Essential resource for system administrators managing Dell EMC PowerEdge YX4X servers. This whitepaper demystifies complex memory health by detailing advanced Reliability, Availability, and Serviceability (RAS) features. It educates users on soft vs. hard memory errors, guides best practices for maximizing server uptime, and provides actionable troubleshooting steps when encountering memory issues in Intel Xeon SP platforms.

📄 Preview PAGE OF 21

Page 1 Text Content

Whitepaper

Memory Errors and Dell EMC Power Edge YX4X Server Memory RAS Features Revision: 1.4 Issue Date: 02/28/2022 Issue Date: 3/1/2022 Introduction Memory sub-system errors are some of the most common types of errors seen on modern computing systems. Understanding how memory errors occur and how to prevent or avoid them can be a complex subject – one that has challenged countless numbers of industry researchers and developers over the last 30 years. While Dell EMC Power Edge servers are designed to provide industry leading Reliability, Availability, and Serviceability (RAS) on memory issues, we realize that many of our technically savvy customers may want to know more on what’s happening ‘under the hood’ of their servers. This technical whitepaper is divided in four sections to help Power Edge users to understand about the following memory error topics: • Types of memory errors and how they may affect a server • Dell EMC Power Edge YX4X server memory RAS capabilities • Configuring a Power Edge server to achieve maximum memory up-time • Recommended user actions when encountering memory errors

Important: The content covered in this whitepaper applies to Dell EMC Power Edge YX4X servers with Intel Xeon SP processors. Customers with YX4X servers that utilize AMD EPYC or Intel Xeon E processors should refer to v1.0 of the RAS whitepaper.

The features described in this document assume the user is running the latest versions of Dell EMC Power Edge server firmware, such as BIOS and i DRAC.

1 Memory Errors and Dell Power Edge YX4X Server Memory RAS Features

Page Summary Contents For Dell PowerEdge YX4X Server Memory RAS Whitepaper

Page 1 Whitepaper Memory Errors and Dell EMC Power Edge YX4X Server Memory RAS Features Revision: 1.4 Issue Date: 02/28/2022 Issue Date: 3/1/2022 Introduction Memory sub-system errors are some of the most co...
Page 2 Revisions Date Author Description January 3, 2020 Jordan Chin • Initial release • Removed content for platforms based on AMD EPYC and Xeon E processors • Added more information to primer on uncorrecta...
Page 3 Mark Farley Component Quality Engineering, Senior Principal Engineer, Dell EMC A Primer on Memory Errors To fully understand the memory RAS response capabilities of Power Edge servers, it is first hel...
Page 4 phenomenon such as Variable Retention Time (VRT) [1] and Random Telegraph Noise (RTN) [2]. o Within the server industry, it is an increasingly accepted understanding, shared by Dell, that some correct...
Page 5 A Primer on Dell EMC Power Edge Server Memory RAS Capabilities Previously discussed memory errors are mitigated through Power Edge server memory RAS capabilities which entail fault avoidance, detectio...
Page 6 symbols can vary depending upon the processor architecture. But regardless if the symbol size is 4-bits or 32-bits, as the SSC-DSD name implies, the coding is designed such that a single symbol may be...
Page 7 XX 140X . . . XX 68X Figure 2 - Advanced ECC can correct multi-bit errors in a single symbol… Figure 3 - But Advanced ECC cannot correct errors in multiple symbols As described earlier, SSC-DSD implem...
Page 8 Adaptive Double Device Data Correction (ADDDC) ADDDC Feature Support Table x4 DIMMs:  DIMMs Supported x8 DIMMs:  Memory Configuration • Two or more memory ranks per memory channel Required Adaptive ...
Page 9 Memory Engineering recommends contacting Dell technical support to schedule replacement of the affected DIMM at the next service opportunity. Memory patrol scrubbing is enabled by default and configur...
Page 10 Memory Rank Sparing is a memory RAS feature available on Intel platforms that will reserve one or more memory ranks per channel as spares for failover. When the Power Edge server memory health monitor...
Page 11 o E.g. 4 ranks = 50% reduction • Largest rank in channel is always held as spare o E.g. One 32 GB RDIMM (2Rx4) and one 16 GB RDIMM (2Rx8) installed = two 16 GB ranks and two 8 GB ranks. Both 16 GB ran...
Page 12 Important: Consult your Power Edge server installation and service manual for complete memory population guidelines to properly enable Memory Mirroring. Fault Resilient Memory (FRM) Fault Resilient Me...
Page 13 Memory channels must be populated with all one DIMM or all two DIMMs (for example, 24 DIMM systems should have 12 DIMMs or 24 DIMMs installed). Fault Resilient Memory is disabled by default and must b...
Page 14 Figure 7 - PPR for a row in a bank group of a 4Gb x4 device PPR is always available on Power Edge server platforms that support it and if deemed necessary by BIOS will automatically execute after a sy...
Page 15 • If the impacted data was in user/application/VM memory, then the OS will terminate the associated process or VM without impacting the rest of the system. • If the impacted data was in user/applicati...
Page 16 o Benefit: Patrol scrub will run every four hours (instead of 24); increased frequency will reduce the accumulation of errors in areas of memory with low utilization and thus not being corrected by de...
Page 17 • MEM0804 – This is an indication that the system has successfully performed memory-self healing at the specified DIMM location in the event message. o Recommended Response Action: No response require...
Page 18 • Power Edge XR2* • Power Edge R440* • Power Edge R540* • Power Edge R640 • Power Edge R740 • Power Edge R740xd • Power Edge R740xd2 • Power Edge R840 • Power Edge R940 • Power Edge R940xa • Power Edg...
Page 19 MEM0804 (success) or MEM0805 (failure) in the System Event Log (SEL) for an indication of next steps. • Updates to Self-Healing Algorithm – Prior to this update, Power Edge server BIOS would rigorousl...
Page 20 o Any PPR or Memory retraining failures will be still be logged for corrective action. • Uncorrectable ECC Memory events o All the BIOS triggers for the scheduling of DIMM "self healing" (PP...
Page 21 Legal Notices THIS WHITE PAPER IS FOR INFORMATIONAL PURPOSES ONLY, AND MAY CONTAIN TYPOGRAPHICAL ERRORS AND TECHNICAL INACCURACIES. THE CONTENT IS PROVIDED AS IS, WITHOUT EXPRESS OR IMPLIED WARRANTIES...

Manual Details

Brand Server
Pages 21
File Size 592.13 KB
Published May 31, 2026
76 views

Enter the captcha to get the download link:

captcha

Frequently Asked Questions

Which processors does this whitepaper apply to?

It applies to Dell EMC PowerEdge YX4X servers utilizing Intel Xeon SP processors.

How do soft and hard memory errors differ?

Soft errors are transient disturbances, but hard errors are persistent faults that cannot be resolved by system resets or power cycles.

What firmware versions must be used for updated uncorrectable event guidance?

You must install iDRAC 5.10.10.00 or newer to reflect the latest memory engineering guidance for uncorrectable events.

How are single-bit soft errors managed by the server?

Single and multiple bit soft errors can be corrected using built-in demand or patrol scrubbing mechanisms.