Showing posts with label Hardware. Show all posts
Showing posts with label Hardware. Show all posts

Friday, August 3, 2012

"No POST", "system won't boot", and "no video output" checklist

SkyHi @ Friday, August 03, 2012

This checklist is a compilation of troubleshooting ideas from many forum members. It's very important to actually perform every step in the checklist if you want to effectively troubleshoot your problem.
 

1. Did you carefully read the motherboard owners manual?
 
2. Did you plug in the 4/8-pin CPU power connector located near the CPU socket? If the motherboard has 8 pins and your PSU only has 4 pins, you can use the 4-pin connector. The 4-pin connector USUALLY goes on the 4 pins located closest to the CPU. If the motherboard has an 8-pin connector with a cover over 4 pins, you can remove the cover and use an 8-pin plug if your power supply has one. This power connector provides power to the CPU. Your system has no chance of posting without this connector plugged in! Check your motherboard owners manual for more information about the CPU power connector. The CPU power connector is usually referred to as the "12v ATX" connector in the owners manual. This is easily the most common new-builder mistake.
 
http://i44.tinypic.com/jtkves.jpg
http://i40.tinypic.com/qx9gdy.png
http://i40.tinypic.com/r6zecn.png
 
3. Did you install the standoffs under the motherboard? Did you place them so they all align with the screw holes in the motherboard, with no extra standoffs touching the board in the wrong place? A standoff installed in the wrong place can cause a short and prevent the system from booting.
 
http://i42.tinypic.com/fwq1ps.jpg
http://i39.tinypic.com/98a7u0.jpg
 
4. Did you verify that the video card is fully seated? (may require more force than a new builder expects.)
 
5. Did you attach all the required power connector(s) to the video card? (some need two, some need none, many need one.)
 
http://i43.tinypic.com/2hcq17b.png
http://i44.tinypic.com/9fpds8.png
 
6. Have you tried booting with just one stick of RAM installed? (Try each stick of RAM individually in each RAM slot.) If you can get the system to boot with a single stick of RAM, you should manually set the RAM speed, timings, and voltage to the manufacturers specs in the BIOS before attempting to boot with all sticks of RAM installed. Nearly all motherboards default to the standard RAM voltage (1.8v for DDR2 & 1.5v for DDR3). If your RAM is rated to run at a voltage other than the standard voltage, the motherboard will underclock the RAM for compatibility reasons. If you want the system to be stable and to run the RAM at its rated specs, you should manually set those values in the BIOS. Many boards don't supply the RAM with enough voltage when using "auto" settings causing stability issues.
 
7. Did you verify that all memory modules are fully inserted? (may require more force than a new builder expects.) It's a good idea to install the RAM on the motherboard before it's in the case.
 
8. Did you verify in the owners manual that you're using the correct RAM slots? Many i7 motherboards require RAM to be installed in the slots starting with the one further away from the CPU which is the opposite of many dual channel motherboards.
 
9. Did you remove the plastic guard over the CPU socket? (this actually comes up occasionally.)
 
10. Did you install the CPU correctly? There will be an arrow on the CPU that needs to line up with an arrow on the motherboard CPU socket. Be sure to pay special attention to that section of the manual!
   
11. Are there any bent pins on the motherboard/CPU? This especially applies if you tried to install the CPU with the plastic cover on or with the CPU facing the wrong direction.
 
12. If using an after market CPU cooler, did you get any thermal paste on the motherboard, CPU socket, or CPU pins? Did you use the smallest amount you could? Here's a few links that may help:
       
13. Is the CPU fan plugged in? Some motherboards will not boot without detecting that the CPU fan is plugged in to prevent burning up the CPU.
 
14. If using a stock cooler, was the thermal material on the base of the cooler free of foreign material, and did you remove any protective covering? If the stock cooler has push-pins, did you ensure that all four pins snapped securely into place? (The easiest way to install the push-pins is outside the case sitting on a non-conductive surface like the motherboard box. Read the instructions! The push-pins have to be turned the OPPOSITE direction as the arrows for installation.) See the link in step 10.
 
15. Are any loose screws laying on the motherboard, or jammed against it? Are there any wires run directly under the motherboard? You should not run wires under the motherboard since the soldered wires on the underside of the motherboard can cut into the insulation on the wires and cause a short. Some cases have space to run wires on the back side of the motherboard tray.
 
16. Did you ensure you discharged all static electricity before touching any of your components? Computer components are very sensitive to static electricity. It takes much less voltage than you can see or feel to damage components. You should implement some best practices to reduce the probability of damaging components. These practices should include either wearing an anti-static wrist strap or always touching a metal part of the case with the power supply installed and plugged in, but NOT turned on. You should avoid building or working on a computer on carpet. Working on a smooth surface is the best if at all possible. You should also keep fluffy the cat, children, and fido away from computer components.
 
17. Did you install the system speaker (if provided) so you can check beep-codes in the manual? A system speaker is NOT the same as normal speakers that plug into the back of the motherboard. A system speaker plugs into a header on the motherboard that's usually located near the front panel connectors. The system speaker is a critical component when trying to troubleshoot system problems. You are flying blind without a system speaker. If your case or motherboard didn't come with a system speaker you can buy one for cheap here: http://www.cwc-group.com/casp.html
 
http://i43.tinypic.com/2lsjlzr.jpg
 
18. Did you read the instructions in the manual on how to properly connect the front panel plugs? (Power switch, power led, reset switch, HD activity led) Polarity does not matter with the power and reset switches. If power or drive activity LED's do not come on, reverse the connections. For troubleshooting purposes, disconnect the reset switch. If it's shorted, the machine either will not POST at all, or it will endlessly reboot.
 
http://i42.tinypic.com/2cftmzb.jpg
http://i42.tinypic.com/20fc18g.jpg
 
19. Did you turn on the power supply switch located on the back of the PSU? Is the power plug on a switch? If it is, is the switch turned on? Is there a GFI circuit on the plug-in? If there is, make sure it isn't tripped. You should also make sure the power cord isn't causing the problem. Try swapping it for a known good cord if you have one available.
 
20. Is your CPU supported by the BIOS revision installed on your motherboard? Most motherboards will post a CPU compatibility list on their website.
 
21. Have you tried resetting the CMOS? The motherboard manual will have instructions for your particular board.
   
22. If you have integrated video and a video card, try the integrated video port. Resetting the bios, can make it default back to the onboard video.
 
23. Make certain all cables and components including RAM and expansion cards are tight within their sockets. Here's a thread where that was the cause of the problem.
 

I also wanted to add some suggestions that user jsc often posts. This is a direct quote from him:
 
"Pull everything except the CPU and HSF. Boot. You should hear a series of long single beeps indicating memory problems. Silence here indicates, in probable order, a bad PSU, motherboard, or CPU - or a bad installation where something is shorting and shutting down the PSU.
 
To eliminate the possiblility of a bad installation where something is shorting and shutting down the PSU, you will need to pull the motherboard out of the case and reassemble the components on an insulated surface. This is called "breadboarding" - from the 1920's homebrew radio days. I always breadboard a new or recycled build. It lets me test components before I go through the trouble of installing them in a case.
 
If you get the long beeps, add a stick of RAM. Boot. The beep pattern should change to one long and two or three short beeps. Silence indicates that the RAM is shorting out the PSU (very rare). Long single beeps indicates that the BIOS does not recognize the presence of the RAM.
 
If you get the one long and two or three short beeps, test the rest of the RAM. If good, install the video card and any needed power cables and plug in the monitor. If the video card is good, the system should successfully POST (one short beep, usually) and you will see the boot screen and messages.
 
Note - an inadequate PSU will cause a failure here or any step later.
Note - you do not need drives or a keyboard to successfully POST (generally a single short beep).
 
If you successfully POST, start plugging in the rest of the components, one at a time."
 

If you suspect the PSU is causing your problems, below are some suggestions by jsc for troubleshooting the PSU. Proceed with caution. I will not be held responsible if you get shocked or fry components.
 
"The best way to check the PSU is to swap it with a known good PSU of similar capacity. Brand new, out of the box, untested does not count as a known good PSU. PSU's, like all components, can be DOA.
 
Next best thing is to get (or borrow) a digital multimeter and check the PSU.
 
Yellow wires should be 12 volts. Red wires: +5 volts, orange wires: +3.3 volts, blue wire : -12 volts, violet wire: 5 volts always on. Tolerances are +/- 5% except for the -12 volts which is +/- 10%.
 
The gray wire is really important. It should go from 0 to +5 volts when you turn the PSU on with the case switch. CPU needs this signal to boot.
 
You can turn on the PSU by completely disconnecting the PSU and using a paperclip or jumper wire to short the green wire to one of the neighboring black wires.
   
This checks the PSU under no load conditions, so it is not completely reliable. But if it can not pass this, it is dead. Then repeat the checks with the PSU plugged into the computer to put a load on the PSU. You can carefully probe the pins from the back of the main power connector."
 

Here's a link to jsc's breadboarding thread:
   

If you make it through the entire checklist without success, Proximon has put together another great thread with a few more ideas here:
   
Here's a couple more sites that may help:
 


REFERENCES

Friday, July 20, 2012

Different processors on a dual socket motherboard?

SkyHi @ Friday, July 20, 2012
The answer is that particular combination won't work together. Page 25 and 26 of the Intel Xeon 5000 Series Datasheet Volume 1 discuss mixing processor types. A brief rundown on what's supported and not: 

1. CPUs must have the same QPI and RAM speed to work together. 
2. CPUs must have the same thermal profile (TDP) to work together. 
3. The CPUs must have the same number of physical cores to work together. 
4. The CPUs must have the same number of logical cores to work together. 
5. Stepping does not matter. 
6. Clock speed does not matter. 

If all of those are met, then you can run the CPUs together. The CPUs run at the speed of the slowest CPU. The current mix and match list: 

1. Xeon E5502- can only work with another E5502 since it's the only dual-core Xeon 5500 (#4.) 
2. Xeon E5504 can work with the E5506, both CPUs run at 2.00 GHz. 
3. Xeon E5520 can work with the E5530 or E5540, both CPUs run at 2.26 GHz in either case. 
4. Xeon E5530 can work with the E5520 (CPUs run at 2.26 GHz) or E5540 (CPUs run at 2.40 GHz) 
5. Xeon X5550 can work with the X5560 or X5570, both CPUs run at 2.67 GHz in either case. 
6. Xeon X5560 can work with the X5550 (CPUs run at 2.67 GHz) or X5570 (CPUs run at 2.80 GHz) 
7. Xeon W5580 can work with the W5590, both CPUs run at 3.20 GHz) 
8. Xeon L5506- can only work with other L5506s since it's the only L-series Xeon 5500 without HyperThreading. 
9. Xeon L5518 can work with the L5520 or L5530, both CPUs run at 2.13 GHz in either case. 
10. Xeon L5520 can work with the L5518 (CPUs run at 2.13 GHz) or L5530 (CPUs run at 2.26 GHz.) 
11. Xeon L5508- can only work with other L5508s since it's the only 38-watt TDP Xeon 5500. 





REFERENCES
http://www.tomshardware.com/forum/272899-28-different-processors-dual-socket-motherboard

Wednesday, March 7, 2012

Kernel panics

SkyHi @ Wednesday, March 07, 2012
start_kernel+0x379/0x380


Kernel panics are almost always hardware related:

- Flaky hardware (especially RAM)
- Motherboard/RAM mismatch (i.e. the wrong kind of RAM is installed)
- Overheating
- Immature driver support

Thursday, December 15, 2011

How should I run fsck on a Linux file system

SkyHi @ Thursday, December 15, 2011

Force fsck on boot using /forcefsck

By creating /forcefsck file you will force the Linux system (or rc scripts) to perform a full file system check.
Login as the root:
$ su -
Change directory to root (/) directory:
# cd /
Create a file called forcefsck:
# touch /forcefsckNow reboot the system:
# reboot


Scenario / Question:

I need to check file system for errors using fsck. Can I run fsck on a mounted file system ?

Solution / Answer:

Running fsck on a mounted file system can result in data corruption. The two options are:
1) Change the running state of the system to single user mode and unmount the file system
What if you need to run fsck on the root / file system ?
2) Boot the computer into Rescue Mode using the installation CD

1) Single User Mode and umount the file system

Issue command to change run level and umount the /home file system that is mounted on /dev/sda2
# init 1
# umount /home
Run fsck:
# fsck /dev/sda2

2) Rescue Mode using installation CD ( to run fsck on root /)

Insert the Installation CD into the drive and reboot your system:
# shutdown -r now
After booting from the Installation CD and presented with the installation command prompt type:
linux rescue nomount
Once you are at the system command prompt you need to run mknod. Because we started Rescue Mode with the “nomount” option, no file systems were initialized and no device files were created. If we try to run fsck on a file system it will fail. We need to use mknod to create the block or character special file.
To use mknod we need to know the Minor and Major numbers of the device.
# ls -l /dev/sda
8 0
# ls -l /dev/sda2
8 2
# mknod /dev/sda b 8 0
# mknod /dev/sda2 b 8 2
Run fsck and force the check and attempt to automatically repair:
-y — cause the fs-specific fsck to always attempt to fix any detected filesystem corruption automatically.
-f — force a check even if reported in a clean state
-v — Produce verbose output, including all file system-specific commands that are executed.
# fsck -yvf /dev/sda2
LVM Partitions
In order to be able to run fsck on lvm partitions we need to find the pv’s, vg’s, lv’s and activate them.
# lvm pvscan
# lvm vgscan
# lvm lvchange -ay /dev/VolGroup00/LogVol_home
# lvm lvscan

# fsck -yfv /dev/VolGroup00/LogVol_home
LUKS Partition
In order to be able to access an encrypted LUKS partition user cryptsetup.
cryptsetup luksOpen
– is the device path
– is the name of the unencrypted mount that can be accessed by /dev/mapper/
# cryptsetup luksOpen /dev/VolGroup00/LogVol_home home
# Ener LUKS passphrase for /dev/VolGroup00/LogVol_home
# fsck -yvf /dev/mapper/home


REFERENCES

Smartd Error: Currently unreadable (pending) sectors

SkyHi @ Thursday, December 15, 2011
I am encountering following error in /var/log/messages:




Aug 15 03:55:42 hostname smartd[2366]: Device: /dev/sda, 1 Currently unreadable (pending) sectors
Which cause the / partition to be mounted as read-only. The server is accessible anyway but you cant do anything much inside. Lets troubleshoot this.

Collecting Information/Troubleshooting

I see read-only filesystem mounted when creating a test file in /root directory:
$ touch /root/testfile
touch: cannot touch `/root/testfile': Read-only file system
What is SMART daemon (smartd)?
Self-Monitoring, Analysis and Reporting Technology (SMART) system built into many ATA-3 and later ATA, IDE and SCSI-3 hard drives. The purpose of SMART is to monitor the reliability of the hard drive and predict drive failures, and to carry out different types of drive self-tests. We will use smartctl command to help us find out what is wrong with the disk.
Lets check the overall health of disk /dev/sda:
$ smartctl -H /dev/sda
smartctl version 5.38 [i686-redhat-linux-gnu] Copyright (C) 2002-8 Bruce Allen
Home page is http://smartmontools.sourceforge.net/
 
 === START OF READ SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED
It passed. But it just general information only. We need to go deeper by do self-test to the disk:
$ smartctl -q errorsonly -H -l selftest -l error /dev/sda
ATA Error Count: 2
Error 2 occurred at disk power-on lifetime: 36795 hours (1533 days + 3 hours)
Error 1 occurred at disk power-on lifetime: 31542 hours (1314 days + 6 hours)
 
Num  Test_Description    Status                  Remaining  LifeTime(hours)  LBA_of_first_error
# 1  Short offline       Completed: read failure       60%     39255         -
When I Google up the error above, it seems like the hard disk might have hardware problem. FSCK only might not helping much since it only fix logical error in file system, not the hardware error.
Errors reported by SMARTD is related to power-on lifetime attributes which explain as below (reference):
Count of hours in power-on state. The raw value of this attribute shows total count of hours (or minutes, or seconds, depending on manufacturer) in power-on state. A decrease of this attribute value to the critical level (threshold) indicates a decrease of the MTBF (Mean Time Between Failures).
However, in reality, even if the MTBF value falls to zero, it does not mean that the MTBF resource is completely exhausted and the drive will not function normally.

Backup

Since the hard disk is in read-only mode, we better do backup before proceed with any problem solving process. In this case, SCP to another server is good idea because we cannot write to the local disk at this moment. For me, “home” partition is the most important folder need to be saved:
$ scp -r /home user1@remoteserver:/home/user1/home_backup

Problem Solving Process

1. Remount the / partition:
$ mount -n -o remount /
mount: block device /dev/sda2 is write-protected, mounting read-only
2. Run e2fsck command to check ext3 file system online:
$ e2fsck /dev/sda2
e2fsck 1.39 (29-May-2006)
/: recovering journal
Clearing orphaned inode 31672817 (uid=0, gid=0, mode=0100755, size=157913)
Clearing orphaned inode 31672803 (uid=0, gid=0, mode=0100755, size=3532999)
Clearing orphaned inode 31666625 (uid=0, gid=0, mode=0100755, size=150604)
Clearing orphaned inode 31666619 (uid=0, gid=0, mode=0100755, size=383872)
Clearing orphaned inode 27885882 (uid=0, gid=0, mode=0100755, size=1011760)
Clearing orphaned inode 31666617 (uid=0, gid=0, mode=0100755, size=1141532)
Clearing orphaned inode 31665420 (uid=0, gid=0, mode=0100755, size=398180)
Clearing orphaned inode 31665416 (uid=0, gid=0, mode=0100755, size=71852)
Clearing orphaned inode 31671503 (uid=0, gid=0, mode=0100755, size=1250176)
/: clean, 80179/38273024 files, 2990728/38258797 blocks
Try remounting again the partition like step 1 but same error occurred. Proceed to next step.
3. Run full file system check using FSCK via rescue environment:
$ fsck -f -y /dev/sda2
Even the box remount correctly after that, the smartd status still haunting me up. This has force me to make final decision as my next step.
4. To avoid any sudden breakdown (since the disk already run more than 1000 days), I decided to replace the hard disk and re-install the box. Its better for me to do this as part of my maintenance task so I will not worrying much about ‘urgent’ maintenance when it breakdown during weekend or sleep time!






REFERENCES
http://blog.secaserver.com/2011/08/smartd-error-1-unreadable-pending-sectors/