Storage spaces direct is great, but every once and a while a S2D storage job will get a stuck and just sit there in a suspended state. This usually happens after a reboot of one of the nodes in the cluster.
What you don't want to do is take a different node out of the cluster while a storage job is stuck and while there are degraded virtual disks.
You should make a habit out of checking the jobs and the virtual disk status before changing node membership. You can do this easily with the Get-StorageJob and Get-VirtualDisk commandlets. Alternatively, you could use the script I wrote to continually update the status of both the S2D storage jobs and the virtual disk status.
So what does one do if a storage job is stuck? There are two commandlets that I've found will fix this. The first is Optimize-StoragePool. The second is Repair-VirtualDisk. Start with Optimize-StoragePool and if that doesn't work then move on to Repair-VirtualDisk. Here is how you use them:
Get-StoragePool <storage pool friendly name> | Optimize-StoragePool
Example: Get-StoragePool s2d* | Optimize-StoragePool
Get-VirtualDisk <virtual disk friendly name> | Repair-VirtualDisk
Example: Get-VirtualDisk vd01 | Repair-VirtualDisk
Usually optimizing the storage pool takes care of the hung storage job and fixed the degraded virtual disk but if not target the disk directly.
If neither of those work, give Repair-ClusterStorageSpacesDirect / Repair-ClusterS2D a try. I haven't tried this one yet but it looks like it could help.
Update: I tried Repair-ClusterS2D. It does not appear to help with this scenario. There is limited documentation on it but it looks like it's something you use if a virtual disk gets disconnected or something.
Update: Run Get-PhysicalDisk. If any of them say they're in maintenance mode, this could be the cause of your degraded disks and your stuck jobs. This seems to happen when you pause and resume a node to close together. To take the disks our of maintenance mode run the following:
Get-PhysicalDisk | Where-Object { $_.OperationalStatus -eq "In Maintenance Mode" } | Disable-StorageMaintenanceMode
Another Update: If a disk becomes dettached, try this.
Tuesday, June 20, 2017
Thursday, February 23, 2017
S2D Continually Refresh Job and Disk Status
In storage spaces direct you can run Get-StorageJob to see the progress of rebuilds/resyncs. The following powershell snippet allows you to continually refresh the status of the rebuild operation so that you know when things are back to normal.
function RefreshStorageJobStatus () { while($true) { Get-VirtualDisk | ft; Write-Host "-----------"; Get-StorageJob;Start-Sleep -s 1;Clear-Host; } }
Enter the above in powershell on one line. Then enter "RefreshStorageJobStatus" to start the script. The output should look similar to the following and refresh every second:
Name IsBackgroundTask ElapsedTime JobState PercentComplete BytesProcessed BytesTotal
---- ---------------- ----------- -------- --------------- -------------- ----------
Repair True 00:00:13 Suspended 0 0 7784628224
Repair True 00:00:06 Suspended 0 0 7784628224
FriendlyName ResiliencySettingName OperationalStatus HealthStatus IsManualAttach Size
------------ --------------------- ----------------- ------------ -------------- ----
vd01 OK Healthy True 1 TB
vd03 Degraded Warning True 1 TB
vd02 Degraded Warning True 1 TB
vd04 OK Healthy True 1 TB
You can press ctrl-c to stop the execution.
Update 8/16/2018: Here is an updated RefreshStorageJobStatus function that shows the bytes processed and bytes remaining in gigabytes instead of bytes:
function RefreshStorageJobStatus () { while($true) { Get-VirtualDisk | ft; Write-Host "-----------"; Get-StorageJob | Select Name,IsBackgroundTask,ElapsedTime,JobState,PercentComplete,@{label=”BytesProcessed (GB)”;expression={$_.BytesProcessed/1GB}},@{label=”Total Size (GB)”;expression={$_.BytesTotal/1GB}} | ft;Start-Sleep -s 1;Clear-Host; } }
Run the same as before, enter "RefreshStorageJobStatus" to start the script. Output looks like this:
FriendlyName ResiliencySettingName OperationalStatus HealthStatus IsManualAttach Size
------------ --------------------- ----------------- ------------ -------------- ----
test2dfsb Incomplete Warning True 3.5 TB
vd02b {Degraded, Incomplete} Warning True 1 TB
vd01b {Degraded, Incomplete} Warning True 1 TB
-----------
Name IsBackgroundTask ElapsedTime JobState PercentComplete BytesProcessed (GB) Total Size (GB)
---- ---------------- ----------- -------- --------------- ------------------- ---------------
Repair True 00:00:41 Suspended 0 0 70.25
Repair True 00:00:01 Suspended 0 0 122.25
function RefreshStorageJobStatus () { while($true) { Get-VirtualDisk | ft; Write-Host "-----------"; Get-StorageJob;Start-Sleep -s 1;Clear-Host; } }
Enter the above in powershell on one line. Then enter "RefreshStorageJobStatus" to start the script. The output should look similar to the following and refresh every second:
Name IsBackgroundTask ElapsedTime JobState PercentComplete BytesProcessed BytesTotal
---- ---------------- ----------- -------- --------------- -------------- ----------
Repair True 00:00:13 Suspended 0 0 7784628224
Repair True 00:00:06 Suspended 0 0 7784628224
FriendlyName ResiliencySettingName OperationalStatus HealthStatus IsManualAttach Size
------------ --------------------- ----------------- ------------ -------------- ----
vd01 OK Healthy True 1 TB
vd03 Degraded Warning True 1 TB
vd02 Degraded Warning True 1 TB
vd04 OK Healthy True 1 TB
You can press ctrl-c to stop the execution.
Update 8/16/2018: Here is an updated RefreshStorageJobStatus function that shows the bytes processed and bytes remaining in gigabytes instead of bytes:
function RefreshStorageJobStatus () { while($true) { Get-VirtualDisk | ft; Write-Host "-----------"; Get-StorageJob | Select Name,IsBackgroundTask,ElapsedTime,JobState,PercentComplete,@{label=”BytesProcessed (GB)”;expression={$_.BytesProcessed/1GB}},@{label=”Total Size (GB)”;expression={$_.BytesTotal/1GB}} | ft;Start-Sleep -s 1;Clear-Host; } }
Run the same as before, enter "RefreshStorageJobStatus" to start the script. Output looks like this:
FriendlyName ResiliencySettingName OperationalStatus HealthStatus IsManualAttach Size
------------ --------------------- ----------------- ------------ -------------- ----
test2dfsb Incomplete Warning True 3.5 TB
vd02b {Degraded, Incomplete} Warning True 1 TB
vd01b {Degraded, Incomplete} Warning True 1 TB
-----------
Name IsBackgroundTask ElapsedTime JobState PercentComplete BytesProcessed (GB) Total Size (GB)
---- ---------------- ----------- -------- --------------- ------------------- ---------------
Repair True 00:00:41 Suspended 0 0 70.25
Repair True 00:00:01 Suspended 0 0 122.25
Monday, February 13, 2017
AD-less S2D cluster bootstrapping
AD-less S2D cluster bootstrapping - Domain Controller VM on Hyper-converged Storage Spaces Direct
Is it a supported scenario to run a AD domain controller in a VM on a hyper-converged S2D cluster? We're looking to deploy a 4-node hyper-converged S2D cluster at a remote site. We would like to run the domain controller for the site on the cluster so we don't need to purchase a 5th server. Will the S2D cluster be able to boot if the network links to the site are down (meaning other domain controllers are not accessible)? I know WS2012 allowed for AD-less cluster bootstrapping but will the underlying mechanics uses for storage access in S2D in WS2016 work without AD? Is this a supported scenario? AD-less S2D cluster bootstrapping?
I asked this question in the Microsoft forums. I did not get a definitive answer from anyone. So I set it up and tested it and it appears to work. I don't know if it's officially supported or not but it does work. The S2D virtual disks and volumes comes up with out a domain controller. At which point you can start the domain controller VM if it did not start automatically. I didn't dig into things, but I have a feeling it's using NTLM authentication and would likely fail if your domain requires Kerberos?
Friday, January 29, 2016
3 Node Storage Spaces Direct Cluster Works!!!
I
went through the following URL https://technet.microsoft.com/en-us/library/mt126109.aspx
but instead of creating a 4 node Storage Spaces Direct cluster, I decided to
try and see if a 3 node cluster would work. Microsoft documentation says that they will only
support Storage Spaces Direct with 4 servers but I thought it can't hurt to
try a 3 node... and it worked!!
Storage Spaces and Latent Sector Errors / Unrecoverable Read Errors
I emailed S2D_Feedback@microsoft.com to ask about storage spaces direct and how it handles Latent Sector Errors (LSE), otherwise known as Unrecoverable Read Errors. Here is the email I sent:
My company is in the process of evaluating different options
for upgrading our production server environment. I’m tasked with finding a
solution that meets our needs and is within our budget.
I’m trying to compare and contrast storage spaces direct
with storage spaces utilizing JBOD enclosures. Data resiliency, integrity and
availability are paramount. So I’m primary looking at both of these technologies
from that perspective. Thus, if we go the JBOD route, we’re looking at
implementing 3 enclosures and utilizing the enclosure awareness of storage
spaces. This solution has existed longer then storage spaces direct and I would
think has been tested more thoroughly. I like the scalability and elegance of
storage space direct though. From a conceptual overview and a hardware setup
perspective it just seems easier to grasp and it seems like a better solution.
My question is, how do both of these setups handle
unrecoverable read errors/latent sector errors? Does one solution handle them
better than the other?
There are horror stories about hardware RAID controllers
evicting drives because of URE/LSE and then during RAID rebuilds encountering
additional UREs/LSEs and bricking the storage. This is more worrisome when SATA
disks are used (due to UREs/LSEs occurring more often and sooner with SATA
disks compared to SAS disks.) How does storage spaces/S2D differ in this
regard? I know one of the selling points of S2D is the use of SATA disks. I’m
curious as to how this problem has been addressed since SATA disks are being
promoted. What happens if there is a URE/LSE in end user data? What happens if
there is a URE/LSE in the metadata used by storage spaces/S2D or the underlying
file system?
Here is the response I received:
Both Spaces direct and
Shared Spaces (with JBOD) both rely on the same software raid
implementation, difference is in the connectivity. Software raid implementation
does not throw away the entire drive on failure, we trigger activity to move
the data out of the drive while keeping the copy till data is moved (if we have
copies available). On Write failure we try to move the impacted range
right away while background activity is moving the untouched data out of the
disk, some of the disks fail to write but they can continue to support
reads in which case the data on those drives can still be used to serve user
requests. Until the data on the failed drive is rebuilt on spare capacity
the drive is not removed, user can still force but not automated. On URE
- we trigger rebuilt to recover lost copy, this is triggered both when reads
errors detected while satisfying user error or by back ground scrub process.
Back ground scrub process detects URE by validating sector level checksum
across copies and validating.
So it would appear that if you utilize storage spaces you don't have to worry about a LSE/URE taking out a drive and then a subsequent LSE/URE taking out another drive, thus taking down your array.
Tuesday, July 14, 2015
IIS Application Initialization Quick Reference
- Need to install application initialization module, if IIS 7.5 it's a separate download or WPI, included in IIS 8
- Set startMode to AlwaysRunning on application pool
- Open configuration editor in IIS on server root
- Select system.applicationHost/applicationPools
- Click edit items
- Find app pool, select and then change startMode in lower pane to alwaysRunning
- Hit Apply This changes C:\Windows\System32\inetsrv\config\applicationHost.config
- Set applicationDefault preloadEnabled on site
- Open configuration editor in IIS on server root
- Select system.applicationHost/sites
- Click edit items
- Find site and select
- Expand applicationDefaults in lower pane
- change preloadEnabled to true
- Hit Apply This changes C:\Windows\System32\inetsrv\config\applicationHost.config
- Note: This only changes the defaults for new apps, you may need to change the existing ones yet. You may not need step 3 if apps already exist…
- Set preloadEnabled on site
- Open C:\Windows\System32\inetsrv\config\applicationHost.config
- Find site you're looking for <site name="www.domain.com" id=…
- Add preloadEnabled="true" to first application element tag
- You should see a <applicationDefaults preloadEnabled="true" /> under the site element if you performed step 3
- Restart IIS
- Set initalizationPage on site and doAppInitAfterRestart
- Open configuration editor in IIS on desired site
- Select system.webServer/applicationInitialization
- Change from to ApplicationHost.config if you want to create a location element in C:\Windows\System32\inetsrv\config\applicationHost.config, otherwise the change will go in the web.config
- Set doAppInitAfterRestart to true
- Click edit items
- Click add
- Enter path for initializationPage ( /folder/page?param=something ), leave hostname blank (I think…)
- Click apply
CHANGE IDLE TIME_OUT
IN ADVANCED SETTINGS ON APPPOOL TO 0
Monday, July 13, 2015
DPM 2010 Slow when Selecting Roverypoint to Recover
It took almost an hour to select a date and time to recover in DPM 2010. It appeared the this was because of high cpu usage by SQL. The query that seemed to be responsible for this was:
SELECT Path, FileSpec, IsRecursive
FROM tbl_RM_RecoverableObjectFileSpec
WHERE RecoverableObjectId = @RecoverableObjectId AND
DatasetId = @DatasetId and
iSgcED = 0
I didn't dig into things too much, but it appeared as though it was running this query for every single recovery point for the item selected and it was doing a clustered index scan for each recovery point. I created the following statistic and covering nonclustered index in the DPM db:
CREATE STATISTICS [_dta_stat_1042102753_9_2_3] ON [dbo].[tbl_RM_RecoverableObjectFileSpec]([IsGCed], [RecoverableObjectId], [DatasetId])
CREATE NONCLUSTERED INDEX [_dta_index_tbl_RM_RecoverableObjectFileSpec_7_1042102753__K2_K3_K9_5_6] ON [dbo].[tbl_RM_RecoverableObjectFileSpec]
(
[RecoverableObjectId] ASC,
[DatasetId] ASC,
[IsGCed] ASC
)
INCLUDE ( [FileSpec],
[IsRecursive]) WITH (SORT_IN_TEMPDB = OFF, IGNORE_DUP_KEY = OFF, DROP_EXISTING = OFF, ONLINE = OFF) ON [PRIMARY]
Now the above query changed from a clustered index scan to a index seek and key lookup. The time it took to select a recovery point went from about an hour down to a minute.
SELECT Path, FileSpec, IsRecursive
FROM tbl_RM_RecoverableObjectFileSpec
WHERE RecoverableObjectId = @RecoverableObjectId AND
DatasetId = @DatasetId and
iSgcED = 0
I didn't dig into things too much, but it appeared as though it was running this query for every single recovery point for the item selected and it was doing a clustered index scan for each recovery point. I created the following statistic and covering nonclustered index in the DPM db:
CREATE STATISTICS [_dta_stat_1042102753_9_2_3] ON [dbo].[tbl_RM_RecoverableObjectFileSpec]([IsGCed], [RecoverableObjectId], [DatasetId])
CREATE NONCLUSTERED INDEX [_dta_index_tbl_RM_RecoverableObjectFileSpec_7_1042102753__K2_K3_K9_5_6] ON [dbo].[tbl_RM_RecoverableObjectFileSpec]
(
[RecoverableObjectId] ASC,
[DatasetId] ASC,
[IsGCed] ASC
)
INCLUDE ( [FileSpec],
[IsRecursive]) WITH (SORT_IN_TEMPDB = OFF, IGNORE_DUP_KEY = OFF, DROP_EXISTING = OFF, ONLINE = OFF) ON [PRIMARY]
Now the above query changed from a clustered index scan to a index seek and key lookup. The time it took to select a recovery point went from about an hour down to a minute.
Subscribe to:
Posts (Atom)