If I may suggest: When posting blog articles about code, use a variable-width theme instead of a narrow single column layout.
More importantly, the performance issue with CSV is not about the disk I/O, which is what System.IO.Pipelines is for. Instead, the problem is the GC pressure, which the article mentions, but does not solve! This is why it doesn't perform that much better than the traditional code.
For example, these two lines jumped out at me:
var value = Encoding.UTF8.GetString(line[..comaAt]);
record.DateOfJoining = Convert.ToDateTime(value);
The "var" is hiding a performance footgun: The author is converting a Span<char> to a temporary String, and then parsing that. The performant approach is to parse the Span directly with this call: https://docs.microsoft.com/en-us/dotnet/api/system.double.pa...Similarly, he's "efficiently" passing in each line as a Span<char>, but then throwing away all this effort by converting it into a string and then throwing this away each time:
if (Encoding.UTF8.GetString(line).Contains(ColumnHeaders))
{
return null;
}
Fundamentally, this article is just wrong, and its advice misses the mark entirely.The purpose of System.IO.Pipelines is to enable a different style of performance-oriented coding, but it doesn't enforce it. Just using it instead of System.IO.Stream won't magically make the code faster!
Typically instead of copying the input data into a data structure like "Employee[] employeeRecords", the idea is to "turn this around" and provide each record one at a time as references to Spans. This means that the original parsed data doesn't have to be broken up into millions of small GC allocations, and will stay in the original buffer.
E.g.: The parsing API should look something like this:
bool ProcessCsvLine( int lineNo, ReadOnlySpan<char> Name, ReadOnlySpan<char> email, DateTime dateOfJoining, decimal salary );
The parser would then call this function for each line, stopping if the function returns false. IF correctly implemented, the parser would require zero temporary allocations to do this.This enables fantastically efficient processing in scenarios where you might only be interested in a subset of the data. E.g.: If you want to get the total salary only and don't care about the emails, the emails won't be allocated on the GC heap at all, and are practically free in terms of processing. Similarly, if the consumer of the data needs to copy out the email column, then it doesn't also have to copy out the Name column. Last but not least, programs that only need the first 1,000 lines for a preview can simply stop parsing at that point by returning false.
Have you seen those programs that can only estimate the "type" of an input data column using the first few thousand rows or something trivial like that? This is why. They're too inefficient internally to use a larger set (or the entire file). If they instead used this Span-based style of programming, computing simple integers such as min/max length of each field requires only constant memory and could be done for millions of rows in milliseconds.