Repository navigation
MIMEType Performance #38
Description
Activity
The implementation is different in how they parse the
content-typeheader, being theutils.MIMETypeis slightly more iterative compared to how thefast-content-type-parseparses the same header.Among some other differences in how the util API is made and exported, we can use the
fast-content-type-parseas a baseline for the benchmark comparison while keeping the safeguards provided byutils.MIMEType.I'll dedicate time to this during the week 🙂
Reacted by KaKacc @nodejs/undici
Undici implements its own content-type parser, which like the nodejs one, is 100% spec compatible. As pointed out above it's unlikely that the other userland modules are totally compliant.
Giving some updates, I've been doing some optimizations at both ends, starting with Undici.
I was able to gain a ~20% increase by doing smaller optimizations while being spec-compliant.
As @KhafraDev pointed out, Undici is the implementation that it implements fully compliant accordingly to the spec, meanwhile, Node's is a little bit distant but in essence, is compliant as well.Results:
util#MIMEType x 1,345,116 ops/sec ±0.29% (97 runs sampled) undici#parseMIMEType(optimised) x 1,954,188 ops/sec ±0.27% (93 runs sampled) undici#parseMIMEType(original) x 1,621,463 ops/sec ±0.18% (98 runs sampled) fast-content-type-parse#parse x 5,750,335 ops/sec ±0.34% (98 runs sampled) fast-content-type-parse#safeParse x 5,752,805 ops/sec ±0.30% (96 runs sampled) Fastest is fast-content-type-parse#safeParse,fast-content-type-parse#parseOne important thing that might be good to think about is that being spec compliant might limit the room for optimizations, meaning that most likely any optimization applied at Undici or Node will be able to fully reach or surface
fast-content-type-parse.
Though, IMHO this is not bad, as being spec compliant is important.I'll be aiming to have another round of optimizations, now including
MIMETypeas well to see how much I can push it at the first iteration.Out of Scope
Just as an idea, it might be good to push the boundaries ofutils#MIMETypeto reach Undici's implementation, and as they serve for kinda the same purposes, Undici maybe can make use of it internally. Not so sure if a good idea, but want to just raise it in case it can be helpful 🙂Thanks @metcoder95 for the awesome work. I didnt get any chance to look over the implementation but curious: would it be faster if we implemented the same exact code using C++?
Reacted by Carlos FuentesIt's not a performance concern in undici - it's only used for parsing the content-type in
data:urls.would it be faster if we implemented the same exact code using C++?
The
MIME utilis original written in C++ but turns out change into JavaScript to reduce maintenance burden.spec-compliant
About specification compliant can be either follow the exact step or optimize the step with output compliant.
The former will have little room for improvement, but the latter will have a great room for it.content-typeandfast-content-type-parseis a example for the latter in providingRFC 7231spec-compliant output.
MIMETypeandundiciis the former in providingMIMETypespec-compliant implementation.Thanks @metcoder95 for the awesome work. I didn't get any chance to look over the implementation but curious: would it be faster if we implemented the same exact code using C++?
Sure thing!
Toward the implementation (or re-implementation) into C++, I would say yes (without adding possible overhead on the communication low-JS), but it translates into the costs that @climba03003 pointed out.The former will have little room for improvement, but the latter will have a great room for it
+1
I think it will be a matter of trade-offs as always. Depending on what's the exact goal that wants to be achieved (fulfill one or another spec).Note: Undici uses the
WHATWGspec for MIMEType representation and parsing.I'll continue with the research and open subsequent PRs 🙂
Opened PR as discussed, from the insights gathered of it, I'll proceed with the
MIMETypeexperiments 🙂Reference: Undici#1871
The PR is merged, I'll use the insights from the assessment to continue with
MIMETypeimprovements on Node 🙂Reacted by Rafael GonzagaAfter having played with the
MIMETypeclass implementation in Node and using different approaches, I was able to come up with the following notes:- I think the most evident,
MIMETypeis fully attached to the spec(link) in its own way - The biggest pain for both implementations, but as its most for
MIMETypeclass, is the parsing of the parameters of the MIME string
Baseline:
-> With string'application/json; charset=utf-8'util#MIMEType x 1,349,103 ops/sec ±0.16% (92 runs sampled) undici#parseMIMEType x 2,349,808 ops/sec ±0.14% (99 runs sampled) fast-content-type-parse#parse x 5,663,227 ops/sec ±0.58% (98 runs sampled) fast-content-type-parse#safeParse x 5,719,013 ops/sec ±0.44% (95 runs sampled)-> With string
'application/json'util#MIMEType x 2,463,796 ops/sec ±0.32% (96 runs sampled) undici#parseMIMEType x 4,467,917 ops/sec ±0.10% (100 runs sampled) fast-content-type-parse#parse x 22,576,597 ops/sec ±0.12% (99 runs sampled) fast-content-type-parse#safeParse x 22,565,488 ops/sec ±0.12% (102 runs sampled)Scenario - 1
Over the same string
'application/json; charset=utf-8', and with improvements to Node.js implementation, the following results are shown:util#MIMEType x 1,500,644 ops/sec ±0.39% (94 runs sampled) undici#parseMIMEType x 2,337,932 ops/sec ±0.13% (97 runs sampled) fast-content-type-parse#parse x 5,790,442 ops/sec ±0.13% (99 runs sampled) fast-content-type-parse#safeParse x 5,793,647 ops/sec ±0.13% (97 runs sampled)It showed ~11% improvement.
The improvements were mostly done at the Regex usage by replacingmatchto test for boolean scenarios.Scenario - 2
Over the same string
'application/json', and similar improvements as first scenario to Node.js implementation, the following results are shown:util#MIMEType x 2,834,890 ops/sec ±0.13% (94 runs sampled) undici#parseMIMEType x 5,199,808 ops/sec ±0.17% (102 runs sampled) fast-content-type-parse#parse x 23,378,962 ops/sec ±0.25% (95 runs sampled) fast-content-type-parse#safeParse x 23,192,733 ops/sec ±0.27% (91 runs sampled)Similar improvements in the results.
The biggest difference is how Undici parses the parameters of the MIME string, as Node.js uses more a direct approach by doing constant usage of
slice,search, andindexOfof the String prototype to iterate, cut, and obtain the different parameters of the MIME string. Meanwhile, Undici does it by a more pragmatic approach not exhausting the String prototype usage but rather a more iterative approach with a callback acting as a predicate to stop the iteration.My take is that an iterative approach will make things easier and improve the performance, as will also reduce the usage of Primordials which are well known for adding slight overhead.
I'll prepare a PR next week with the proposal so this is more graphic and will make it easier to provide feedback and discuss 🙂
- I think the most evident,
Hey! To give some heads-up about the progress of the work 🙂
I recently opened node#46607 with the initial changes in pro of trying to improve the perf of theMIMETypeclass.Note: Just so you know, it is important to remark that it is in a draft state as cleanup and ensure that tests are required.
After several evaluations I was able to come up with the following results:
Compared against Node v18
confidence improvement accuracy (*) (**) (***) util/mime-parser.js n=100000 strings='application/json; charset="utf-8"' *** 12.07 % ±2.78% ±3.72% ±4.87% util/mime-parser.js n=100000 strings='text/html ;charset=gbk' *** 8.13 % ±2.18% ±2.91% ±3.79% util/mime-parser.js n=100000 strings='text/html; charset=gbk' *** 4.20 % ±2.35% ±3.13% ±4.08% util/mime-parser.js n=100000 strings='text/html;charset= "gbk"' *** 10.10 % ±2.02% ±2.69% ±3.50% util/mime-parser.js n=100000 strings='text/html;charset=GBK' *** 11.53 % ±2.61% ±3.48% ±4.57% util/mime-parser.js n=100000 strings='text/html;charset=gbk' *** 8.37 % ±2.23% ±2.96% ±3.86% util/mime-parser.js n=100000 strings='text/html;charset=gbk;charset=windows-1255' *** 14.77 % ±1.66% ±2.21% ±2.88% util/mime-parser.js n=100000 strings='text/html;x=(;charset=gbk' *** 12.14 % ±3.32% ±4.41% ±5.75%
The improvements were between ~8-12% on average, which is good but still not sure if enough to call it successful.
On ops/sec, the results were unprecise as with the previous baseline of:* util#MIMEType x 1,349,103 ops/sec ±0.16% (92 runs sampled) * undici#parseMIMEType x 2,349,808 ops/sec ±0.14% (99 runs sampled) * undici#parseMIMEType(original) x 1,520,777 ops/sec ±0.26% (100 runs sampled) * fast-content-type-parse#parse x 5,663,227 ops/sec ±0.58% (98 runs sampled) * fast-content-type-parse#safeParse x 5,719,013 ops/sec ±0.44% (95 runs sampled)the new baseline varied to:
util#MIMEType x 2,075,831 ops/sec ±0.14% (197 runs sampled) undici#parseMIMEType x 2,729,673 ops/sec ±0.15% (194 runs sampled) fast-content-type-parse#parse x 5,751,725 ops/sec ±0.13% (195 runs sampled) fast-content-type-parse#safeParse x 5,766,105 ops/sec ±0.10% (195 runs sampled) Fastest is fast-content-type-parse#safeParseAlmost ~50%, which seems to be quite out compared to the results of Node benchmarks. I wanted to put them on the table but wouldn't ensure they are precise as they were varying by ~20% on several iterations.
Looking forward to your feedback and seeing what points can be improved that I could missed 🙂
Reacted by Marvin Hagemeister and Benjamin Gruenbaumreally good! good job @metcoder95
Reacted by Carlos FuentesHi guys! Is there a way to move nodejs/node#46607 forward? 🙂
Hmm. Didnt know this thread exists. Didnt think that fast-content-type-parse is not spec compliant. Can somebody point to a test suite for spec compliance?
thx
Hmm. Didnt know this thread exists. Didnt think that fast-content-type-parse is not spec compliant. Can somebody point to a test suite for spec compliance?
BTW, I believe that when referring to
fast-content-type-parseis not compliant, it is about to MIMESniff standard from the WHATWG, not the RFC-1341 which I believe the library in fact respects.LOL, I forgot about this issue and created #120
Dang.
Here my post from the other issue:
Maybe we should check how we can improve this overall?
I think that MIMEType can be improved significantly.
I created rn a PR regarding lazily parsing the MimeParams.
Also i could improve toASCIILower with this snippet:
const ASCII_LOOKUP = new Array(127).fill(0).map((v, i) => { if (i >= 65 && i <= 90) { return StringFromCharCode(i + 32); } return ''; }) function toASCIILower(str) { let result = ''; for (let i = 0; i < str.length; ++i) { const code = StringPrototypeCharCodeAt(str, i); if (code > 90 || code < 65) { result += str[i]; } else { result += ASCII_LOOKUP[code]; } } return result; }
Also why do we use SafeStringPrototypeSearch in encode? Cant we just use RegexPrototypeExec? Cant we manually inline the encode function into toString?
Is there a faster way to iterate the SafeMap? Why do we use a SafeMap and not a NullObject (because of the generator functions?
Can we avoid the generator fns? Does it make sense to avoid the generator fns?
Can we optimize parseTypeAndSubtype? Maybe also lazily parse the values? We need to throw errors in special cases, but other than that, we just have to store the string.
Why do we call in MIMEType toString
FunctionPrototypeCall(MIMEParamsStringify, this.#parameters);and not
this.#parameters.toString()?Questions over questions...
No worries, happy to see there is progress on this front. I really want to continue my work over nodejs/node#46607, maybe now that @Uzlopak is a member, we can make it move forward with reviews and adjustments? 👀
Also why do we use SafeStringPrototypeSearch in encode? Cant we just use RegexPrototypeExec?
SafeStringPrototypeSearchis slightly faster thanRegexPrototypeExec: nodejs/node#46607 (comment)Is there a faster way to iterate the SafeMap? Why do we use a SafeMap and not a NullObject (because of the generator functions?
I think it was mostly commodity, didn't find counters while working on the PR of using a plain object. Can be sealed with
NullObjectif we want. The spec does not define a requirement for the params to be a properMapinstance, but ratherMap-like data structure. We can even store them as we like meanwhile we can parse them as strings correctly.Can we avoid the generator fns? Does it make sense to avoid the generator fns?
Which generators?
Can we optimize parseTypeAndSubtype? Maybe also lazily parse the values? We need to throw errors in special cases, but other than that, we just have to store the string.
I think we can lazy it at much until is required to provide the type and subtype back; but that won't help much as usually, it is within the hot path so it cannot be delayed too much the parsing.
Why do we call in MIMEType toString FunctionPrototypeCall(MIMEParamsStringify, this.#parameters); and not
this.#parameters.toString()?Sealing the call I'd say?
nodejs/node#49889 got merged.
@metcoder95
Should we focus now on your PR?Sure thing, let me put it up-to-date and run the benchmarks agreed. Didn't had the time yet to take another look, I'd try to do it this week 👍
@metcoder95 Have you benchmarked it after your recent changes? Is it still a bottleneck?
Noup, I haven't (I actually missed this comment).
I'll have some time this weekend to try it out, let me give it a try
I know that
MIMETypeis recently added and underExperimentalflag. But I want to draw a attention on it's performance.It is not the slowest compare to the userland module, but it can do better.
I am considering replacing
MIMETypefor theContent-Typeparsing infastifybut the performance is the main drawback.The main reason behind using
MIMETypeis the.essenceproperty did a great job in unifying theBrowserandServerin terms ofContent-Typeguessing / matching and prevent the similar security issue like GHSA-3fjj-p79j-c9hhRefs fastify/fastify#4502