Signing TLS handshakes inside a TPM (bschaatsbergen.com)

32 pointsby bschaatsbergen13 hours ago41 comments

duk3luk3 7 hours ago

Sounds interesting; too bad all we get is text made up by an LLM rather than any of the author's insights.

abound 7 hours ago

Yeah I was interested for the first few paragraphs, then all of a sudden I get hit with two "genuinely"s and a

> That’s the third property, and it’s the one that decides this.

and I gave up at that point.

bschaatsbergen 4 hours ago

Author here. All of it is mine, the library (https://github.com/bschaatsbergen/go-tpm-tls) and the benchmarks (https://github.com/bschaatsbergen/go-tpm-tls-bench) and the working notes. English isn't my first language, so I edit a lot, and I can see how that comes out flat. I've been over it once more; hopefully it reads better now. Thanks for saying so rather than just closing the tab.

nilsherzig 2 hours ago

Still flags as 100% LLM written https://www.pangram.com/history/cc9b6131-c772-4b3b-9a5d-8e81...

bob1029 2 hours ago

I suppose the value of this depends on your threat model.

The TPM will give you stronger assurance that a machine owns a key, but it's likely that a dedicated HSM would be much harder to extract the key material from.

TPM being inside the machine is a double edged sword. On one hand it makes attestation feasible, but on the other you now have the security black box inside the same physical domain as the machine that uses it. Risk of side channel extraction goes up dramatically when these systems coexist. It's a lot harder to instrument an HSM across the network.

mjg59 2 hours ago

A dedicated HSM will give you stronger trust that the private key material can't be extracted, but there's no real way to bind an HSM to a specific client and that's a very easy thing to do in the vTPM case.

bob1029 2 hours ago

How often do we need to bind a specific machine to a specific key in the case of TLS?

In every case of TLS I've seen we are concerned with organizational identity, not machine identity. This effectively extends to client certificates in cases like B2B & vendor integration.

In both scenarios you would definitely want to use an HSM style solution.

Protecting the HSM from inappropriate use (proving you are allowed to sign using the keys within) is a problem orthogonal to protecting the key material. In cloud HSM applications, you often combine the cloud vendors managed identity solution and HSM policies to effectively bind a set of machines to a set of keys.

mjg59 an hour ago

In the given case - you want to bind communication to a given confidential compute instance, which means you want to be able to ensure that the communication is coming from within the confidential compute instance, which means you want to be able to prove that the private key is only accessible from within that instance. An HSM buys you nothing more there.

jon-wood 11 minutes ago

Any time you've got hardware and want to attest that it hasn't been tampered with before allowing it to interact with something like an API endpoint.

At work we deploy industrial IoT gateways, these are very much not end-user devices. We are actually concerned about the device's identity, and more specifically about being able to attest that the device is in fact the one we thought it was and it hasn't been tampered with. By putting the key for TLS client certificate in the device's TPM, locked behind attestation that what's been booted is what we expected to boot, we can have a reasonable degree of confidence that we're communicating with the device we thought we were rather than just someone who managed to copy the private key off disk.

flippingheck 2 hours ago

TPM is just a spec, it isn't necessarily a black box.

ARM TrustZone, for example, can run this OSS TPM: https://github.com/OP-TEE/optee_ftpm

I expect there are equivalents for Intel/AMD.

KaiserPro 3 hours ago

For a company I work for I needed to ship a machine through unknown channels and have some confidence that it wasn't fiddled with.

my threat model was reasonably technical engineer swapping drives for some reason, or someone claiming that the machine is "different". (no nation state shit)

after the machine was imaged, it would connect to our central config server, get its hostname and exchange keys which would be embedded in the TPM.

once the machine is shipped and booted, it'll check in and sign a challenge. any kind of action on the central API could have a challenge. Each machine is attested at least once an hour.

I'm not sure how "secure" it all is, but it seems to work.

mjg59 2 hours ago

I helped design the attestation framework for https://docs.cloud.google.com/transfer-appliance/docs/4.0/re... - the goal was to ensure that the device you're about to copy a bunch of sensitive information onto is actually the device you were shipped and is running the expected software. This is definitely used in the real world.

thomashabets2 3 hours ago

Looks like speeds have picked up since I last looked at this, when a signature in TPM took 0.7s and no concurrent capacity.

https://blog.habets.se/2012/02/Benchmarking-TPM-backed-SSL.h...

https://blog.habets.se/2012/02/TPM-backed-SSL.html

Well, it's been over 14 years so I should hope so.

mjg59 2 hours ago

The benchmarks are from GCP, where the vTPM is implemented in the hypervisor rather than on something that's plausibly an 8051[1]. Doing this on actual client hardware is going to be a bunch slower.

[1] Typically ARM these days, but most system vendors aren't picking TPM vendors based on performance

bschaatsbergen 2 hours ago

What mjg59 says. The benchmarks are against a vTPM, that was what I had access to, and it's the environment I'm implementing the RATS side in.

Worth adding that not every outbound connection needs to go through the TPM (IMO). It's for the handful of services where the machine-identity actually matters, a secret store, or an HSM releasing key material onto an attested confidential VM, in my case.

mjg59 2 hours ago

You didn't really go into actually verifying the machine identity - obviously if you have a trusted mechanism to do that in advance then that's easy enough, but otherwise you'd want something like https://github.com/google/go-attestation and then to use control plane APIs to identify the vTPM EK to tie the TPM to the VM.

flippingheck 2 hours ago

Yeah, a typical TPM chip has much lower throughput than OP.

Not suitable for servers, since it's such an easy DoS vector.

ivlad an hour ago

I might be missing something: is this any conceptually different from using PKCS11 provider for TPM in OpenSSL?

Also, with real TPM, the key could be locked to a specific configuration register value, which makes less sense for VMs. “Quote” is mentioned and I guess author means that, but did not elaborate further.

flippingheck 10 minutes ago

> is this any conceptually different from using PKCS11 provider for TPM in OpenSSL?

PKCS11 doesn't allow you to attest that the key is resident in the PKCS11 provider, which as you say, the author alludes to, but doesn't cover.

> with real TPM, the key could be locked to a specific configuration register value, which makes less sense for VMs.

A vTPM is as real as a physical TPM chip.

The question is which TPM endorsement certificate CAs you are willing to trust.

For some that might the manufacturer of TPM chips, for others it might be their VM provider. (For some, none: for some both!)

Trusting their VM provider isn't so crazy if the VM provider is able to influence the guest code anyway.

ted_dunning 3 hours ago

The link between attestation and the key is nicely made with TAS. TAS gives you a cert and Spiffe then requires a cert like that to give a SVID that you use as a certificate for mTLS.

This means that the root of trust threads through software (TAS) that verified that your attestation evidence matches the live policy. This works with no changes to Spiffe.

This doesn't really meet your requirements to keep the key out of memory since the resulting SVID lasts for several minutes in memory, but it does meet most people's needs.

https://github.com/TEE-Attestation/tas

bschaatsbergen 2 hours ago

Thanks for sharing Ted!

mjg59 2 hours ago

This feels like a somewhat odd design choice - you have a TEE, most TEEs (outside TPMs) are fast so there's little overhead in pushing your signing through there, why bother with short-lived credentials instead of just attesting to private key material ownership and having that be what the SPIFFE cert is issued to? Bearer token SVIDs are an awful thing that we should be getting as far away from as possible.

yusufmotiwala 3 hours ago

Isn't this a well-discussed issue already, and not specific to TPM?

We faced a similar issue (we use OpenSSL). OpenSSL does have OPENSSL_secure_malloc() which prevents sensitive memory from being dumped. However, the problem is that not all paths use the secure allocator. For example, this issue: https://github.com/openssl/openssl/issues/27603

Not sure if this has changed in OpenSSL 4.x, but it is certainly something desirable.

lesspassiveobse an hour ago

I wonder, at what point will it be cheaper to kidnap and ransom those remote attestation engineers' families for key material than to work around those schemes with technical measures. Keeping in mind that people set up bot farms with physical phones just for attestation keys, it seems like tightening it all too much will just shift the balance towards the $5 wrench approach...

bschaatsbergen an hour ago

Joke's on them, I don't have the key either.

lesspassiveobse an hour ago

Well I guess AMD/Intel/Qualcomm/Infineon/Google people are at the greatest risk, with places like TSMC also in play. Infineon TPMs and smartcards in particular had so many flaws that I wonder if it already happened. Also, note that even if you don't have the keys you may still control the implementation, and potentially introduce flaws. Keep safe.

ram_rattle 6 hours ago

Nothing new here, attested TLS was being discussed in IETF for quiet sometime right?

https://datatracker.ietf.org/doc/draft-fossati-tls-attestati... https://www.youtube.com/watch?v=MF9AwkMJOlw

bschaatsbergen 4 hours ago

That's right, I'm learning in public here. That draft is a different direction though, they change the handshake: new TLS extensions carry the evidence, and the far end appraises the platform during the connection.

What I'm doing changes nothing on the wire, the verifying side has no idea a TPM is involved. In RATS (https://www.rfc-editor.org/rfc/rfc9334.html) we prove a machine is sound by measuring it and appraising the evidence. But after attestation the usual thing is to hand the machine a short-lived identity saying it is attested, and when that machine then authenticates over mTLS to something like an HSM, the thing that gives that machine its identity is a private key in a file. That bothered me. What I want is to tie the key in the TPM to the evidence of the confidential VM at issuance time, and let that be the identity the machine carries afterwards. Working notes while implementing RFC 9334.

ram_rattle 4 hours ago

Never mind, I had no idea who you where, looked you up, please take a bow, apologies if the comment came out rude, more power to your work and agree learning in public and publishing more will what will make this idea better.

Hat Tip!

madduci 4 hours ago

Exactly, you could do this also with the Microsoft Cryptographic Provider long time ago, which is the basic Provider called by the go-tpm library, when running under Windows

ted_dunning 2 hours ago

Attested TLS has had some rough patches lately which can be attributed to making big changes to a complex protocol.

It really better to separate the attestation, the check against policy and then the TLS stuff. Solve one problem at a time, sign that progress and move on.

jauntywundrkind 3 hours ago

Oh great, a new fresh hell against users, keeping them from being able to see the world or understand computing. Fantastic.

The War Against General Purpose Computing ticks on.

psanford 4 hours ago

I wish the author provided some latency numbers for this. One issue with tpms is that they are slow relative to performing the same operation on a modern CPU.

donavanm 4 hours ago

Thats the “what it costs” section? Im a bit impressed if they are down to ~3ms per handshake. When I last looked at TPM signing (many years ago) it was more like single digit transactions per second.

That said, even 3ms TPM signatures are going to be for special cases or novelty. Plain old CPU tls will do about 1ms cpu time per request which will scale by cpu core count. One or two orders of magnitude more throughput per host.

bschaatsbergen 4 hours ago

Author here. There's a benchmark table further down the post, the numbers come from this repo if you want to run them yourself: https://github.com/bschaatsbergen/go-tpm-tls-bench

ranger_danger 6 hours ago

Let's hope this doesn't get picked up by the (corporate) masses... the last thing I want is my browser offering personal TLS certificates to every server I visit as some kind of identity verification or fingerprint/tracking.

It's bad enough that ssh does this by default with all your keys.

altairprime 5 hours ago

Client TLS is rather unusable on the Internet by a typical random end user visiting a random public site, so that should at least keep the specific scenario you describe at bay.

ranger_danger 5 hours ago

Currently yes, but there's not much stopping Chrome etc. from adding a new feature that has a way of presenting a client certificate to a website in a backwards-compatible manner.

Of course the website itself would need to support that, but it's all possible in time.

altairprime 4 hours ago

Chrome would be more likely to implement a persistent and identifiable (to Google alone) tracking cookie replacement and ship it worldwide, which iirc they did — and then cancelled, of course. They seem to be focusing instead on improved tracking of Android users from the kernel up, rather than browsers from the headers down; GrapheneOS is, presumably, viewed as a serious threat to their advertising revenue.

https://privacysandbox.google.com/blog/update-on-plans-for-p...

_flux an hour ago

Wouldn't this in practice be a lot like Passkeys? But it might be more difficult to integrate this kind of approach to the stacks we use, whereas Passkeys fits in relatively easily.

I suppose client cert would protect against from a MitM attack, if the client failed to notice it, or if the MitMer has the website keys to make a perfect attack.

zx8080 4 hours ago

This is most probably where it's going in less than a year. The recent campaign "Safer with Google" in Chrome hints to this.