Next in thread → Next in month →

RE: [cti-taxii] Query Use Cases Needed! - Privacy Preserving Data Sharing

From
Chet Ensign <>
Date
2015-08-18T13:34:00+00:00
ID
Thread
RE: [cti-taxii] Query Use Cases Needed! - Privacy Preserving Data Sharing
Chris, Patrick, et al - Just want to make sure you are all aware of the privacy work being done at OASIS. In particular...
- OASIS Privacy by Design Documentation for Software Engineers (PbD-SE)

TC -  https://www.oasis-open.org/committees/pbd-se
- OASIS Privacy Management Reference Model (PMRM)

TC -  https://www.oasis-open.org/committees/pmrm
- OASIS Cross-Enterprise Security and Privacy Authorization (XSPA)

TC (this one specifically focused on health care use cases) -  https://www.oasis-open.org/committees/xspa
There are identity and trust groups as well. I'm happy to make introductions if anyone wants to learn more.
/chet
On Tue, Aug 18, 2015 at 8:34 AM, Patrick Maroney  < 
>  wrote:
Chris,
Great set of topics/ideas.  Need to fully digest same and review related papers. However, in terms of engagement, wanted to share some initial musings.
(1) The algorithmic costs need to be considered.  In other words, generation of the Ciphers and calculations on the back end (at scale) would need to be addressed.
(2)  Connecting dots would be harder as "I" have to know what compound questions to ask ahead of time (especially as it pertains to relationships and ) .  In other words, what happens if I don't ask the exact compound question?  in a simple "T/F" model  I could test 100 assertions, get 99/100 "right", but still get a "F" return.  As one continues the thought experiment, adding relationships and their attributes, temporal context, subjective measures/values, etc., some interesting things emerge.  In other,  other words some iteration through probabilistic outcomes should be envisioned  (i.e., "Your'e getting Warmer...").
(3)  I still see Source and Source Path Determinism (obfuscated where/as required) as key elements to success in any model where one is (1) trying to measure things like value, relevance, effectiveness, etc., (2) provide for RFIs, (3) provide a pathway  for sharing of sightings/observations/analysis/assertions/...
Patrick Maroney
_____________________________
From: Chris O'Brien < 
>  Sent: Tuesday, August 18, 2015 3:59 AM
Subject: RE: [cti-taxii] Query Use Cases Needed! - Privacy Preserving Data Sharing
To: < 
>
Hi all,

Sorry for the thread roll back - I've been thinking about the True/False use case, and I think there's an interesting avenue to explore here...

Combined with this use case:
https://github.com/TAXIIProject/TAXII-Specifications/wiki/TAXII-2.0-Use-Cases#government-wants-to-share-intel-but-no-one-to-know-it-was-them

...which is obviously close to my heart  : )  the True/False use case not only allows organisations to protect the 'who' but also potentially the 'what'. There's some great work going on in academia (eg:  http://arxiv.org/pdf/1502.05337.pdf ) that investigates the concept of privacy preserving data sharing, allowing organisations to share potentially sensitive data with others without actually revealing what that  data is (eg: storing an indicator in encrypted form, encrypting incoming indicators and comparing the cipher text to the stored indicator). The effect is similar to the True/False use case, but allows for a more peer-distributed set up. The advantage being  that, when it's identified, the indicator is already shared. Those who work in a classified environment may appreciate the elegance of, what is effectively, automated parallel evidence procedures - it's almost like a massive game of CTF (which would, similarly,  need to be really well locked down)! Governments (and other organisations with sensitive / classified data sources) are getting better at data sharing, but can always do more (us included). This might help remove some of the barriers.

There's a nice by-product of this use case too (as described in the UCL paper linked above) that organisations can mathematically estimate the value of data sharing. Apart from just proving that sharing data is a good thing (some people  still need to be reminded of that) it allows users to assess the 'value' of feeds based on how much they can tell the recipient that they don't already know / create links between objects / <insert your calculation of 'value' here>. With the anticipation of  having more feeds than a user has processing power, this could allow them to prioritise feeds based on empirical evidence rather than reputation.

Thoughts? This is something we're hoping to experiment with in CERT-UK against our Edge setup. If we can make that help the community then let me know!

Cheers,  Chris
Next in thread → Next in month →