lexidma — archive
[Date Prev]
| [Thread Prev]
| [Thread Next]
| [Date Next]
— [Date Index]
| [Thread Index]
| [Month Index]
| [List Home]
Module-by-module proposal
OK, so, what I'm hearing from John and Iztok is support for the radical option:
- parts of speech, pronunciation, inflected forms and other formal properties belong at the entry level,
- semantics and pragmatics belong at the sense level,
- and there is to be no mixing between them. Any putative occurrence of sense-specific formal properties should trigger a reanalysis as homonymy. So, there would be two entries for "jeÅÃb", for example.
The more I mull it over in my head the less I'm opposed to it. So, if there are no more opinions one way or another, I'll rewrite the proposal like that. The flip side is that it will be leaner, simpler data model.
Give me a few days for the rewrite. I want to include a few other things I've heard here, namely: replace Label with a more specific types, and rephrase "attributes" as "properties" to avoid confusion with XML attributes.
M.
Iztok < [email protected]
> schrieb am Mi., 2. Juni 2021, 13:03:
My
comments
below .
Best
Iztok
From
:
[email protected] [mailto: [email protected] ] On Behalf
Of
Michal MÄchura
Sent
:
Tuesday , June 1, 2021 3:37 PM
To:
[email protected]
Subject
: Re: [ lexidma ] Module- by -module proposal
2/ I am not sure about
allowing
PartOfSpeech , Pronunciation
and
InflectedForm as children
of
Sense . Having
senses
of
an
entry
with
different part- of - speech
values is something
that some models
explicitly
avoid , we
would
also
need to figure out how this
inherits
from
the
Entry's
PartOfSpeech . I don't
think
we
should
have
Pronunciation
and
InflectedForm at all , as senses
with
different
pronunciations
or
inflections are homographs
and
we
really
should
insist
that
homographs are distinct at the
Entry
level .
I am not sure of it either , it does
smell
of
bad
practice : sloppy
separation
between form and
function . We
could go the
radical
route
and
prohibit it altogether . On the
other
hand it is what
lexicographers
sometimes
want to do. This
always
makes me think
of
the
Czech
word " jeÅÃb " ' crane
which
has
two
different
plurals
depending on whether
it's
the
animal
or
the
machine : I don't
think
lexicographers
would
welcome
the
idea
of
having to create
two separate entries
for
these . This is not an
exception , I could
probably
dig out other
examples
from
oher
languages . So the
question is: do we
want to force
lexicographers to re- analyze
all
sense-specific
morphosyntax as homonmy ?
I would
say so... but
what do others
think
here ?
If
we go that
way , then
we
wlll
have a data model which
makes
dictionaries more easily
machine-understandable ( because
of
the
clear
separation
between form and
meaning ) but
less human- friendly .
As an alternative proposal , how about
if
we
do
have
all
morphosyntax at the
entry-level (as you
wish ) but , in addition , we
invent some relational
mechanism
for
expressing
the
fact
that some of
the morphosyntactic properties ( such as this plural or
that
pronunciation ) only
apply to some senses
and not to others . Or
would
that make the data model even
less
machine-understandable
instead
of more?
[IK]
I would
agree
with John, we
should separate such
entries
with
different
pronunciations , plurals
etc . In general, I think
the model will " force " some lexicographers
into
many
things
they
might not be used too , and it is better to stay
consistent
with
what
we
think is best
practice .
J
3/ There
seems to be no way to record
properties
of
entries
such as noun
gender in the model.
That's
what
the
Label
object
type
would be for , in my
proposal .
4/ Pronunciation
probably
needs a scheme
and a variety
property . See
this
recent
paper (Sec 3.3) for a discussion
of
this : https://www.aclweb.org/anthology/2021.gwc-1.11.pdf
For
scheme , probably , yes . For
variety , if
by
that
you
mean
things like ' British '
versus ' American ', I enviaged
that , again , the
Label
object
type
would do that .
Okay ... so label is very
overloaded in terms
of
what it represents . My
feeling is I would
prefer some more specific
categories
for
common
annotations .
On second
thought , I agree
with
you : we
should
have
specific
types
for
specific
kinds
of
annotations , instead
of a catch-all
Label
type .
Why
am I changing
my
mind ? Because it makes
sense
that
all
corners
of
the data model should be on the same level
between
specific
and
abstract . If
we
have
Entry , Sense
and
SenseGroup
instead
of
just one abstract "Segment" type (as I was
proposing
earlier ), then
by
the same token
we
should
have ( for
example ) Register, Region
and Time instead
of
just one abstract
Label . The Segment and
Label
types are valid
abstraction
and
belong in a meta-model, but not belong in DMLex
because
DMLex
wants to be less "meta" and more immediately
implementable .
Also , I wasn't
even
consistent
with
myself
because
PartOfSpeech is really
just a Label
too .
[IK]
Perhaps
we
should be clear
what a label is
to me, it is an
indication
of some sort of
restriction
of
the
sense . Part of
speech
does not fall
under
this
category as it is the
key
distinguishing
info
separating
entries . Relatedly , I also
agree
we
should not have part of
speech , pronunciation
and
inflected form as children
of
Sense .
M.
[Date Prev]
| [Thread Prev]
| [Thread Next]
| [Date Next]
— [Date Index]
| [Thread Index]
| [Month Index]
| [List Home]